Now liveThe Skillselion MCP - thousands of ranked skills, loaded into your agent mid-task. No install.Get it →
huaweicloud avatar

Huawei Cloud Ascendc Operator Performance Optim

  • 47 installs
  • 19 repo stars
  • Updated July 31, 2026
  • huaweicloud/huaweicloud-skills

Develop and optimize custom AscendC operators on Ascend NPU, analyzing bottlenecks and validating optimizations with the CANN toolkit.

About

Guides developing and optimizing custom operators in the AscendC language for Ascend 910B NPUs, using performance analysis, bottleneck identification, and optimization validation. A developer uses it when inference performance depends on tuning or writing operators for specific workloads.

  • Workflow: performance analysis to bottleneck ID to operator development to validation
  • Built on AscendC and CANN toolkit, validated with Ascend Profiler

Huawei Cloud Ascendc Operator Performance Optim by the numbers

  • 47 all-time installs (skills.sh)
  • +4 installs in the week ending Aug 2, 2026 (Skillselion tracking)
  • Ranked #945 of 2,101 Data Science & ML skills by installs in the Skillselion catalog
  • Data as of Aug 2, 2026 (Skillselion catalog sync)
npx skills add https://github.com/huaweicloud/huaweicloud-skills --skill huawei-cloud-ascendc-operator-performance-optim

Add your badge

Show developers this skill is listed on Skillselion. Paste this into your README.

Listed on Skillselion
Installs47
repo stars19
Last updatedJuly 31, 2026
Repositoryhuaweicloud/huaweicloud-skills

What it does

Develop and optimize custom AscendC operators on Ascend NPU, analyzing bottlenecks and validating optimizations with the CANN toolkit.

Files

SKILL.mdMarkdownGitHub ↗

Huawei Cloud AscendC Operator Performance Optimization

Overview

This skill provides guidance for developing and optimizing custom operators using AscendC programming language.

Architecture: Performance Analysis → Bottleneck Identification → Operator Development → Optimization → Validation

Related Skills:

  • huawei-cloud-ascend-profiler-db-explorer - Performance data analysis and bottleneck identification
  • huawei-cloud-ascend-small-model-migrate - Migration workflow that may require operator optimization

Architecture Components

This skill involves the following cloud services and components:

  • AscendC: Programming language for custom operator development
  • CANN: Huawei Cloud AI Computing Platform for NPU
  • Ascend 910B: Target NPU hardware for operator deployment
  • Ascend Profiler: Performance analysis tool for validation

Architecture Diagram:

┌─────────────────────────────────────────────────────────────┐
│            AscendC Operator Optimization Skill             │
├─────────────────────────────────────────────────────────────┤
│  ┌──────────────┐    ┌──────────────┐    ┌──────────────┐ │
│  │  Performance │───▶│  Bottleneck  │───▶│  Operator    │ │
│  │  Analysis    │    │  Identification│   │  Development │ │
│  └──────────────┘    └──────────────┘    └──────────────┘ │
│         │                   │                   │          │
│         ▼                   ▼                   ▼          │
│  ┌──────────────┐    ┌──────────────┐    ┌──────────────┐ │
│  │  Profiling   │    │  Optimization│    │  Validation │ │
│  │  Data        │    │  Techniques  │    │  & Testing  │ │
│  └──────────────┘    └──────────────┘    └──────────────┘ │
└─────────────────────────────────────────────────────────────┘

Use Cases

Typical Problem Scenarios:

  • Optimizing performance-critical operators on Ascend NPU
  • Developing custom operators for specific workloads
  • Improving model inference performance through operator optimization
  • Fixing operator bottlenecks identified during profiling
  • Implementing missing operators for NPU deployment

Typical User Phrases:

  • "Optimize my custom operator for Ascend"
  • "Develop AscendC operator for GEMM"
  • "Improve inference performance on NPU"
  • "Fix bottleneck operator"
  • "Implement custom operator using AscendC"
  • "AscendCOperator"
  • "OptimizationAscendOperatorPerformance"
  • "OperatorPerformance"

Scope

Supported:

  • Custom operator development in AscendC
  • Performance optimization for existing operators
  • Operator validation and testing

Not supported:

  • Non-AscendC operator development
  • Framework-level optimizations

Core Workflow

1. Performance Analysis

  • Use profiling tools to identify performance bottlenecks
  • Analyze operator execution time and resource utilization

2. Bottleneck Identification

  • Identify operators with high execution time
  • Determine optimization opportunities

3. Operator Development

  • Implement custom operators using AscendC
  • Follow AscendC best practices

4. Optimization Techniques

  • Memory optimization
  • Compute optimization
  • Data layout optimization

5. Validation

  • Verify functional correctness
  • Validate performance improvement

Reference Documents

DocumentDescription
Acceptance CriteriaFunctional acceptance criteria
Verification MethodVerification approach
TroubleshootingCommon issues and solutions

Prerequisites

  • CANN >= 7.0.0 installed
  • AscendC >= 1.0.0 installed
  • Ascend NPU driver installed and working properly
  • Operator code or performance data to be optimized

Core Commands

# Analyze operator performance bottlenecks
msprof --output=/path/to/output ./my_operator

# Optimize operator using AscendC
# Refer to CANN development guide for operator development

Parameter Confirmation

ParameterDescriptionRequired
Operator code pathOperator source code to be optimizedYes
Output directoryPerformance analysis result output pathYes
Optimization strategyPerformance optimization scheme selectionNo

Output Format

Performance analysis results are saved in the specified output directory:

output/
├── summary.json           # Performance summary
├── operator_stats.csv     # Operator execution statistics
├── timeline.json          # Execution timeline data
└── recommendations.md     # Optimization recommendations

Summary JSON Structure:

{
  "total_time_ms": 1234.56,
  "operator_count": 42,
  "top_operators": [
    {"name": "CustomGEMM", "time_ms": 456.78, "percentage": 37.0},
    {"name": "VectorAdd", "time_ms": 123.45, "percentage": 10.0}
  ],
  "optimization_candidates": ["CustomGEMM", "DataTransfer"]
}

Validation Method

Functional Validation

1. Run operator with test inputs 2. Compare outputs with reference implementation 3. Verify numerical accuracy (tolerance: 1e-5 for FP32, 1e-3 for FP16)

Performance Validation

1. Benchmark operator before optimization 2. Apply optimization changes 3. Benchmark operator after optimization 4. Calculate speedup ratio: speedup = time_before / time_after

Acceptance Criteria

  • Functional correctness: Output matches reference within tolerance
  • Performance improvement: Speedup >= 1.2x (20% improvement)
  • No regression: Other operators not affected

Best Practices

Memory Optimization

  • Use GM (Global Memory) for large tensors
  • Use L1/L0A/L0B for intermediate results in matrix operations
  • Align memory access to 32-byte boundaries
  • Reuse memory buffers when possible

Compute Optimization

  • Vectorize operations using AscendC intrinsics
  • Use MMA (Matrix Multiply Accumulate) for matrix operations
  • Parallelize independent operations
  • Minimize synchronization points

Data Layout Optimization

  • Use NZ format for matrix operations
  • Use ND format for vector operations
  • Avoid unnecessary format conversions
  • Consider memory coalescing for data access

Code Structure

  • Separate compute logic from memory operations
  • Use template metaprogramming for flexibility
  • Document optimization assumptions
  • Profile before and after each optimization

Notes

Common Pitfalls

  • Memory bank conflicts: Ensure data is distributed across memory banks
  • Unaligned access: Check 32-byte alignment for all buffers
  • Excessive synchronization: Minimize barrier usage between kernels
  • Wrong data format: Match format to operation type (NZ for matmul, ND for vector)

Performance Tips

1. Profile first to identify real bottlenecks 2. Focus on hot paths (operators with >10% total time) 3. Consider algorithmic changes before micro-optimizations 4. Test with realistic input sizes 5. Validate correctness after each optimization

Debugging Tips

  • Use ASCENDC_DEBUG=1 for verbose logging
  • Check CANN log files in /var/log/npu/
  • Compare with CPU reference implementation
  • Use msprof for detailed performance breakdown

Limitations

  • AscendC operators are hardware-specific (910B)
  • Not all PyTorch operators have AscendC equivalents
  • Custom operators require CANN recompilation for deployment

Related skills

This week in AI coding

Five minutes, every Monday - the tools, releases and tactics for developers.

unsubscribe anytime.