Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Core Performance Concepts and Metrics
- Key indicators: latency, throughput, power consumption, and resource usage
- Distinguishing between system-wide and model-specific bottlenecks
- Profiling strategies for inference versus training phases
Profiling with Huawei Ascend
- Leveraging CANN Profiler and MindInsight for insights
- Analyzing kernel and operator behavior
- Understanding offload patterns and memory mapping
Profiling on Biren GPU
- Utilizing Biren SDK for performance monitoring
- Addressing kernel fusion, memory alignment, and execution queues
- Implementing power and temperature-aware profiling
Profiling on Cambricon MLU
- Using BANGPy and Neuware performance utilities
- Gaining kernel-level visibility and interpreting logs
- Integrating the MLU profiler with existing deployment frameworks
Graph and Model-Level Optimization
- Strategies for graph pruning and quantization
- Reconstructing computational graphs and fusing operators
- Standardizing input sizes and tuning batch configurations
Memory and Kernel Optimization
- Refining memory layout and data reuse
- Managing buffers efficiently across different chipsets
- Applying platform-specific kernel tuning techniques
Best Practices Across Platforms
- Achieving performance portability through abstraction strategies
- Creating shared tuning pipelines for multi-chip environments
- Case study: Optimizing an object detection model across Ascend, Biren, and MLU
Wrap-up and Future Actions
Requirements
- Hands-on experience with AI model training or deployment pipelines
- Solid grasp of GPU/MLU compute mechanisms and model optimization techniques
- Foundational knowledge of performance profiling tools and key metrics
Target Audience
- Performance engineers
- Machine learning infrastructure teams
- AI system architects
21 Hours