Get in Touch

Course Outline

Core Performance Concepts and Metrics

  • Key indicators: latency, throughput, power consumption, and resource usage
  • Distinguishing between system-wide and model-specific bottlenecks
  • Profiling strategies for inference versus training phases

Profiling with Huawei Ascend

  • Leveraging CANN Profiler and MindInsight for insights
  • Analyzing kernel and operator behavior
  • Understanding offload patterns and memory mapping

Profiling on Biren GPU

  • Utilizing Biren SDK for performance monitoring
  • Addressing kernel fusion, memory alignment, and execution queues
  • Implementing power and temperature-aware profiling

Profiling on Cambricon MLU

  • Using BANGPy and Neuware performance utilities
  • Gaining kernel-level visibility and interpreting logs
  • Integrating the MLU profiler with existing deployment frameworks

Graph and Model-Level Optimization

  • Strategies for graph pruning and quantization
  • Reconstructing computational graphs and fusing operators
  • Standardizing input sizes and tuning batch configurations

Memory and Kernel Optimization

  • Refining memory layout and data reuse
  • Managing buffers efficiently across different chipsets
  • Applying platform-specific kernel tuning techniques

Best Practices Across Platforms

  • Achieving performance portability through abstraction strategies
  • Creating shared tuning pipelines for multi-chip environments
  • Case study: Optimizing an object detection model across Ascend, Biren, and MLU

Wrap-up and Future Actions

Requirements

  • Hands-on experience with AI model training or deployment pipelines
  • Solid grasp of GPU/MLU compute mechanisms and model optimization techniques
  • Foundational knowledge of performance profiling tools and key metrics

Target Audience

  • Performance engineers
  • Machine learning infrastructure teams
  • AI system architects
 21 Hours

Number of participants


Price per participant

Upcoming Courses

Related Categories