Get in Touch

Course Outline

Introduction to Biren GPU Architecture

  • Overview of Biren and its primary use cases
  • Hardware components: cores, memory structures, and compute clusters
  • Comparative analysis with NVIDIA and AMD GPU architectures

Establishing the Biren Programming Environment

  • Installation of the Biren SDK and runtime components
  • Insights into the toolchain and compiler model
  • Understanding basic project structures and build workflows

GPU Programming via the Biren Stack

  • Exploring thread and block models
  • Managing memory and handling data transfers
  • Developing kernels and establishing launch patterns

Migrating from CUDA to Biren

  • Techniques for translating CUDA codebases
  • Mapping common APIs and making necessary adaptations
  • Practical labs and exercises for code conversion

Debugging and Profiling Strategies

  • Utilizing Biren’s debugger and profiling tools
  • Detecting performance bottlenecks
  • Optimizing memory access patterns

Advanced Optimization Techniques

  • Thread scheduling and instruction pipelining
  • Loop unrolling and efficient use of shared memory
  • Fine-tuning kernels for maximum throughput

Case Studies and Application Examples

  • Training models using Biren accelerators
  • Porting and profiling vision or NLP models
  • Benchmarking performance against CUDA/NVIDIA solutions

Summary and Future Steps

Requirements

  • A solid understanding of GPU architecture and parallel processing concepts
  • Practical experience with CUDA, OpenCL, or comparable GPU programming frameworks
  • Proficiency with deep learning frameworks like PyTorch or TensorFlow

Target Audience

  • HPC developers
  • AI infrastructure engineers
  • Performance optimization specialists
 21 Hours

Number of participants


Price per participant

Testimonials (2)

Upcoming Courses

Related Categories