Rik Kisnah - Blog

#GPU

Design the Training Network

Design the Training Network

Diagnose a Slow NCCL Job

Diagnose a Slow NCCL Job

Design a Topology-Aware GPU Scheduler

Design a Topology-Aware GPU Scheduler

All-Reduce Explained

All-Reduce Explained

Number Formats: FP32 to FP8

Number Formats: FP32 to FP8

Checkpointing at Scale

Checkpointing at Scale

Why GPUs for AI

Why GPUs for AI

GPU Faults: XID, ECC, and What They Mean

GPU Faults: XID, ECC, and What They Mean

Data, Model, and Pipeline Parallelism

Data, Model, and Pipeline Parallelism

RDMA Explained

RDMA Explained