Rik Kisnah - Blog

#GPU

Design an Inference API

Design an Inference API

OCI Powers America's AI Future at NVIDIA GTC 2025: Supercomputers, AI Factories, and Strategic Leadership

OCI Powers America's AI Future at NVIDIA GTC 2025: Supercomputers, AI Factories, and Strategic Leadership

From First Principles to Zettascale: How OCI's GPU/RDMA Architecture Redefines AI Infrastructure

From First Principles to Zettascale: How OCI's GPU/RDMA Architecture Redefines AI Infrastructure

Mixture of Experts and All-to-All Traffic

Mixture of Experts and All-to-All Traffic

Continuous Batching for Inference

Continuous Batching for Inference

Size the KV Cache

Size the KV Cache

Power and Cooling for GPU Racks

Power and Cooling for GPU Racks

Design a Burn-In Pipeline

Design a Burn-In Pipeline

Straggler Detection in Training Jobs

Straggler Detection in Training Jobs

Measure GPU Utilisation and MFU

Measure GPU Utilisation and MFU