<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Burn-In on Rik Kisnah - Blog</title><link>https://www.rik-kisnah.ai/tags/burn-in/</link><description>Recent content in Burn-In on Rik Kisnah - Blog</description><generator>Hugo</generator><language>en</language><lastBuildDate>Tue, 08 Sep 2026 09:00:00 -0700</lastBuildDate><atom:link href="https://www.rik-kisnah.ai/tags/burn-in/feed.xml" rel="self" type="application/rss+xml"/><item><title>GPU Burn-In at Scale: Commands, Gates, and Troubleshooting for NVIDIA and AMD Clusters</title><link>https://www.rik-kisnah.ai/posts/gpu-burn-in-at-scale-reference-guide/</link><pubDate>Tue, 08 Sep 2026 09:00:00 -0700</pubDate><guid>https://www.rik-kisnah.ai/posts/gpu-burn-in-at-scale-reference-guide/</guid><description>A plain-language operator guide to burn-in testing GPU clusters: the L10, L11 and L12 levels, east-west and north-south networks, the commands that find bad parts, and the 2026 NVIDIA and AMD hardware you will meet.</description></item><item><title>Design a Burn-In Pipeline</title><link>https://www.rik-kisnah.ai/teach/gpu-ai/design-a-burn-in-pipeline/</link><pubDate>Tue, 09 Apr 2024 09:00:00 -0700</pubDate><guid>https://www.rik-kisnah.ai/teach/gpu-ai/design-a-burn-in-pipeline/</guid><description>Brand new hardware fails. Find the bad parts on your time, not the customer&amp;rsquo;s. Test one GPU, then one server, then one rack, then the whole cluster, with a gate at every step. Fail fast, log everything, never let a stage run without a stop rule.</description></item></channel></rss>