<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Fleet Health on Rik Kisnah - Blog</title><link>https://www.rik-kisnah.ai/tags/fleet-health/</link><description>Recent content in Fleet Health on Rik Kisnah - Blog</description><generator>Hugo</generator><language>en</language><lastBuildDate>Tue, 09 Jun 2026 09:00:00 -0700</lastBuildDate><atom:link href="https://www.rik-kisnah.ai/tags/fleet-health/feed.xml" rel="self" type="application/rss+xml"/><item><title>Design a Fleet Health Score and Repair Loop</title><link>https://www.rik-kisnah.ai/teach/gpu-ai/design-a-fleet-health-score/</link><pubDate>Tue, 09 Jun 2026 09:00:00 -0700</pubDate><guid>https://www.rik-kisnah.ai/teach/gpu-ai/design-a-fleet-health-score/</guid><description>Thousands of GPUs, each with fifty signals. Boil them into one number per node that says &amp;lsquo;schedule on me&amp;rsquo; or &amp;lsquo;do not&amp;rsquo;, and a loop that takes sick nodes out, fixes them, proves they are fixed, and puts them back. The metric that matters is usable compute.</description></item><item><title>Design a Burn-In Pipeline</title><link>https://www.rik-kisnah.ai/teach/gpu-ai/design-a-burn-in-pipeline/</link><pubDate>Tue, 09 Apr 2024 09:00:00 -0700</pubDate><guid>https://www.rik-kisnah.ai/teach/gpu-ai/design-a-burn-in-pipeline/</guid><description>Brand new hardware fails. Find the bad parts on your time, not the customer&amp;rsquo;s. Test one GPU, then one server, then one rack, then the whole cluster, with a gate at every step. Fail fast, log everything, never let a stage run without a stop rule.</description></item><item><title>Straggler Detection in Training Jobs</title><link>https://www.rik-kisnah.ai/teach/gpu-ai/straggler-detection/</link><pubDate>Tue, 16 Jan 2024 09:00:00 -0800</pubDate><guid>https://www.rik-kisnah.ai/teach/gpu-ai/straggler-detection/</guid><description>In a synchronous job, the whole cluster runs at the speed of its slowest GPU. One card throttling at 80 percent makes a thousand cards run at 80 percent. Find it in minutes, automatically, and take it out.</description></item><item><title>GPU Faults: XID, ECC, and What They Mean</title><link>https://www.rik-kisnah.ai/teach/gpu-ai/gpu-fault-taxonomy/</link><pubDate>Tue, 13 Apr 2021 09:00:00 -0700</pubDate><guid>https://www.rik-kisnah.ai/teach/gpu-ai/gpu-fault-taxonomy/</guid><description>A GPU tells you it is unhappy in a few specific ways. Some mean &amp;lsquo;retry&amp;rsquo;. Some mean &amp;lsquo;reboot&amp;rsquo;. Some mean &amp;rsquo;this card goes back in the box&amp;rsquo;. The job is to read the message and sort it into the right bin fast, without a human staring at logs.</description></item></channel></rss>