<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Reliability on Rik Kisnah - Blog</title><link>https://www.rik-kisnah.ai/tags/reliability/</link><description>Recent content in Reliability on Rik Kisnah - Blog</description><generator>Hugo</generator><language>en</language><lastBuildDate>Tue, 10 Feb 2026 09:00:00 -0800</lastBuildDate><atom:link href="https://www.rik-kisnah.ai/tags/reliability/feed.xml" rel="self" type="application/rss+xml"/><item><title>Design an On-Call Paging System</title><link>https://www.rik-kisnah.ai/teach/systems/design-an-on-call-paging-system/</link><pubDate>Tue, 10 Feb 2026 09:00:00 -0800</pubDate><guid>https://www.rik-kisnah.ai/teach/systems/design-an-on-call-paging-system/</guid><description>When the alert fires, wake exactly the right person, make sure someone acknowledges it, and escalate if nobody does. The system that pages must be the last thing standing when everything else is down, so it must not depend on anything else.</description></item><item><title>Checkpointing at Scale</title><link>https://www.rik-kisnah.ai/teach/gpu-ai/checkpointing-at-scale/</link><pubDate>Tue, 07 Dec 2021 09:00:00 -0800</pubDate><guid>https://www.rik-kisnah.ai/teach/gpu-ai/checkpointing-at-scale/</guid><description>A thousand-GPU job will be interrupted. The only question is how much work you lose when it is. Save often enough that a failure costs minutes, cheaply enough that saving does not cost more than the failures.</description></item></channel></rss>