<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Checkpointing on Rik Kisnah - Blog</title><link>https://www.rik-kisnah.ai/tags/checkpointing/</link><description>Recent content in Checkpointing on Rik Kisnah - Blog</description><generator>Hugo</generator><language>en</language><lastBuildDate>Tue, 07 Dec 2021 09:00:00 -0800</lastBuildDate><atom:link href="https://www.rik-kisnah.ai/tags/checkpointing/feed.xml" rel="self" type="application/rss+xml"/><item><title>Checkpointing at Scale</title><link>https://www.rik-kisnah.ai/teach/gpu-ai/checkpointing-at-scale/</link><pubDate>Tue, 07 Dec 2021 09:00:00 -0800</pubDate><guid>https://www.rik-kisnah.ai/teach/gpu-ai/checkpointing-at-scale/</guid><description>A thousand-GPU job will be interrupted. The only question is how much work you lose when it is. Save often enough that a failure costs minutes, cheaply enough that saving does not cost more than the failures.</description></item></channel></rss>