<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>RoCE on Rik Kisnah - Blog</title><link>https://www.rik-kisnah.ai/tags/roce/</link><description>Recent content in RoCE on Rik Kisnah - Blog</description><generator>Hugo</generator><language>en</language><lastBuildDate>Tue, 15 Aug 2023 09:00:00 -0700</lastBuildDate><atom:link href="https://www.rik-kisnah.ai/tags/roce/feed.xml" rel="self" type="application/rss+xml"/><item><title>Design the Training Network</title><link>https://www.rik-kisnah.ai/teach/gpu-ai/design-the-training-network/</link><pubDate>Tue, 15 Aug 2023 09:00:00 -0700</pubDate><guid>https://www.rik-kisnah.ai/teach/gpu-ai/design-the-training-network/</guid><description>Two networks, not one. East-west is where GPUs talk to GPUs and must never drop a packet. North-south is everything else. Draw the rails, keep it lossless, and never let the two share a wire.</description></item><item><title>RDMA Explained</title><link>https://www.rik-kisnah.ai/teach/gpu-ai/rdma-explained/</link><pubDate>Tue, 10 Mar 2020 09:00:00 -0800</pubDate><guid>https://www.rik-kisnah.ai/teach/gpu-ai/rdma-explained/</guid><description>Normal networking hands every packet to the operating system, which copies it, thinks about it, and copies it again. RDMA lets one machine write straight into another machine&amp;rsquo;s memory with nobody in the middle. That is why GPU clusters use it.</description></item></channel></rss>