Watch this Intel keynote from Computex starting at 43:30 through 51:10. You will see a live demo that ran on a OCP AI network, which enabled the SambaNova RDUs to share KV cache memory with NVIDIA B200 GPUs to disaggregate pre-fill and decode. The result you see on the screen is low latency inference that is 3x faster than the NVIDIA B200 GPUs alone. https://www.youtube.com/live/1h_zY377urU?si=FLKaPhmHKT1KiaGw&t=2595
Watch this Intel keynote from Computex starting at 43:30 through 51:10. You will see a live demo that ran on a OCP AI network, which enabled the SambaNova RDUs to share KV cache memory with NVIDIA B200 GPUs to disaggregate pre-fill and decode. The result you see on the screen is low latency inference that is 3x faster than the NVIDIA B200 GPUs alone. https://www.youtube.com/live/1h_zY377urU?si=FLKaPhmHKT1KiaGw&t=2595
Watch OCP educational webinar on AI network reference architectures https://www.opencompute.org/events/past-events/ocp-educational-webinar-new-ocp-reference-architectures-for-ai-networking
Check out the OCP reference architectures for AI networking in open clusters https://www.opencompute.org/documents/open-cluster-designs-aligned-ai-inference-fabric-reference-architecture-pdf
https://substack.com/@rossboulton1/note/p-199734470?r=2leuaj&utm_medium=ios&utm_source=notes-share-action