Discussion about this post

User's avatar
Marc Austin's avatar

Watch this Intel keynote from Computex starting at 43:30 through 51:10. You will see a live demo that ran on a OCP AI network, which enabled the SambaNova RDUs to share KV cache memory with NVIDIA B200 GPUs to disaggregate pre-fill and decode. The result you see on the screen is low latency inference that is 3x faster than the NVIDIA B200 GPUs alone. https://www.youtube.com/live/1h_zY377urU?si=FLKaPhmHKT1KiaGw&t=2595

2 more comments...

No posts

Ready for more?