
Memory in the AI Era, Part 5: Exploring CXL Workloads
We examine how CXL is used in real workloads through KV cache offload, memory pooling, and memory tiering.

We examine how CXL is used in real workloads through KV cache offload, memory pooling, and memory tiering.

A hands-on walkthrough of Transformer-based LLM internals — from each module’s role to key optimization techniques.

We explore the technical principles behind NVIDIA’s ICMS — a new storage tier designed to solve the KV cache capacity bottleneck in LLMs — and the Bluefield-4 DPU that manages it.