Dell and NVIDIA Expand the Horizons of AI Inference

Discover how Dell & NVIDIA redefine AI inference with KV Cache offloading, boosting speed, efficiency, and scalability for LLMs.

Key takeaways: Dell and NVIDIA are driving the next evolution of AI inference, now advancing KV Cache with innovations like the Context Memory Storage Platform (CMS) and the NVIDIA BlueField-4 data processing unit (DPU). This collaboration enables faster, more efficient processing for Large Language Models (LLMs), helping organizations optimize speed, reduce latency, and improve cost efficiency. Dell’s high-performance storage solutions, including Dell PowerScale, Dell ObjectScale, and Project Lightning, are engineered to support these advancements, providing the flexible foundation needed for current and future AI workloads. Together, Dell and NVIDIA are building the infrastructure to power the next generation of AI innovation.


Artificial Intelligence is advancing rapidly, with Large Language Models (LLMs) becoming increasingly intelligent and complex. For organizations deploying these models, the challenge often shifts from training to agentic inference, delivering fast, context-aware responses while optimizing infrastructure and accelerating token generation. A key solution to this challenge is Key-Value (KV) Cache offloading.

When an LLM processes a prompt, it generates “attention” data, Keys and Values, that help it understand inference context. Storing this data in the GPU’s high-bandwidth memory (HBM) enables quick token generation, a process known as KV Caching. However, as conversation history or document length grows, the cache expands, forcing costly recomputation when it can no longer be kept in GPU memory. This bottleneck slows response times and increases power consumption.

The answer lies in offloading KV Cache to more abundant resources, freeing GPUs to focus on computation. NVIDIA BlueField-4 and Dell Technologies provide the performance and scalability needed to tackle these challenges, ensuring efficient AI inference at scale.

Introducing Context Memory Storage Platform with NVIDIA BlueField-4

NVIDIA’s latest AI advancement, the NVIDIA BlueField-4 data processor, brings the concept of CMS to the forefront. It’s a dedicated memory tier designed to handle the growing demands of the “reasoning reservoir” in AI workloads. Dell is developing storage solutions that are purpose-built to complement and fully leverage the data processor’s capabilities for CMS to further accelerate inference.

NVIDIA BlueField-4 optimizes KV caching with a specialized acceleration engine, bridging the gap between fast, but limited GPU memory and traditional storage to further accelerate inference performance.

Key Benefits of NVIDIA BlueField-4 for KV Cache

    • Optimize GPU Utilization and Throughput: The data processor is designed to optimize data paths, reduce stalls and recomputation, improving throughput and utilization for long-thinking inference.
    • Accelerate Agentic Inference: For active reasoning and real-time conversations, every millisecond counts. The data processor’s low latency improves responsiveness and minimizes the time it takes to fetch cached context.
    • Improve Power Efficiency: By optimizing data movement, the solution improves performance per watt, making it a sustainable choice for scaling AI factories.

Scalable performance for every architecture

Dell’s proven storage and data management expertise ensures customers don’t have to wait for tomorrow’s hardware to see massive performance and efficiency gains right now.

Dell is dedicated to supporting the latest NVIDIA AI innovations by developing storage solutions designed to unlock and amplify the capabilities of NVIDIA BlueField-4. Our goal is to help organizations harness the full power of this new platform for a wide range of AI workloads, providing seamless, scalable infrastructure that builds on NVIDIA’s advancements.

At the same time, for organizations today, Dell already delivers highly scalable, high-performance KV Cache offload solutions that seamlessly accelerate inference, delivering a 19x improvement in TTFT (Time to First Token) and provide up to 5.3x improvement in the number of queries per second.

Dell’s approach: Flexibility meets performance

For environments without NVIDIA BlueField-4, or for those needing to scale capacity into the petabyte range, Dell provides a robust software stack. By integrating technologies like LMCache and NVIDIA NIXL (NVIDIA Inference Transfer Library) with our industry-leading AI storage engines, we turn your storage infrastructure into a high-speed extension of your GPU memory.

This solution allows the KV Cache to be offloaded efficiently to Dell’s file or object storage engines using RDMA technology, bypassing the CPU to maintain high-speed data flow.

The power of choice: Dell AI storage engines

We support this offloading capability across our diverse portfolio, giving you the freedom to choose the right storage for your specific needs:

    1. Dell PowerScale: Ideal for those who need the simplicity of NAS with high-performance parallel access. Using NFS-over-RDMA, PowerScale delivers low-latency access to massive amounts of cached data.
    2. Dell ObjectScale: For organizations building cloud-native applications, ObjectScale provides high-performance object storage. With our unique S3-over-RDMA technology, you get the scalability of object storage with the speed typically reserved for file systems.
    3. Project Lightning (in private preview): For the most demanding workloads, our breakthrough parallel file system, designed for the AI era, leverages NVMe-over-Fabrics to transfer data directly from drives to GPU memory, minimizing latency and maximizing throughput.

Why this matters for your business

The ability to offload KV Cache effectively—whether through a specialized DPU like NVIDIA BlueField-4 or a scalable storage engine—transforms the economics of AI.

    • Cost Efficiency: You no longer need to buy more expensive GPUs just to get more memory. You can expand your context window using cost-effective storage or specialized DPUs.
    • Enhanced User Experience: By retaining long context windows, your AI models can remember more of the conversation, summarize larger documents, and provide more accurate, personalized responses even for long, multi-turn conversations across multiple user sessions.
    • Future-Proofing: As models grow to support millions of tokens in a single prompt, the KV Cache will become too large for any single server, increasing the importance of distributed inference. Dell’s scalable storage solutions ensure that your infrastructure can grow alongside your AI ambitions.

Advancing AI infrastructure together

At Dell Technologies, we believe in the power of an open ecosystem. By collaborating with NVIDIA, we are building a comprehensive AI Factory that empowers you to innovate faster.

Whether you are deploying NVIDIA BlueField-4 for ultra-low latency context extension or leveraging the massive scalability of Dell PowerScale, ObjectScale, and Lightning for your enterprise inference stack, we have the solutions to help you move forward.

The future of AI is context-aware, fast, and efficient. Together, we are building the infrastructure that makes it possible.

About the Author: Rajesh Rajaraman

Rajesh Rajaraman is responsible for technology strategy, architecture, and innovation across the portfolio.   Proven technology leader with over 25 years of experience in storage and distributed systems. He possesses tremendous breadth and technical depth in storage, data protection, public cloud technologies and has authored many patents. He also holds an engineering degree in electronics and communication. Prior to joining Dell, he held positions at NetApp, Cohesity and DEC.