ResourcesSolution brief

Backend.AI + VAST Data:
Line-Rate KV Cache Offloading

A benchmark look at line-rate KV cache offloading for heavy agent workloads

Backend.AI + VAST Data: Line-Rate KV Cache Offloading

See how Backend.AI sessions offload KV cache directly to VAST Data storage over GPUDirect Storage so agents re-enter long contexts without recomputing them, and what the benchmarks say about response time.

KV Cache Offloading That Halves TTFT for Agent Workloads

Coding agents and other workloads that revisit the same working context across turns reuse a long prefill holding a large codebase and tool-call context every turn. When multiple sessions take turns occupying one inference server, KV cache blocks in GPU memory get evicted and the next turn runs prefill again, stretching time to first token (TTFT). Perceived performance hinges less on raw token throughput than on how quickly a session re-enters the same context.

A Backend.AI session mounts the VAST KV cache folder through Storage Proxy, and vLLM with LMCache inside the session references the VAST mount directly. Using NVIDIA Magnum IO GPUDirect Storage, KV blocks move between GPU memory and storage without passing through host RAM, so contexts not in active use shift to the storage tier while GPU memory keeps only the active ones.

See the full brief for the measurement conditions, per-turn results, and guidance on which workloads benefit from offloading.

Related Services

Backend.AI

Backend.AI is a vendor-agnostic accelerated workload hosting platform based on our own home-grown orchestration and job scheduler, running on top of either cloud or on-premises (air-gapped) clusters.

Explore service →
VAST Data

VAST Data is the AI Operating System company. The VAST AI OS unifies data services, compute services, and agentic runtime into a single scalable platform, built on the DASE parallel distributed architecture to remove trade-offs between performance, scale, simplicity, and resilience.

Learn more →

We're here for you!

Complete the form and we'll be in touch soon

Contact Us
lablup

Headquarter & HPC Lab

KR Office: 8F, 577, Seolleung-ro, Gangnam-gu, Seoul, 06143, Republic of Korea US Office: 3003 N First st, Suite 221, San Jose, CA 95134

  • facebook
  • youtube
  • Linkedin
  • GitHub

© Lablup Inc. All rights reserved.

We value your privacy

We use cookies to analyze site traffic, understand how visitors use our website, and improve our services. Necessary cookies for basic site functions are always active. Learn more

By clicking "Accept All", you agree to the storage of analytics cookies on your device. Click "Reject All" to keep only necessary cookies, or "Customize" to choose for yourself. You can change your settings at any time.