Jul 29, 2026

News

Lablup - FuriosaAI RNGD Whitepaper Released

  • Lablup

    Lablup

    Lablup

Jul 29, 2026

News

Lablup - FuriosaAI RNGD Whitepaper Released

  • Lablup

    Lablup

    Lablup

Lablup and FuriosaAI have published "RNGD meets Backend.AI," a whitepaper documenting the performance and operational efficiency of the RNGD and Backend.AI combination under real LLM workloads. Backend.AI has supported unified operation of diverse GPUs and AI accelerators from a single platform. In this whitepaper, Lablup and FuriosaAI present benchmark results covering throughput, latency, power efficiency, and multi-GPU scalability when operating large language models with FuriosaAI RNGD in a Backend.AI environment.

When RNGD Meets Backend.AI

Backend.AI is designed to manage a range of accelerators, including GPUs, NPUs, and TPUs, from a single platform. This whitepaper focuses specifically on the integration results with FuriosaAI RNGD. It explains how RNGD is reliably placed through Backend.AI's session-based execution environment and Sokovan orchestration, and how LLM workloads with mixed prefill and decode phases are operated efficiently.

RNGD is an NPU with 48 GB HBM3 and a 180 W TDP, targeting both the high memory bandwidth and power efficiency required for AI inference. The whitepaper shows what advantages emerge from a production operations perspective when these hardware characteristics are combined with Backend.AI's resource management and isolation capabilities.

RNGD Benchmark

Lablup and FuriosaAI configured a benchmark test using the Qwen3-32B FP8 model on an environment with four FuriosaAI RNGD cards. In the benchmark environment, RNGD demonstrated competitive throughput at lower power consumption compared to the reference configuration, with particular strengths in power efficiency and throughput as concurrency increased. RNGD also showed stable response characteristics at high concurrency levels in TTFT and TPOT measurements.

What the Whitepaper Covers

  • Backend.AI and RNGD integration architecture
  • FuriosaAI RNGD hardware characteristics and LLM suitability
  • Performance validation results based on Qwen3-32B FP8
  • Throughput, TTFT, and TPOT comparison by concurrency level
  • Analysis from the perspectives of power efficiency and multi-GPU operations

For a closer look at what the combination of Backend.AI and FuriosaAI RNGD means in practice, read the full whitepaper at the following link.

RNGD meets Backend.AI: AI Inference Infrastructure for the Era of Sovereign AI

RNGD meets Backend.AI: AI Inference Infrastructure for the Era of Sovereign AI

See how Backend.AI operates FuriosaAI's RNGD inference accelerator in a single control plane, and what matched-condition benchmarks show about throughput and power efficiency.

We're here for you!

Complete the form and we'll be in touch soon

Contact Us
lablup

Headquarter & HPC Lab

KR Office: 8F, 577, Seolleung-ro, Gangnam-gu, Seoul, 06143, Republic of Korea US Office: 3003 N First st, Suite 221, San Jose, CA 95134

  • facebook
  • youtube
  • Linkedin
  • GitHub

© Lablup Inc. All rights reserved.

We value your privacy

We use cookies to analyze site traffic, understand how visitors use our website, and improve our services. Necessary cookies for basic site functions are always active. Learn more

By clicking "Accept All", you agree to the storage of analytics cookies on your device. Click "Reject All" to keep only necessary cookies, or "Customize" to choose for yourself. You can change your settings at any time.