Jul 29, 2026
News
Lablup - FuriosaAI RNGD Whitepaper Released
Lablup
Lablup
Jul 29, 2026
News
Lablup - FuriosaAI RNGD Whitepaper Released
Lablup
Lablup
Lablup and FuriosaAI have published "RNGD meets Backend.AI," a whitepaper documenting the performance and operational efficiency of the RNGD and Backend.AI combination under real LLM workloads. Backend.AI has supported unified operation of diverse GPUs and AI accelerators from a single platform. In this whitepaper, Lablup and FuriosaAI present benchmark results covering throughput, latency, power efficiency, and multi-GPU scalability when operating large language models with FuriosaAI RNGD in a Backend.AI environment.
When RNGD Meets Backend.AI
Backend.AI is designed to manage a range of accelerators, including GPUs, NPUs, and TPUs, from a single platform. This whitepaper focuses specifically on the integration results with FuriosaAI RNGD. It explains how RNGD is reliably placed through Backend.AI's session-based execution environment and Sokovan orchestration, and how LLM workloads with mixed prefill and decode phases are operated efficiently.
RNGD is an NPU with 48 GB HBM3 and a 180 W TDP, targeting both the high memory bandwidth and power efficiency required for AI inference. The whitepaper shows what advantages emerge from a production operations perspective when these hardware characteristics are combined with Backend.AI's resource management and isolation capabilities.
RNGD Benchmark
Lablup and FuriosaAI configured a benchmark test using the Qwen3-32B FP8 model on an environment with four FuriosaAI RNGD cards. In the benchmark environment, RNGD demonstrated competitive throughput at lower power consumption compared to the reference configuration, with particular strengths in power efficiency and throughput as concurrency increased. RNGD also showed stable response characteristics at high concurrency levels in TTFT and TPOT measurements.
What the Whitepaper Covers
- Backend.AI and RNGD integration architecture
- FuriosaAI RNGD hardware characteristics and LLM suitability
- Performance validation results based on Qwen3-32B FP8
- Throughput, TTFT, and TPOT comparison by concurrency level
- Analysis from the perspectives of power efficiency and multi-GPU operations
For a closer look at what the combination of Backend.AI and FuriosaAI RNGD means in practice, read the full whitepaper at the following link.

RNGD meets Backend.AI: AI Inference Infrastructure for the Era of Sovereign AI
See how Backend.AI operates FuriosaAI's RNGD inference accelerator in a single control plane, and what matched-condition benchmarks show about throughput and power efficiency.