RNGD meets Backend.AI:
AI Inference Infrastructure for the Era of Sovereign AI
A benchmark look at RNGD on Backend.AI
Purpose-built inference accelerators are designed to handle the same serving work with less power, lowering the operating cost of always-on inference services. See how Backend.AI operates FuriosaAI's RNGD inference accelerator in a single control plane, and what matched-condition benchmarks show about throughput and power efficiency.
LLM Inference Infrastructure Built on a Korean-Designed NPU
FuriosaAI's RNGD is a datacenter inference accelerator (NPU) with a TDP of 180W per card, trading the general-purpose compute needed for training for a design optimized for inference-stage latency, throughput, and power efficiency. The combination of an inference chip designed in Korea and an open-source-based operations platform developed in Korea answers the practical requirements of sovereign AI: air-gapped operation, supply chain options free of overseas vendor lock-in, and code-level verifiability.
In a benchmark serving the Qwen3-32B model at FP8 precision, four RNGD cards recorded about 95% of the throughput of four latest-generation GPUs of comparable class at 256 concurrent requests, while total accelerator power stayed at 56-70% of the GPU setup. Converted to throughput per watt, RNGD came out about 1.3-1.5x higher across every measured concurrency level, and time to first token (TTFT) was also shorter across the board.
Backend.AI abstracts RNGD as a session-level resource just like a GPU through its accelerator plugin architecture. Allocation, isolation, model serving, and lifecycle management work the same way as an existing GPU cluster, so GPUs and NPUs can run side by side in one cluster without learning a new operations toolchain.
See the full whitepaper for the measurement environment, detailed results by concurrency level, workload scenarios, and the environments where adoption is recommended.
Related Services
Backend.AI is a vendor-agnostic accelerated workload hosting platform based on our own home-grown orchestration and job scheduler, running on top of either cloud or on-premises (air-gapped) clusters.
Explore service →FuriosaAI designs data-center AI accelerators purpose-built for inference. RNGD, built on the Tensor Contraction Processor architecture, serves large language models at high power efficiency with a 180W TDP per card.
Learn more →