Backend.AI + Intel Gaudi 3:
Benchmarks and Deep Integration
Comprehensive whitepaper on Backend.AI and Intel Gaudi 3 AI Accelerators performance
A comprehensive whitepaper with detailed benchmarks showing how Backend.AI and Intel Gaudi 3 deliver remarkable performance for both small and large language model inference workloads.
Operating Gaudi 3 Infrastructure for NeoCloud Providers
Intel Gaudi 3 is an AI accelerator built on a 5nm process, carrying 8 MMEs, 64 TPCs, 128GB of HBM2e memory, and 3.7TB/s of memory bandwidth, offered in both OAM and PCIe form factors. Backend.AI operates Gaudi 3 clusters as a multi-tenant environment with card-level accelerator allocation, NUMA-aware resource scheduling, external storage integration, and per-user quota management. Single-card inference sessions and multi-node distributed training jobs are placed by the same scheduler, so how you operate does not change as workload scale does.
The whitepaper includes benchmarks run with the Llama 3.1 8B and 70B models, organizing results for small and large models along with deployment recommendations and optimization strategies. See the full whitepaper for the measurement conditions and detailed per-model results.
Download Resource
Please fill out the form below.
Related Services
Backend.AI is a vendor-agnostic accelerated workload hosting platform based on our own home-grown orchestration and job scheduler, running on top of either cloud or on-premises (air-gapped) clusters.
Explore service →
Bringing choice to gen AI with performance, scalability and efficiency. Meet your new high-performance option built to handle your AI workloads, your way.
Learn more →