Powerful infra management with Backend.AI® on Tenstorrent Wormhole™ AI accelerator
Lablup Backend.AI links up with Tenstorrent Wormhole
Meet the combination of Lablup Backend.AI® and Tenstorrent Wormhole™, enabling enterprises to seamlessly develop, deploy, manage, and scale generative AI solutions.
Running LLM Inference on Tenstorrent Wormhole
Modern AI workloads demand massive parallelism, high memory bandwidth, and low latency all at once, and purpose-built AI chips are designed around those demands from the start. Drawing the full value out of an accelerator, though, takes an operations layer on top of it. Allocating resources dynamically to raise utilization, preventing resource contention, and guaranteeing fair scheduling across users and projects is the job of orchestration software, not the hardware.
Tenstorrent Wormhole is a PCIe card AI accelerator built on high-performance Tensix Cores. Each card integrates compute units, a network-on-chip, local caches, and RISC-V cores to move data efficiently inside the chip, and ships in two models, n150s and n300s. Backend.AI's Sokovan orchestrator handles accelerator allocation, storage mounts, and network connectivity at scheduling time, so training and inference serving run without separate setup work in on-premise and air-gapped environments.
The brief includes performance benchmarks run with the Llama 3.1 8B model on a Wormhole LoudBox. It organizes results and recommended configurations for three usage scenarios: real-time chatbots and assistants, RAG-based question answering, and long-form content generation, laying out operating baselines that fit each workload. See the full brief for the measurement conditions and per-scenario recommendations.
Related Services
Backend.AI is a vendor-agnostic accelerated workload hosting platform based on our own home-grown orchestration and job scheduler, running on top of either cloud or on-premises (air-gapped) clusters.
Explore service →Tenstorrent designs high-performance AI accelerators with the Wormhole architecture, delivering massive parallel processing and low latency for enterprise AI workloads.
Learn more →