ResourcesSolution brief

Powerful infra management with Backend.AI® on Tenstorrent Wormhole™ AI accelerator

Lablup Backend.AI links up with Tenstorrent Wormhole

Powerful infra management with Backend.AI® on Tenstorrent Wormhole™ AI accelerator

Meet the combination of Lablup Backend.AI® and Tenstorrent Wormhole™, enabling enterprises to seamlessly develop, deploy, manage, and scale generative AI solutions.

Running LLM Inference on Tenstorrent Wormhole

Modern AI workloads demand massive parallelism, high memory bandwidth, and low latency all at once, and purpose-built AI chips are designed around those demands from the start. Drawing the full value out of an accelerator, though, takes an operations layer on top of it. Allocating resources dynamically to raise utilization, preventing resource contention, and guaranteeing fair scheduling across users and projects is the job of orchestration software, not the hardware.

Tenstorrent Wormhole is a PCIe card AI accelerator built on high-performance Tensix Cores. Each card integrates compute units, a network-on-chip, local caches, and RISC-V cores to move data efficiently inside the chip, and ships in two models, n150s and n300s. Backend.AI's Sokovan orchestrator handles accelerator allocation, storage mounts, and network connectivity at scheduling time, so training and inference serving run without separate setup work in on-premise and air-gapped environments.

The brief includes performance benchmarks run with the Llama 3.1 8B model on a Wormhole LoudBox. It organizes results and recommended configurations for three usage scenarios: real-time chatbots and assistants, RAG-based question answering, and long-form content generation, laying out operating baselines that fit each workload. See the full brief for the measurement conditions and per-scenario recommendations.

Related Services

Backend.AI

Backend.AI is a vendor-agnostic accelerated workload hosting platform based on our own home-grown orchestration and job scheduler, running on top of either cloud or on-premises (air-gapped) clusters.

Explore service
Tenstorrent

Tenstorrent designs high-performance AI accelerators with the Wormhole architecture, delivering massive parallel processing and low latency for enterprise AI workloads.

Learn more

We're here for you!

Complete the form and we'll be in touch soon

Contact Us
lablup

Headquarter & HPC Lab

KR Office: 8F, 577, Seolleung-ro, Gangnam-gu, Seoul, 06143, Republic of Korea US Office: 3003 N First st, Suite 221, San Jose, CA 95134

  • facebook
  • youtube
  • Linkedin
  • GitHub

© Lablup Inc. All rights reserved.

We value your privacy

We use cookies to analyze site traffic, understand how visitors use our website, and improve our services. Necessary cookies for basic site functions are always active. Learn more

By clicking "Accept All", you agree to the storage of analytics cookies on your device. Click "Reject All" to keep only necessary cookies, or "Customize" to choose for yourself. You can change your settings at any time.