Engineering

Serving Solar Open 2 with Two DGX Spark Systems
By Kyujin Cho, Jinho HeoWe share how we enabled Solar Open 2 NVFP4 on dual DGX Sparks with vLLM by resolving FlashInfer b12x Expert Parallelism and checkpoint limitations.24 July 2026

Is Korean really a low-resource language?
By Wonik Cho, Youngsook SongThe claim that Korean is "low-resource" is only half true: the data is not missing, just scattered, closed off, and poorly known. Drawing on the Open Korean Corpora report, this article sorts 100 open Korean corpora into 10 categories and marks each by three criteria—documentation, usage, and redistribution. The result is a single map that lets anyone see, at a glance, what Korean data exists and how it can actually be used.1 July 2026

Backend.AI on DGX Spark: Open source installation guide
By Kyujin Cho and 3 othersWe provide a step-by-step tutorial on how to install the open-source version of Backend.AI on NVIDIA DGX Spark and easily launch a model.22 June 2026

The era of diverse AI accelerators: Backend.AI's heterogeneous GPU operation strategy
By Jinho HeoThe AI accelerator market is diversifying across NVIDIA, AMD, Intel, and domestic NPUs, enabling heterogeneous GPU operations. Backend.AI solves software stack and performance challenges by separating resources into groups and applying workload-specific policies to maximize each accelerator's strengths.19 June 2026

VLM Safety Failures: Why safe scenes get flagged as dangerous
By Youngsook Song, Dasol ChoiVision-language models detect real emergencies reasonably well, but they often overreact by misclassifying safe scenes as dangerous. This post shares results from the VERI benchmark, which measures this limitation using visually similar emergency and safe scenes with different meanings.18 June 2026

Agent coding at long context: What KV cache offloading on VAST Data & Backend.AI buys you
By Jinho Heo and 2 othersA joint benchmark by Lablup and VAST Data shows that KV cache offloading improves TTFT by up to 3.3x and reduces overall latency by more than half in agent-based coding workloads.16 June 2026

Intel Arc meets Backend.AI: What the Arc Pro B70's 32GB memory buys for agentic AI
By Jinho Heo and 2 othersBackend.AI now officially supports Intel Arc Pro B70, expanding its Intel lineup beyond Gaudi 2/3 AI accelerators to Arc graphics. From datacenter Gaudi to workstation Arc Pro, teams can manage Intel AI hardware in one platform.12 June 2026

eBPF security audit for GPU clusters with Backend.AI
By Kyujin Cho, Jinho HeoThis blog introduces eBPF, a method for implementing security audits while maintaining the performance of GPU training clusters. It covers the principles of eBPF, which directly processes events within the kernel, audit application cases, and security precautions.27 May 2026

How to save GPU memory in LLM serving: Principles and operating conditions of KV cache offloading
By Kyujin Cho, Jinho HeoHow KV cache offloading works in LLM serving for agentic AI: the architecture, data paths, and when offloading actually helps inference performance.27 April 2026

Building Production RAG Systems: Lessons from Tariff Support
By Sergey LeksikovLablup's research team shares lessons from two production RAG systems, the HSense tariff classifier (92.4% Top-1) and a Backend.AI support assistant, including why retrieval quality matters more than model choice.23 April 2026

Writing Stories for 50 Components: Foundation, Automation, and AI
By Seunghyun LimWriting Storybook stories for 50+ Backend.AI WebUI components: setting up i18n, theming, and branding, then automating with a 1,000-line guideline, Claude-based generation, and GitHub Actions CI.5 March 2026

The Pulse of 500+ GPUs: Monitoring Large-Scale AI Training Clusters
By Hanjeong LeeHow Lablup built proactive fault detection for a 500+ NVIDIA B200 GPU cluster while supporting Solar Open 102B pretraining, including its data collection and real anomaly cases.20 February 2026