Sep 28, 2026

Engineering

The Age of Metered Intelligence: From Data Centers to Token Factories

  • Jinho Heo

    Jinho Heo

    Technical Writer

Sep 28, 2026

Engineering

The Age of Metered Intelligence: From Data Centers to Token Factories

  • Jinho Heo

    Jinho Heo

    Technical Writer

For a long time, we have relied on standardized units of measurement. One of the most familiar examples comes from the horse. The power of a single horse is known as one horsepower. More precisely, horsepower is a unit of power that James Watt devised while looking for a way to express the performance of his steam engines. This gave us a unit for describing the power of steam engines and other machinery, and horsepower later became the standard measure of automobile engine output. Yet if we take the name literally, as "the power of a horse," horsepower is not a precise unit at all. Individual horses naturally differ in how much work they can do, but we set those differences aside and defined an arbitrary value as "one horsepower" in order to create a benchmark that commerce could rely on.

As industries developed, new units began to emerge across every field. Trade between nations grew, and with it came the need to measure and count what was being shipped. Cargo crossing the ocean, however, was remarkably hard to count. From a ship's point of view, a sack of grain and a crate of machine parts are both just cargo, yet they differ in both the space they occupy and the weight they carry. With every kind of goods loaded together, there was no simple way to say how much a ship could carry or how much a port could handle in a day. So in the 1960s, the International Organization for Standardization (ISO) set standard dimensions for shipping containers. Once the volume of a single 20-foot container became known as one TEU, the world finally had a way to measure the flow of cargo.

What Measurement Makes Possible

In this sense, units are a prerequisite for commerce. Once a unit exists, we can compare, put a price on things, and design around that unit. As the TEU took hold in global trade, ships and cranes, trucks and railcars, and port container yards were all designed to the same specifications. Neither horsepower nor the TEU was created for precise measurement. Both deliberately ignored individual differences, and in exchange, we gained a standard everyone could use.

We can ask the same question of AI. What unit can we use to count the work AI does? Writing text, composing music, generating video, and writing code are so different from one another that they seem as hard to count as the mixed cargo that filled ships before the standard container arrived.

Tokens: A Unit for Counting Intelligence

The most practical answer we have today is the token. If horsepower measured physical power and the man-month measured knowledge work, the token represents the next step. A fragment of text produced by a language model, a unit of sound generated by a speech model, and the information describing a single video frame are all tokens. We can count the tokens that go in and the tokens that come out, and we can attach a cost to them. Just as calculating TEUs does not depend on what is inside the container, tokens do not ask what they contain. For the purpose of counting, it simply doesn't matter.

Comparing tokens to electricity makes the role of a unit clear. Consumers pay for electricity in kilowatt-hours, while the companies that build power plants plan in kilowatts. Consumers pay for how much they use, but the infrastructure is designed around how much power it can deliver at once. AI works the same way. Users pay in tokens, and the infrastructure that produces those tokens is designed around how many tokens it can generate per second. For the first time, both the amount of work and the rate of work can be expressed in the same unit.

How New Units Give Rise to Industries

Units do more than measure what people can do; they can also change where money goes and how it flows. The money that follows a new technology is much like venture capital. It invests in potential and in the story. But once an industry reaches the scale-up stage, things change. The capital that goes into building power plants and seaports is no longer venture capital. An industry can only move forward when it can calculate unit costs and forecast demand, and that requires a unit. Consider a few examples. Electricity became a pay-as-you-go commodity only after meters were introduced to measure usage. Cloud computing, too, became a resource that anyone could freely buy and sell only once it had a price list for renting server instances by the hour. A technology can be born without a unit, but to become an industry, it eventually needs one of some kind.

This is exactly what is happening in AI right now. NVIDIA CEO Jensen Huang calls today's AI data centers 'AI factories.' One could dismiss this as marketing rhetoric from a company that needs to sell a lot of GPUs, but it is worth noting because it does reflect a real shift in perspective. Seen in terms of what it produces, an AI factory is really a token factory. Traditionally, a data center was a computing facility that stored data and performed calculations on request. The token factory turns that view on its head. It is a place where electricity, hardware, and trained models go in, and a product called tokens comes out. Raw materials enter, equipment runs, and a countable product ships, so the factory analogy fits quite literally. Being able to count power made it possible to build factories, and being able to count cargo made it possible to design ports. In the same way, being able to count intelligence now allows us to talk about factories that produce intelligence.

The shift in AI infrastructure's center of gravity from training to inference, which everyone seems to be discussing lately, can be understood in the same light. Training is closer to product development, while inference is closer to mass production. Development only needs to finish once, but production cannot stop as long as orders keep coming in, and as orders grow, unit cost and delivery time become critical. As the focus of AI infrastructure moves from large-scale training to running real-world inference services, technologies that efficiently handle requests across many models and users are becoming increasingly important.

So let us return to the idea of the token factory. When we evaluate a factory, we don't judge it solely by how many machines it owns. We look at output, unit cost, on-time delivery, and yield. Since a token factory ultimately takes in electricity and produces tokens, it will be measured by how many tokens it delivers per watt and per dollar.

What Makes Up a Token Factory

Does stocking up on accelerators, then, give you a token factory? Installing machines alone does not make a finished factory. You also need operations that fill the gaps between machines and keep each machine running as it should. The reason we never called data centers factories is that what they did was not manufacturing but leasing. Until now, what organizations with infrastructure could offer was equipment uptime, in other words, renting out machines. In the mainframe era, you paid for the CPU time you used, and GPU clouds have priced their service by the hours a GPU is rented. But this unit only tells you how long a machine was rented, not what was produced or how much. On top of that, whoever rented the equipment had to deploy models on it, build the serving environment, and tune performance on their own. That is certainly not a factory. To turn this structure into a factory, both input and output must be measurable, and someone must also take charge of routing each order to the right production line.

token-factory-en-v2@2x.png

A token factory can be broken down into three main components: the production line, the dispatcher, and the control room. The production line is where tokens are actually produced. It refers to the accelerators that generate tokens and the models running on them. Because the production line must run nonstop and at peak efficiency at all times, an operations layer that can manage the equipment efficiently is essential. This is the layer handled by software such as Lablup's Backend.AI.

Backend.AI can divide a single physical GPU into multiple fractions and allocate them to different models and users. Keeping the line running without idle time, and making the scale more fine-grained than counting GPUs one whole card at a time, is what this layer does.

The second component is the dispatcher. If inference requests from users are the orders coming into the factory, a token factory faces something close to chaos, with orders pouring in from countless customers. Running the factory properly therefore requires an operating strategy that distributes these orders efficiently while minimizing bottlenecks. Lablup's Continuum Router handles this layer. Continuum Router provides a single OpenAI-compatible endpoint for these requests, manages load balancing and automatic failover, and routes each request to the appropriate backend. And because every order passes through this point, this is also where tokens are metered.

The third component is the control room. Just as a factory needs a control room to oversee all of its equipment and every shipment in and out, a token factory needs a place where every aspect of the operation is visible. Lablup's Continuum Hub serves this role, centrally managing the policies and costs of the requests processed by the Routers. It records who used which model, how much, and which team and project each cost belongs to. Most importantly, it follows the principle that oversight must never become a bottleneck, so the path along which requests are actually processed is kept separate from the management path. As a result, even if something goes wrong in the control room, the production line keeps running.

With all three components in place, the factory meets the conditions described earlier. Backend.AI knows how much GPU capacity was used, and Continuum knows how many tokens were produced in that time. Combining the two makes it possible to measure both input and output. The result of bringing these pieces together is what we call a token factory.

Counting Different Tokens the Same Way

At this point, one problem comes into view: the token is not yet a finished unit. Token sizes differ from model to model, so the same sentence might be 100 tokens in one model and 120 in another. For now, the tokens produced by a small model and those produced by a large model after deep reasoning also accomplish different amounts of work. Shipping containers were no different at first. Before ISO set the standard, box sizes varied from company to company. The unit for counting intelligence is still going through that same phase.

So when a single factory runs multiple models, token counts cannot simply be added together; they need to be converted. That conversion is the job of Hub, which serves as the control room. Continuum Hub lets you set prices for each model. By applying model-specific prices to the tokens tracked by the Router, usage across different models becomes measurable in a common form: cost. Like horsepower and the TEU, it is a rough method of counting, but at least within a single factory, tokens from different models can now be counted in one unit.

Using a Token Factory

For users, the token factory already looks familiar. Services like OpenRouter are a good example. Connect to a single endpoint and you can choose from hundreds of models, each priced per a set number of tokens. What stands out is that the same model is offered by multiple providers at different prices. Once the token became a common unit, a market formed on top of it, and each provider listed in that market is itself a token factory.

Using "metered intelligence" in this way changes how we work with AI. Just as we don't think about the power plant when we use electricity, we start to ask how much intelligence a given task requires. Lightweight tasks such as summarizing meeting notes can go to a small model, while complex analysis goes to a large one, so each task gets just as much intelligence as it needs. We are billed for exactly the intelligence we choose to use, and we can order as much as we need without knowing which machine produced the tokens.

Closing Thoughts

Horsepower was never meant to measure a horse's strength precisely, and the TEU was never meant to measure the value of cargo. Both were standards established to create momentum for what came next. The same is true of tokens. Once tokens could be counted, the data center, which had simply rented out machines by the hour, became a factory that ships a standardized product. And because the data center became a factory, we can now talk about production cost, price, and reliability of supply, concepts that have long defined traditional manufacturing. At lab|up >/conf/5 in September 2025, Lablup announced its goal of becoming a company that quantitatively measures the amount of intelligence a given task requires, much like electricity, and delivers it reliably. Lablup's token factory, built on Backend.AI and Continuum Router & Hub, will be the first step on that path. See that future for yourself at lab|up >/conf/6 on September 29.

We're here for you!

Complete the form and we'll be in touch soon

Contact Us
lablup

Headquarter & HPC Lab

KR Office: 8F, 577, Seolleung-ro, Gangnam-gu, Seoul, 06143, Republic of Korea US Office: 3003 N First st, Suite 221, San Jose, CA 95134

  • facebook
  • youtube
  • Linkedin
  • GitHub

© Lablup Inc. All rights reserved.

We value your privacy

We use cookies to analyze site traffic, understand how visitors use our website, and improve our services. Necessary cookies for basic site functions are always active. Learn more

By clicking "Accept All", you agree to the storage of analytics cookies on your device. Click "Reject All" to keep only necessary cookies, or "Customize" to choose for yourself. You can change your settings at any time.