One dinner in late June, the guy across from me was a GPU owner.
I won't name him. A few thousand cards, H-series and domestic brands, scattered across data centers in two or three cities. He'd built the stash from the IDC business years earlier, and for the past two years he'd been trying to find a way out.
After a few drinks, he slid his phone across the table. On the screen was a photo of Jensen Huang at GTC.
"Tell me the truth," he said. "Am I sitting on a gold mine?"
His math was straightforward: token demand is exploding, and Jensen says every token is revenue. Models: DeepSeek open-sourced those, free. Inference engines: vLLM, SGLang, open-source, free. Cards: he had plenty. Put the three together, and wasn't he a token factory? Lying back and printing tokens for money.
He was waiting for me to say "you're set." We'd been working on the inference layer for years, so he was asking me as an insider.
I said: Don't.
He froze for a few seconds. "This is exactly what you do every day. You're telling me to stay out?"
Because we do this every day. The dinner stretched from seven to eleven, and I broke the whole business down for him, piece by piece. This article is the full version of that night: why three open-source blocks don't make a factory, and why the hottest category—text-to-video—is something we won't touch.
Get the word "factory" right
I didn't invent the term "token factory." "AI factory" is Jensen Huang's own phrase. At GTC 2025 he said the data center has only one job: "generating these incredible tokens." The blunter version, "token factory," came from Satya Nadella. At Microsoft's October 2025 earnings call, he asked his team: "How efficient is our planet-scale token factory?" By March 2026, NVIDIA's own technical blog put "token factory" in the headline. At GTC Taipei in June, Jensen's line was even blunter: "Every token is profit. Every token is revenue."
Many people hear "factory" and think it's an energy or compute business: if you have power and cards, you can start production. That's the biggest misunderstanding of the word. Having equipment is only the threshold. A factory is about wringing maximum efficiency from the whole production chain.
A token factory breaks down into three layers of technology stacked together, each layer holding efficiency levers of several times to several dozen times.
Layer one: chips. Between generations, the jump is orders of magnitude. NVIDIA's official line is that the GB200 NVL72 is 30× faster than H100 for large-model inference; running DeepSeek R1, Blackwell cuts the per-token cost to one-tenth that of H200. Vera Rubin, announced in March, cuts per-token cost to another one-tenth of Blackwell, and by June it was in full production. At GTC, Jensen said the line that's now quoted everywhere: after Blackwell volume production, nobody would want Hopper even if it were free. He calls himself the "Chief Revenue Destroyer."
Layer two: the model. The model determines the ceiling of intelligence a factory can produce. That's easy to grasp. What's easy to miss is the other half: the model architecture itself determines inference efficiency. How sparsity is handled, how attention is designed, how parameters are distributed—at the same intelligence level, different architectures can make production costs differ several-fold. Designing models is increasingly like designing industrial products: performance and manufacturability must be locked in together on the drawing board.
Layer three: the inference system. The full stack where software meets hardware, and the layer the industry most underestimates. On the same hardware, running the same model, the inference system can create several-fold differences. When vLLM came out, it delivered 2–4× the throughput of the best systems at the time; SGLang's launch claimed up to 5×. The most extreme example is NVIDIA's Dynamo: on the same GB200 NVL72 running DeepSeek R1, turning the system on or off changes per-GPU token output by 30×, according to NVIDIA's official figure. Same hardware. Same model. 30×.
The inference system has a more famous footnote. In March 2025, DeepSeek opened its inference-system books: GPU rental costs ran about 562,000, theoretical cost-profit ratio 545%. They themselves stated that actual revenue was far lower than this theoretical figure, but the numbers still shook the industry. The cards were the same; the inference system was squeezing every last drop from each one.
Only when these three layers are stacked together does it count as a factory. Chip choice, model architecture, and inference system—each layer must be dialed in near its optimum. That and "stacking three open-source blocks" are separated by a productivity gap of several dozen times.
What was my friend's plan? Last-generation cards, open-source engines running default parameters, models designed by others. He wasn't trying to start a factory; he was setting up a workshop next to someone else's industrial park.
Efficiency is life or death in a commodity market
A workshop can survive, right? Small restaurants thrive next to chains.
No. Because tokens are commodities.
I wrote earlier in Token Is Not One Thing that tokens are a matrix of intelligence tier by speed tier and shouldn't be mixed together. But within the same cell of the matrix—same model, same SLA tier—tokens are a pure commodity. A DeepSeek V4 token you produce and a DeepSeek V4 token I produce are identical, byte for byte. Buyers look at only two numbers: price and speed.
The iron rule of commodity business, anyone who's traded commodities knows it: the selling price is set by the market; the only thing you decide is cost. Efficiency gaps turn straight into profit gaps, with no brand, user experience, or story to hide behind.
And this market is almost built for perfect competition. Supply is globally connected, there's no regional protection, and prices change by the day. Open OpenRouter right now and you'll see the same open-source model listed by four or five suppliers at once. Input prices differ by nearly 5×, speeds by several times. The expensive ones are fast; the cheap ones are slow. It looks exactly like a commodity trading board. In May, DeepSeek simply made the V4-Pro 75%-off promotional price permanent. Over a longer horizon, the numbers are scarier: a16z calculated at the end of 2024 that the price of equivalent capability falls 10× per year; Epoch AI's latest figure is about 40× per year.
In such a market, if you hold a 50% efficiency advantage, you have two options: keep prices steady and pocket the extra as profit, or cut prices and take demand from others. Thicker margins let you lock up more upstream cards and grab the raw material too. Squeezed from both sides, players without a technical edge quickly discover that running their own factory is worse than renting their cards to people who know how to run one.
That's the core reason I told him not to build it. Token factories will not stay small and scattered; like all commodity industries, they will become highly concentrated.
Capital has already voted on this judgment. Just look at the last half-year: Fireworks at a 1 billion, processing 40 trillion tokens a day; Baseten at 8.3 billion; Cerebras' first-day market cap hit around 20 billion to acquire Groq's technology and core team. Money doesn't lie; they're buying the same thing: the ability to push token production efficiency to its limit.
By the way, where is this consolidation happening fastest? AI coding. In the OpenRouter and a16z 100-trillion-token study, coding consumption rose from 11% at the start of 2025 to over 50% by year-end. Where demand is most standardized, SLAs clearest, and buyers most professional, the factory model always takes shape first.
In 1987, someone already told this story
Halfway through the meal, he asked a genuinely good question: the model companies themselves are selling APIs and running factories. Why would a third party stand a chance?
I told him the story of TSMC.
When Morris Chang founded TSMC in 1987, he made a decision no one understood at the time: only manufacturing, no design, never competing with customers. That promise was later written, almost verbatim, into TSMC's annual report every year. Before that, the industry norm was integrated design and manufacturing; a chip company was supposed to look like Intel. Manufacturing for others? That was grunt work.
But it was precisely that promise of neutrality that created a new species from nothing: fabless. Design companies no longer had to pour billions into fabs. Draw the blueprint, and get capacity from TSMC. When NVIDIA was founded in 1993, it didn't own a single fab floor, and it still doesn't today.
The token world is splitting into the same structure.
The closed-source giants are Apple. OpenAI, Anthropic—they don't hand out their models, they grind their own inference stack, they co-design chips with suppliers, and they run the whole chain from design to retail themselves, setting their own margins. That's a coherent path, as long as the model stays ahead.
The open-source camp fights as a team. Model labs release the design and open the weights; anyone can produce. And third-party token factories are the TSMCs of this camp. They don't make models, they don't bet on who wins; they hold what everyone needs: neutral, maxed-out capacity.
Why is neutrality valuable? Ask one question: Can one model company's compute become another model company's production capacity? No. This month B's model tops the charts; A would never free its cluster to help its worst rival ship. A model company's capacity is locked to its own model: model succeeds, capacity is valuable; model fails, capacity goes down with it.
Third-party factories don't have that lock. Capacity switches to whichever model is good. Switching speed is faster than intuition would suggest: last July, when Kimi K2 launched, two inference providers had it listed within six hours; Groq went live on day three; three days after that, measured speed hit 40× that of Moonshot AI's official API. Models are theirs, capacity is yours, the windfall is everyone's.
This structure is also an antidote for model companies. For a new model company today, the heaviest burden is inference compute: training is something you can rent if you grit your teeth, but inference grows with users and is a bottomless pit. But if the capacity layer is neutral and open, the playbook changes completely: focus on training the model, open-source to open the market, then turn around and book capacity. Like Apple booking the first run of TSMC's 3nm capacity (reportedly 90% of the initial output went to Apple), or the entire 750,000-unit annual supply of Ascend 950PRs being neatly snapped up by ByteDance, Alibaba, and the other giants. Model companies become fabless: light assets, bet on design, hand capacity to neutral factories. A new leader can go from launch to scale in months.
For factories, neutrality is the only risk-resistant posture. How model companies swing with each generation is a textbook case in the last 30 days: in mid-June Zhipu released GLM-5.2; a week later its stock jumped over 42% intraday, market cap breaking a trillion HKD. Just yesterday, Moonshot AI released Kimi K3, and Zhipu fell as much as 30% intraday, with MiniMax down 16%. The same company went through heaven and hell in one month.
Binding yourself to any single ecosystem is betting it wins every generation. I wrote about how brittle this binding is in Intelligence as a Strategic Resource: even national-level bets can fail. TSMC never bets on which design company wins; it only bets on one thing: the world will always need chips. Token factories only bet on the same thing: the world will always need tokens.
There will always be another catfish
By this point he was basically convinced, but he added a sharp question: this whole logic rests on models being open-source. What if everyone stops one day?
That concern deserves a serious answer. My judgment: there's no going back.
Let's replay the 2025 domino chain.
On December 26, 2024, DeepSeek V3 was open-sourced, weights released right after training finished, performance matching overseas closed-source flagships, better than every domestic closed-source flagship. Domestic closed-source models were suddenly awkward: the best one was free. What gave the paid ones the right to charge?
The dominoes fell faster than anyone expected. On January 15, MiniMax, which had never open-sourced a flagship, rushed out the full weights of MiniMax-01—note: this was before R1 was released. On January 20, R1 landed under the MIT license and made a splash. By AI Product Rankings' count, it added over 100 million users in the following seven days, while ChatGPT took two months to hit that mark. On January 27, NVIDIA dropped nearly 17% in a single day, erasing $589 billion in market cap—the largest single-day market-cap loss in U.S. stock market history at the time.
Then came the rest of 2025: in April Zhipu moved the GLM series to MIT; in June MiniMax M1 used Apache 2.0; in July Kimi K2 came under a modified MIT license and GLM-4.5 under plain MIT. Within a year, "flagships must be closed-source" became "flagships are embarrassed to launch without open-source weights."
This wasn't driven by idealism. It was a cold marketing calculus. You train a good model. How do you tell the world? Buying ads and holding launches are slow and absurdly expensive. Release the weights, and the open-source community's distribution machine runs for free: downloads, word of mouth, inference providers rushing to list it, benchmarks flooding the feed. 100 million users in seven days. That kind of result has never happened in advertising history. It can only come from open source.
The race hasn't slowed today. The bomb that dropped Zhipu 30% yesterday is itself the latest move in open source: Kimi K3, 2.8 trillion parameters, weights released directly, the largest ever.
So even if DeepSeek closes one day, some unknown team with less compute and a worse seat at the table will release the best thing directly, because that's the most efficient way for an underdog to flip the table. As long as one catfish does this, everyone else has to follow. Moreover, in the current landscape, Chinese models account for 17.1% of global open-source model downloads, surpassing the U.S. share of 15.8%. This flywheel won't stop.
Open source exists, so the recipe exists; the recipe exists, so the production layer exists; the production layer exists, so the foundation of token factories exists.
Why we don't touch text-to-video
Finally, one of our own decisions, and also the question that came up at the end of the night: text-to-video is so hot, demand for Seedance is everywhere. Why doesn't a token factory do video too?
The answer is simple: we can't. It's not that it's hard. The business structurally doesn't exist.
A factory can only start production if you hold the recipe. On the text side, the recipe is open, so the whole supply chain has grown. Video is the opposite. The models with the highest demand are all closed-source: ByteDance's Seedance, Google's Veo and Gemini, Kuaishou's Kling. Seedance is on version 2.5 and has never released weights once.
Wanxiang makes the point even clearer: it was the biggest open-source exception in this space, with Wan 2.1 and 2.2 under Apache 2.0 and over 30 million downloads. But starting with 2.5, Alibaba pulled it back, leaving only an API. The more demand rises, the tighter the grip. On the video track, even the only exception has disappeared.
Models are the vendors' lifeblood. They run them closed-source and keep closed-source profits. Why would they give them to you?
Without the recipe, the only role left is renting out compute. The vendor takes your cards, deploys, operates, and sets its own prices; the inference optimization is theirs too. The inference techniques you're proud of have nowhere to stand in this chain. Without a production layer, there's no efficiency gap, and no excess profit. This isn't a business; it's collecting rent.
So whenever someone asks if we should do video inference, I shoot back the same question. It's also the first question we ask internally about any new direction:
Who holds the recipe?
In the end
At the end of the dinner, the owner asked: "So what do I do with these thousands of cards?"
I told him the truth: while the market is hot, rent the cards to people who know how to run factories, and reinvest the proceeds where you actually have an efficiency advantage.
What he finally decided, I still don't know. But I'll write down for you what I said to him that night.
This industry manufactures the illusion of "owning resources" every day: those with cards think they own compute, those with money think they have a ticket. But resources are the fastest-depreciating assets in the industry: cards drop a tier every three months, models release a generation every three months, today's scarcity is inventory in six months.
The only thing that survives cycles is the ability to use the same resources more efficiently than others.
The year TSMC was founded, contract manufacturing was seen as undignified grunt work. More than thirty years later, the proudest design companies line up for its capacity.
Position matters more than effort. Find a position where everyone needs you and you don't bet on any single winner, and grind efficiency until others can't catch up.
The wind will shift, and thrones will change.
Efficiency always has buyers.
References
Token Factory Concept
- NVIDIA GTC 2025 keynote transcript ("AI factories ... generating these incredible tokens") — Rev
- Jensen Huang Computex 2025 speech ("You apply energy to it, and it produces ... tokens") — NVIDIA Blog
- Satya Nadella "planet-scale token factory" — Microsoft FY26 Q1 Earnings Call Transcript
- NVIDIA Technical Blog puts "Token Factory" in the headline (2026-03-25)
- Jensen Huang GTC Taipei 2026: "Every token is profitable. Every token is revenues." (transcript)
Chip Layer
- GB200 NVL72 up to 30× H100 inference — NVIDIA Product Page
- Jensen Huang: Blackwell vs. H200 on DeepSeek R1 per-token cost 1/10 — NVIDIA FY2026 Q3 Earnings Call
- "you couldn't give Hoppers away" — GTC 2025 keynote transcript
- Vera Rubin platform launch: per-token cost about 1/10 of Blackwell — NVIDIA Press Release (2026-03-16)
Inference System Layer
- vLLM launch blog: up to 24× throughput vs. HF Transformers
- PagedAttention paper: 2-4× throughput vs. then SOTA systems (SOSP 2023)
- SGLang launch blog: up to 5× throughput
- NVIDIA Dynamo: GB200 NVL72 running DeepSeek-R1, per-GPU token output improved over 30× (projected) — NVIDIA Press Release
- DeepSeek-V3/R1 inference system overview: daily cost 562K, theoretical cost-profit margin 545% — DeepSeek Official GitHub (2025-03-01)
Pricing and Market
- a16z LLMflation: equivalent-capability inference price falls ~10× per year
- Epoch AI: fixed-capability inference price falls ~40× per year (February 2026 update)
- OpenRouter example of the same model listed by multiple providers (Qwen3-Coder, 6 providers, prices differ by nearly 5×)
- OpenRouter × a16z State of AI: An Empirical 100 Trillion Token Study: coding category rose from ~11% to over 50%
- DeepSeek makes V4-Pro 75% discount permanent — Official Pricing Page
Third-Party Factories: Funding and Capacity Switching
- Fireworks AI Series D: 17.5B valuation, ARR over $1B, 40 trillion tokens per day (2026-07)
- Baseten Series F: 13B valuation (2026-06)
- Together AI Series C: 8.3B valuation (2026-07)
- Cerebras NASDAQ IPO: first-day market cap about $95B (CNBC, 2026-05-14)
- NVIDIA acquired Groq core technology and team for about $20B (CNBC, 2025-12-24)
- Kimi K2 listed by Novita/Parasail within 6 hours of launch — OpenRouter Wayback Archive (2025-07-11 21:16 UTC)
- Groq launches Kimi K2 (day 3, 185 tok/s)
- Artificial Analysis measured Groq running K2 at over 400 tok/s, 40× Moonshot AI's official API
- SiliconFlow and Huawei Cloud Ascend launched DeepSeek R1/V3 (2025-02-01, GeekPark repost of official release)
- Huawei Ascend 950PR annual 750K-unit capacity locked by ByteDance/Alibaba/Tencent/Baidu/enterprises and government (East Money)
Open-Source Domino
- DeepSeek V3 open-sourced on release (official announcement, 2024-12-26)
- DeepSeek R1 weights repository (MIT, 2025-01-20)
- AI Product Rankings: DeepSeek added 100 million users in 7 days (Web + App cumulative, non-deduped, Sina repost)
- ChatGPT reached 100 million monthly users in two months (Reuters, 2023-02-01)
- NVIDIA lost $589 billion in market cap in one day, then a U.S. stock market record (CNBC, 2025-01-27)
- MiniMax-01 open-source (official announcement, 2025-01-15, five days before R1)
- Zhipu GLM-4-32B-0414 series switched to MIT (2025-04-14)
- MiniMax-M1 open-source (Apache 2.0, 2025-06-16)
- Kimi K2 modified MIT license text (2025-07-11)
- GLM-4.5 weights MIT (2025-07-28)
- Chinese open-source models' share of global downloads reached 17.1%, surpassing the U.S. at 15.8% (People's Daily Online, 2025-12-08)
- Moonshot AI releases Kimi K3: 2.8 trillion parameters, the largest open-weight model ever; Zhipu fell as much as 30% intraday (CNBC, 2026-07-17)
Foundry Model and Text-to-Video
- TSMC company profile: founded professional wafer foundry model in 1987 — TSMC website
- TSMC annual report original wording: "does not compete directly with its customers"
- NVIDIA company history: founded in 1993, no own fab, first chip manufactured by SGS-Thomson — FundingUniverse
- Apple reportedly secured ~90% of TSMC's initial 3nm capacity (MacRumors citing DigiTimes, 2023-05-15)
- ByteDance Seedance 2.0 official launch (2026-02-12, closed-source)
- Seedance 2.5: native 30-second single-clip video (The Decoder, 2026-06)
- Artificial Analysis text-to-video leaderboard (checked July 2026: top models are closed-source)
- Tongyi Wanxiang official: Wan 2.1/2.2 open-source, over 30 million downloads (2026-01); weights not open from 2.5 onward
- After GLM-5.2 release, Zhipu's market cap broke HK$1 trillion, intraday up 42% (Caixin Global, 2026-06-23)
