When the Model Is Free, the Harness Is the Business

DeepSeek V4 Flash started a token price war, and the market didn’t stop there. OpenAI cut prices on GPT-5.6 Luna by 80 percent, and Meta’s Muse Spark quoted rates that sit at a few cents per million tokens, provided you accept training on your data (like everyone else probably, but at least they wrote it down). The data is worth more to them than the tokens you consume.

When intelligence costs less than a rounding error, the margin has to live somewhere else.

Commodity.

You will read the word cheap several times in this article, because comparing prices is part of the argument. I want to be clear about what I mean. When I say a model is cheap, I mean capable models that sit below the absolute frontier of raw power. The frontier still has a role. Competent teams reserve it for the initial planning of a new project, and for the moments where complexity has to be structured before anything else can run.

Day-to-day execution doesn’t need that horsepower. Once the plan and the structure exist, the routine workload routes to models that cost a fraction of the price and deliver results that are close enough. The frontier is becoming the architect. The affordable middle class does the building.

I wrote some weeks ago that Large Language Models would follow the path of relational databases, operating systems, and web hosting. The barriers to access would collapse, and the value would shift away from the raw model toward the systems around it. That argument was a hypothesis at the time. A few weeks of releases have been moving in that direction, even if the picture is still incomplete.

The Price War Nobody Won

The price war is the clearest signal that the model is no longer the product. When a frontier release forces competitors to slash token prices by 80 percent within days, the pricing power has already left the building. Laboratories are now fighting on a margin that keeps shrinking, while their capital requirements keep growing. That combination is structurally broken, and every player in the market knows it.

This is exactly what a commoditizing technology looks like before the consolidation phase. The same pattern played out with relational databases, with operating systems, and with web hosting. Access costs collapse, performance curves flatten, and the companies that survive are the ones selling something around the core product, never the core product itself.

The West Is Opening Too

The most decisive evidence doesn’t come from the Chinese labs that started this. DeepSeek, Qwen from Alibaba, Kimi, and MiniMax keep releasing open-weight frontier models at API prices the American giants cannot match. The truly revealing move happened on the other side of the ocean.

Meta released Muse Glimmer as open weights under the Apache 2.0 license and committed to opening Muse Spark 1.2. NVIDIA ships the Nemotron family as open weight. Thinking Machines Lab, founded by Mira Murati after leaving OpenAI, released Inkling with the full weights on Hugging Face, a 975-billion parameter model they trained from scratch.

An ex-CTO of OpenAI drops her first model as open weights, complete with training papers. Meta returns to the open-weights strategy it abandoned a year ago. NVIDIA pushes open models because open models need more silicon. The strategy conversation is settled. Every player is reading the same room, and the model itself is where the value stops living.

The Harness Is the Product

If the model is cheap, the differentiation must sit in the shell that makes the model useful. That shell is the harness.

The harness is the software layer that turns a raw language model into a working system. It handles tool calls, memory, context management, and evaluation loops. In our world, it defines how the agent talks to your ERP, your warehouse systems, and your external data sources. It enforces guardrails and decides which model handles which task. You can swap the engine and keep the machine running.

Think of a model as an engine and the harness as the chassis around it. Two cars with the same engine feel completely different when the steering, the suspension, and the electronics differ. The same holds for an open-weight model running in a good harness versus a bad one. Companies that treat the harness as a commodity will keep paying in wasted integration work, repeated mistakes, and unstable output. Companies that treat it as the product build systems that survive a model swap.

This framing matters in the enterprise more than anywhere else. The models that dominate the benchmarks today are already outdated three months later, and your data pipeline shouldn’t care. When the harness is done well, the model underneath is a replaceable component. When the harness is done poorly, the whole system is rebuilt every time the vendor changes its pricing.

Thinking Machines shipped Inkling with an explicit note that the model runs inside the OpenCode harness. Even a frontier lab doesn’t sell the model alone. The product is the model plus the environment that controls it. That detail confirms the shift better than any benchmark.

If you have used tools like Codex or Claude Code, you already know what a harness does, even if nobody ever called it that. Those products wrap a raw model with file access, command execution, code review loops, and a memory of what you’re building. You’re paying for that wrapper as much as for the engine underneath. Swap the engine and the workflow keeps running. Remove the wrapper and the model alone turns back into a chat box that types words but does nothing.

The Hardware Race Was Already Running

The local inference movement made this visible earlier. OpenClaw, the open-source personal assistant, sparked a Mac mini shortage in January 2026 as enthusiasts built personal AI servers out of machines never designed for inference. People bought unified-memory Macs, learned about quantization, and started treating the SSD as a first-class citizen for model weights.

That instinct was ahead of the mainstream, and the industry noticed. Apple pushed the unified-memory story with the M-series line, quietly turning a desktop appliance into an inference box. NVIDIA answered with the DGX Station line, and then with the DGX Spark, a device that looks like a workstation and carries a price tag above five thousand dollars. That number looks absurd for a consumer device, and that is exactly the point.

Before calling that figure stratospheric, think about the hardware people already buy. A solid workstation for heavy video editing or 3D rendering sits in the same range, and nobody blinks. The DGX Spark is effectively an editing-grade workstation that runs inference instead of timelines. If that price is already acceptable for creative work, imagine the reaction once memory costs drop and the same box gets cheaper. The ceiling is not where the market is aiming.

The trend accelerated. Inference engines written in C, SSD streaming, and asymmetric quantization now run models with hundreds of billions of parameters on consumer hardware. Salvatore Sanfilippo’s engine squeezes DeepSeek V4-class models out of Apple Silicon. The software side of local inference is solved enough to matter.

The hardware side is where the real race begins. NVIDIA sells the DGX Spark at a price tag above five thousand dollars, and the machines people buy to run open models locally are becoming a category of their own. The suppliers of that category will sell it or rent it. Expect the same leasing logic Apple explored when memory prices pushed device costs upward, because a five-thousand-dollar box is a harder sell than a monthly fee. And that fee buys a box that runs the whole workflow in-house.

The Capital That Needs a Return

The money though, isn’t all flowing toward open models. Data centers are being built at an industrial pace, with dedicated energy plants attached. That capital was raised on the promise that intelligence sold by the token would pay it back. A market where models converge in performance and collapse in price makes that promise harder to keep.

The correction, when it comes, won’t spare everyone. It will separate the companies that built their business on selling tokens from the ones that built it on the harness and the silicon. The token sellers face a shrinking margin. The harness builders and the chipmakers face growing demand from every company that decides to run open models where it controls them. Sovereignty, fixed costs, and predictable bills are exactly what the commodity trend delivers.

Meanwhile the Chinese ecosystem keeps building its own independent supply chain for lithography, chips, and memory. Export restrictions no longer bind the way they once did, and that removes the last structural barrier to a fully open market. When hardware, software, and distribution all become available from multiple independent suppliers, the only defensible position left is the layer that sits between the model and the business process.

The model is cheap. The harness is the business. The silicon is the moat.

Written by Andrea Guaccio 

August 18, 2026