Agents kills Inference Margins

One of the questions I have been very curious in the rise of AI since 2022 is where does the value accrue.

In 2023: Was perplexity a wrapper or a value creator on top of Open AI’s API
In 2024: Can Cursor, Windsurf, Cognition and others be able to create value or Labs will compete it away
In 2026: Can Agents like Muse and Instinct forever be free - what is the emerging business model?

All of this while we see Nvidia make record profits and great margins.

As I think deeper about this problem, few things are obvious to me
a) Human traffic of World Wide Web peaked in 2025

b) Agents is how we will primarily consume information, tasks, shopping, coordination, digital presence at large, therapy and almost all digital activity.

Our need for intelligence is going to grow 100x in the next few years (Insane) but soon we won’t be choosing which model to use, the agent layer does. . What happens when the application becomes the lab’s purchasing department?

If you are a lab selling inference, agents pose the a significant and deep margin risk. Agents and routers will squeeze every cent they can to make services cheaper for users.

The AI value chain

The useful map runs from TSMC → NVIDIA → data centers → model labs → inference platforms → applications → agents. I don’t think every part of the stack will capture value at the same margin and I expect competition to be higher - higher up the stack you go keeping customer prices cheap and that ripples back down the value chain.

Apoorv Agrawal’s April analysis, published in his personal newsletter, estimates annualized gross profit at $225 billion for semiconductors, $40 billion for infrastructure and $20 billion for applications. Chips account for roughly 79% of that estimated pool. [1]

Figure 1. Agrawal’s April 2026 estimates, not audited industry totals. His application category includes model labs application layers too.

That last detail matters. OpenAI and Anthropic account for roughly three-quarters of his application revenue currently. This is not a chart showing independent agents winning against labs.

The hardware economics are stark though,. NVIDIA reported $96.2 billion of quarterly revenue and a 75% GAAP gross margin for its quarter ended July 26, 2026. Those are company-wide figures, not inference-only results. [2]

So the starting point much of the money is still being earned below them and as agents and applications layer grows, much of the money will be earned above them.

So the inference business of the labs may not be as attractive in the future, and there is likely going to be continued push by labs into application layers and maybe layers below them too.

Adoption IS NOT pricing power

Ramp’s September release contains two facts that belong together. In August, 43.8% of businesses in its measure paid Anthropic and 39.8% paid OpenAI. Meanwhile, its effective token-price index fell from a March peak of $1.15 per million tokens to $0.68 in early September—about 41%. [3]



Figure 2. Published endpoints, not a reconstructed historical curve.


At a fixed mix, a 41% price cut requires roughly 69% more volume just to preserve revenue. That is a lot. A 41% price cut in 6 months is massive, but I believe they are seeing more than 69% volume jumps.

AI adoption increases, the cost will fall and inference might look like energy, telecom, internet - a mass utility priced as such.


China is not going away


I love the saying by Nassim Nicholas Taleb — 'The three most harmful addictions are heroin, carbohydrates, and a monthly salary.' Similarly, I believe cheap prices is the cocaine of the mass capitalism and it makes the world go around.

OpenRouter’s country chart is striking. Between the weeks of September 1, 2025 and June 8, 2026, Chinese-authored models went from 20% to 55% of its token volume. US-authored models went from 76% to 43%. [5]

Figure 4. Model-developer origin; OpenRouter input-plus-output tokens. Not customer geography, revenue share or worldwide inference share. Rounded endpoint labels are transcribed from the original chart.

The platform’s companion volume chart also shows American-model token usage growing over the period. Losing share did not mean losing absolute demand. [5]

But it does show one thing, the difference between top models is narrowing, and inference pricing power is slowly evaporating. China is not winning inference, but proving buyer flexibility to move for prices and open weights. It is proving demand and establishing a price elasticity curve.

Agent pays LESS

A person buying a chatbot chooses a brand. An agent executing a workflow can make a different choice at each step - chose a different model, different providers, different approaches. All of it to maximize utility while reducing the cost of serving us.

Use an inexpensive model to classify the inbox. A stronger one to interpret an ambiguous instruction. Deterministic software or model like Jev to check the arithmetic. A premium model to handle the exception and complex reasoning.

That is not the same as declaring models interchangeable. It is allocating differentiated models more carefully.

The RouteLLM researchers demonstrated the principle in 2024. Routing between GPT-4 Turbo and Mixtral 8x7B reduced costs by more than 85% on MT-Bench, 45% on MMLU and 35% on GSM8K, while attaining 95% of the stronger model’s benchmark performance. These were historical benchmark results—not a promise of equal quality or current production savings. [7]

Figure 6. Cost savings at a specified performance target. The MT-Bench bar marks the reported lower bound.

The behavior is appearing in practitioner discussions, too. One Reddit user described using GPT for specifications, DeepSeek for implementation, and several models for review. Their results were self-reported and AI-graded; they are not reliable comparative benchmarks. The workflow is the interesting part: keep the premium model, but stop giving it every task. [8]

When tinkerers start doing this, just know that within 12-24 months, it going to be productized and be available to all of us. Actually, given how fast OpenClaw led to Muse - the time gap might be 6 months or less.

The agent company’s valuable asset, then, is not merely a router. Routing can itself be copied. It is the combination of customer context, permissions, workflow integration and evidence about which completed work is actually acceptable.

The metric that matters is:

Cost per successful task = (all model and tool charges + human review and rework costs) ÷ successfully completed tasks.


My counter Argument - Agents wants MORE


Social media companies got us addicted to doom scrolling. Agent companies will have us all wanting MORE. Shop more, travel more, do more - because friction of doing is reduced, we must all must DO MORE.

I used to read maybe 100 pages of a book on long flights. Now I often find myself writing articles like this or doing deep research in my topics of internet or building tools or websites for my company.

My agents are getting me addicted to DO MORE. And I love doing more. Not a single person in silicon valley who is agent maxxing is getting to save time in a day to do liesurely stuff, in fact, they most often talk about 3am AI psychosis. Agents wanting MORE mean they will be 10x or 100x in their usage of inference that we can possibly do in our prompt led sessions.

[Please don’t have a sloppy argument, that silicon valley tech bros want to do more and goal maxx, but average people just want to chill. Everyone wants more - Everyone - might be different things, but greed is universal]



Agents are excellent customers for labs. Anthropic reported that, in its own data, agents used about four times the tokens of chat interactions, and multi-agent systems about 15 times. Those are observations about its systems, not universal multipliers. But they point to a real possibility: efficiency creates room to attempt much more work. [9]

OpenRouter’s summer discount analysis shows how large the demand response can be. During the July 27–August 14 promotion, average daily token usage was 13.8 times the pre-period level for Luna and 5.6 times for Terra. Sol, undiscounted during that window, was at 1.11 times. [10]

Figure 7. Before: July 8–26. Promotion: July 27–August 14, 2026. Observational comparison, not a randomized experiment or a profit chart.

The dates matter: launches and additional list-price changes overlapped. The chart supports strong price-sensitive demand, not a clean causal estimate or proof that the discounts were profitable. [10]

It also shows why the “cheap intelligence kills the labs” argument is incomplete. Lower prices can expand the market as well as redistribute it.

Andrey Fradkin’s research using early-2025 OpenRouter data finds both substantial multi-model use and persistent differentiation between models. Customers did not simply converge on the cheapest option. [11]

All of this to say, if Agents increase inference demand 100x in the next 5 years (likely faster), we can see a whole spectrum of labs amd models emerge - from the cheap to the cutting edge expensive use-cases for science, maths, biology, wars and more.


So the labs might look like a suite of products - some which represent utility like pricing and margins and higher end models that hold software like margins and both might expand exponentially.


Agents will accrue the highest value


Muse and Instinct make this debate tangible - I am sure Google and Open AI are cooking and won’t let Mark walk away with the spoils here. If I was Sergey and Larry, I would be putting everything I have on line to not lose another category (social) to Mark.

In V1 of AI, the labs held our context layer. Who am I ? What do I like ? What did I do yesterday ? Where are my digital files etc.

I think that meaningfully shifts to the Agents being the context layer and without context, the underlying reasoning loses a lot of teeth.

Muse is a pefect example of a model builder competing through an agent. My position is that in both in consumer and enterprise getting people used to your agents might be the best way to monetize your inference.. [12][13]

For an independent agent to weaken a lab’s position, three things must hold: alternative models must meet its quality threshold; switching must be practical after integration and evaluation costs; and the agent—not the lab—must retain the customer relationship.

I am betting all 3 will happen. Agents will pass savings to customers. Agents will monetize my every transaction. if it did booking in a second versus me taking 5 minutes - will I pay $1 - yes. Will I pay Agents $1 on every meal it buys me to keep me healthy, fit and on-time - hell yes. and the list goes on. Agents will make me addicted to MORE.

That is where I land: agents will 100x inference market in next 5 years or sooner, but will kill margin profile of inference. The labs will fight back by becoming agents themselves. Some will win. But they will win by owning the customer relationship, not merely renting out the best weights.

Sources

[1] Apoorv Agrawal, “The Economics of Generative AI: Two Years Later,” April 1, 2026.

[2] NVIDIA, Q2 fiscal 2027 results, August 26, 2026.

[3] Ramp Economics Lab, September AI Index, September 9, 2026.

[4] Ramp, “How Ramp data works,” April 3, 2026.

[5] OpenRouter, “DeepSeek V4 Is Earning Agentic Token Share,” June 30, 2026; original country charts.

[6] Ramp Economics Lab, “Our latest data on China vs. the American AI Labs,” July 8, 2026.

[7] RouteLLM research and author summary, 2024.  Author summary

[8] Reddit, r/codex, self-reported multi-model coding workflow and subsequent caveats.

[9] Anthropic, “How we built our multi-agent research system,” June 13, 2025.

[10] OpenRouter, “GPT 5.6 Discounts & Jevons Paradox,” August 25, 2026.

[11] Andrey Fradkin, “Demand for LLMs: Descriptive Evidence on Substitution, Market Expansion, and Multi-Homing.” Working paper; observations January 11–April 11, 2025.

[12] Meta, Muse launch announcement, September 8, 2026.

[13] Reuters, “Wall Street expects Meta’s AI agent to shape into a new revenue engine,” September 22, 2026.

[14] Instinct public product page, accessed September 22, 2026.

[15] OpenRouter’s public X-post feed and source-analysis links, accessed September 22, 2026.

Original OpenRouter charts: Country token share  |  US and China absolute token volume

Next
Next

The one belief that runs our life