The price for a given level of AI abilities decreases by 5-10x per year. (eg. epoch link to 10-900x, plus a 5-10x deflation paper). This will greatly reduce the competitive advantage of future frontier models. Including realistic responses it may not even be valuable to train.

In 2026 Fable-5 cost say 10B to train and 50/million tokens to run. In 2027 Fable-6 would cost say 100B to train. For Fable-6 to generate 100B in inference profit, not only does it have to be 100B better than the 2027 competition, it also has to be 100B better than using Fable-5 at 2027 prices. At 10x/year deflation that’s 5 per million tokens. In 2028 to train Fable-7 it must be 1T better than Fable-5 at 0.5/million token AND Fable-6 at 5/million tokens. To fund this Anthropic will have much less secure cash flows. Think of cash cohorts based on an AI’s capabilities for that year. 100B in revenue in 2026 is not ARR, it is 100B for 2026 capabilities in 2026, 10B for 2026 capabilities in 2027 and 1B for 2026 capabilities in 2028. There could be increased demand for 2026 capabilities if the price was 1/10th as much, but considering how much AI spend is already specutlative, it is explicitly investments preparing for when AI is even better and cheaper, I doubt there’s a Jevon’s paradox. Only recently there were lots of talk of not seeing an ROI on AI installs, and most of the cost was setting up the system not in token spend.

But if new use cases emerge that were not previously possible then the demand trend could continue. METR says each year frontier AIs can reliably do tasks taking 7x more unaided human hours than the year before (131-day doubling, 2^(365/131)≈6.9x). An AI that can work for 7x as long could be worth 10x as much. It also might be cheaper to divide a big task into smaller tasks cheaper models can do, even with the extra human planning and review. Question: Is x-axis “unadied human time” or “human time”? AI didn’t accelerat humans in Spring of 2025, summer of 2025 has many people talking about AI acceleration everywhere but in productivity statistics, and then agents at start of 2026 become actually good. If x-axis is “human time” that includes time managing AI, tools call tools, and AI accelerates in total abilities.

I have a new parameter, Beta which models this, I estimate it’s X with 90% CI range Y-Z which means Q. Various parameter sweeps of modeling of integration costs. (Do this as a function encompassing all terms taking only task length.)

A digression where log-lines are true but the economic impact changes. An example is Moore’s Law and Denard Scaling.

Is Metr x-axis “unaided human time” or “human time” because even in early 2025 they found an AI slowdown of 20% for developers. Only with Agents in late 2025 was AI consistently a net acceleration to developers, and Mythons in early 2026 models is about when the trendline broke. How does this change all of this?

Metr - Twitter hype

METR benchmark performance versus model price deflation

There’s almost no business usecase for an AI right 50% of the time. Include the graph of Metr logistic where 15%=twitter type and 85%=can use as coding agent or wrap into harness and 95%=vibecode. Also the recent hacks show that with 0.1% success rate is enough to do scary things, or that agents together get a lot more done.

Practical jumps of 80% are what matter. Very few use-cases for code that’s 50-50 to be right. The word for someone who’s reliable 80% of the time is “unreliable”. Personal experience: asking for 15m complex feature works way better, 3h independent data science projects, and all night if after fixing the data I tell it to go wild.

Current AIs can do incredible things and are very intelligent. However they have spikey abilities, see Gary Marcus for all the simple things they can do. Say LLM intelligence is 80-98% correlated with human intelligence. Talk of RSI should use the right terms. “AI deployments will have no substantial limitations because they have enough of a factor that’s 80-98% correlated with human intelligence” isn’t obvious. The whole point is that AI isn’t as generally capable as a human. They can at times do AGI things, but we haven’t seen if it is all there yet. It also indicates why a great deal of alignment research on the current system of AIs isn’t as useful: for true ASI much of the others will be handled.

Aside: I think a lot of RSI worries would disappear if projecting forward 5 years “The intelligence of 1m Terry Taos in a datacenter begets even more intelligence than all humans combined” was replaced by “An intelligence which correlates 80-98% with humans and 80% of the time can do human tasks which would take 4h, when working together in 1million copies will become even more intelligence than all humans combined”. Still rather possible, but no longer self-evident. And I think it encourages better expectations about what real near-RSI systems look like.

Deflation Business impacts - Cohort Math

Aside: It also means that GPT-wrappers have a bad business model for a different reason. Sell 2026 AI tokens for 50 cents on the dollar, get a ton of customers. In 2027 those same 2026 capability tokens only cost 5 cents and you have 90% margins. Even if your product isn’t much better than codex on their desktop, the biggest service you sell is the enterprise sales, integration, and white glove setup so people can actually use AI this year. Many companies will have a corporate AI/IT policy of “everyone just run codex on your desktop”, implicit or not, but few of those will be companies with large amounts of Brand Equity. You gift the Brand Equit-ies an AI policy and monetize by stopping them from using a more expensive AI next year.

No longer are you worried that next year ChatGPT will make your product a feature. You’re worried that the next OAI releases (or default enables) a model router which automatically does the financial controls your business was banking on. Given all the math above and even worse chip shortage they will.

This is why most AI integrations are fixed fee: a markup on AI spend does not capture much of the future value to the client when the integration spend should be deflating at 10x per year. Charging per task completed could work, though since the client is paying way more than true marginal cost it substantially alters behavior.

Conclusion:

The model cost deflation trendline is important and underconsidered. It determines the relative advantage of building the next generation of models which determines scaling, and the next tech breakthrough is which has large business and alignment impacts.

Predictions:

Most SWEs stop using frontier models day to day. Most benchmark gains come from Harness. 4-8 agent Swarm of Frontier generation minus 1 becomes standard. Fable-7 will be a 3x larger model than the existing trendlines predict. If the frontier model no longer solves tasks directly, it’s total inference demand goes down. With less inference demand Sardana et al. “Beyond Chinchilla-Optimal Inference-aware scaling laws show an even bigger model is economical. When balancing training and lifetime inference costs with less lifetime you can spend more on training and serve a bigger model.

Humans: https://docs.google.com/document/d/1ijQOas5tlP875-NgVmQvInJZ2gHRRYWmj-8IBYZUnKM/edit?tab=t.0 New non-interrupting GUI Delayed Agent Start will be the next big optimization. Make the plan ahead of time, check there’s no silly permission issues or underspecified details, and only start it sometime in the next 6 hours. For AI to work on really hard tasks it has to run really fast or the human becomes the bottleneck again since they have too many concurrent projects.

I predict the primary UI next year will be a queue which on each AI to review or direct your screen completely re-arranges to show all the context that you need, not today’s agent manager. It will automatically and dynamically sequence the next task for your review. After you submit your comments on the current task the agent will no longer immediately run and no longer run at the same latencies. It will be scheduled possibly on a higher-throughput lower latency server, or delayed till off-peak.

Since you should build the priority task queue (as a first step to fully voice control setup), it would also unlock these performance optimizations.

Bullets

Deflation worked example never use Fable-6. (Trival, done but In the example lets use f=8, d=9, g=7 to make math clear and then talk of 10x-ing training runs total as seperate from the time per completed task.)

Metr lines, thoughts on reliablity, llm intel r=0.8-98, thoughts on pricing. (Was done, but I need to include your note on “Correct the cohort argument with demand elasticity” - 2026 ARR is only “same-work spending cohort” so allow for some more tasks to be done. Don’t estimate this or anything it’s not the goal of the essay But I think this doesn’t matter since with all the AI rush now very few tasks are happeneing that wouldn’t if cheaper. If anything people are using AI for things that don’t make sense now but will because they expect AI to get better. )

Model, limits: (Reliability : 1/1-q means only linear cost and assume catches on last run, no verification costs for now)

introducing beta, link to vide-coded estimates of beta and dispersion graph, implications for future based on param)

How 1-param beta model changes if we include a cost for verifying work chunk, and an estimate for human time for each block.

“But is METR x-axis time for unaided human, or time for human? Based on their spring 2025 metric AI actually made developers worse” (and can link to people talking about how there’s no dev acceleration in 2025. ) Footnotes to people estimating scaling runs out by 2030-2032 arbitary cap at Fable-10.

An example is Denard vs Moores Law Example; log-line true but meaning broke. Ask if true x-axis is AI managing AI?

Proposed Changes to model: before treated it as laying line segments where cost proportional to segment. Now we’ll treat it as stacking N-dim blocks where: blocks of volume V have a cost proportional to radius to plan and review for humans, have a cost proportional to volume for the AI to to do the work, and building with blocks of surface area SA has a cost of making sure works with touhcing piece/modularity reduces performance/AI has this cost to manage or work with other AIs. Model version changes the cost for “material” of each work, task decomposition changes both size and dimension of the shape we’re making or how much total cost is volume vs surface area. Info-graphic of nested squares. Better models have a limit to the amount of “permitier” they can handle before managerial capabiltlies tap out

Range of estimates of volume vs SA vs radius costs based on real coordination and real past, what they mean. Also if x-axis is unaided human time or human time period.

Model Implications (check if math makes true and from what params): Swarms make since now since task size lines up with that well, but won’t in future.

Fable-6 is 10x larger than trend since no longer have to do as much inference if it’s managing Fable-5 instances not doing the work directly.

Question: but do you get 1 big Fable-6 managing Fable-5s, or a cloud of Fable-5’s working together? (Note concern that Fable-6 capabilities are required to distil Fable-5 to be even cheaper, but don’t let affect model). And how does this change as we go up to Fable-10 (I assume scaling runs out in 5 more generations) Or 1 Fable-10 managing 10 country blocks which are run by a council of Fable-9s etc on done

Other expectations on organization

Proposals on how to get data to fit My model.

Note Theorietical Concerns: how much deflation is distillation of power next range

Implications on future:

Harness/swarm engineering: it’s only avenue to scale up 4oom, and capabitlies are about at place this model would predict.

Possibility Ai doesn’t pencil due to lower priced competition.

Brief note on but Link out to the human limits, Ai wall clock, WIP, latency, etc essay but say none of these things in the rest of this essay

Please stop using Beta for everything, use different variables.

Model: Reworked to have task be a function of dim D for task, original beta which is model generation costs per volume depending on model and year, the model generation limits on surface area to model limited capabitlies in managerial/interaction/integration work, and human wages per radius (human touch/ai wall clock time are in the seperate WIP essay and don’t get mentioned here). Even if Couplling is more realistic I’d rather have everything use fewer inscruptible variables that combines multiple effects. Ai can have time-horizons grow longer but if they can’t do managerial work then they choke on cost vs cheaper: if N is high and surface area capability limits are low.

Aside from the math about 80% reliability needing 25% more total time we don’t talk about it. But we keep the math of the 16h 80% time horizons so we’re grounded to start.

This is fine footnote for current state of swarms “Meta reports preliminary 40% fewer tool calls after constructing a model-agnostic repository knowledge layer, illustrating “context capital.”

Muse Spark is explicitly trained to adapt to harnesses, plan, delegate subagents, and compact context, making model versus harness improvement increasingly inseparable. StrongDM’s factory relies on scenario-based validation and a digital twin, illustrating how strong verification can replace source review in a bounded domain. Kimi’s reported 100-subagent system supports over 1,500 tool calls and claims up to 4.5× latency reduction, but its own examples emphasize highly parallelizable research and document-processing workloads.”