Will AI Cost Less in the Future, or More?

AI Cost

There is one question executives always ask about AI spending: how much will this cost next year? The answer is not simple. The price per token has clearly fallen over the past two years, and yet at many companies the AI bill keeps climbing. This article explains why costs rise even as unit prices fall, and three reasons the price per token could actually rise from here.

Unit prices are falling.
Your bill still goes up.

First, the facts: unit prices really have come down.

Price cuts are real. Over the past year, the major AI providers have lowered their prices repeatedly. The driver is better inference efficiency. At one large AI company, gross margin improved from minus 94% in 2024 to roughly 60% in 2026. What matters here is that this improvement did not come from raising prices. Margins rose because the same work now takes fewer computing resources.

-94% → about 60%Gross margin at a major AI company (2024 to 2026)
about 7.6xIts annualized revenue (last 7 months)
Not disclosedTokens included in a flat monthly plan

And yet the bill keeps growing.

Costs rise even as unit prices fall because usage is growing faster than prices are falling. That same company's annualized revenue grew about 7.6x in seven months. Growing revenue that fast while cutting prices means the industry is consuming even more tokens than before. The same thing happens inside your company: more people use it, more work is handed to it, and each person uses it longer. If the price halves but consumption triples, the bill still goes up.

The smarter it gets, the more each task consumes.

This is the part people miss. Each new generation of AI is smarter, but that intelligence shows up as extra token consumption. Even if the price list stays the same, you pay more when a job needs more tokens. Here are three changes that are already happening.

The same text now counts as more tokensCountingWhen a model generation changes, the way tokens are counted can change too. In one such transition, the exact same input began to count as 1 to 1.35 times as many tokens as before. The price per token was unchanged, but the bill went up.
Thinking deeper costs moreThinking tokensRecent models have a thinking step before they answer. You can choose how deep that step goes, but deeper thinking consumes more thinking tokens. The more accuracy you want, the more each request costs.
One image now costs about 3x moreImage inputNow that high-resolution images are supported, a single image went from about 1,600 tokens to up to about 4,784 tokens. Accuracy improves, but processing the same number of images costs nearly three times as much.

In short, cost equals price multiplied by consumption, and only the first of the two has fallen in the past two years. The second grows as models get smarter. Seeing a price cut announcement is not yet a reason to relax.

A flat monthly plan is only today's terms.

Flat monthly plans carry an element of introductory pricing, meant to last until AI takes root in daily work. Three facts are worth knowing.

The included token allowance is not publishedGround rulesNone of the major providers publish how many tokens a plan includes. All you get is a multiplier. You cannot build next year's budget on an allowance nobody publishes.
Both limits and prices have changed beforeTrack recordLimits have been raised and prices have been cut, more than once. The fact that they have changed means they can change again. The direction is decided by the provider.
Anything past the cap goes back to list priceMarginal costA flat plan lowers your average price, but usage past the cap is metered. The price of your next token never gets cheaper, no matter how much you use.

On top of that, the major providers are heading toward public listings, a phase where the pressure to show profitability is high. Which way prices move next is not something the customer decides.

You cannot choose the price. You can choose the consumption.

Given all this, the move available to a company is clear. The price is set by someone else. The only thing you control is how many tokens you consume.And if you bring consumption down, it matters less whether prices go up or down. You do not have to bet on a hike, or wait for a cut.

Three practical ways to cut consumption1. Stop re-reading the same material — record what was researched as a summary, and have the next person or AI read that summary instead.2. Use an in-house AI server — routine work such as organizing, classifying and summarizing can run on a model you host yourself, with no metered external charges.3. Cache the context you keep resending — the part where you send the same preamble every time gets far cheaper with caching.

We built this idea into something we actually use in house: a system where the team's AI shares what it researched as summaries, cutting the tokens that used to go into re-reading. It is called SmartOptimizer. Because the summaries are organized on an AI server installed inside our own company, the organizing itself costs nothing in external fees. The tokens saved can be checked as numbers, by month and by project.

No one can say for certain which way AI pricing will move. The only certainty is that the company that reads less twice comes out ahead either way.

Note: figures in this article are approximations as of August 2026, based on public disclosures and press reports. Prices and terms are subject to change. Please check each provider's official information for current terms.

Category:

Tags:

🌐 English