$0.58 vs $2.98: Glean Just Moved the Enterprise AI Argument to the Invoice

Announced today at Glean:GO 2026 in San Francisco Announced today at Glean:GO 2026 in San Francisco. Image credit: Glean.

Reporting live from Glean:GO 2026 at Fort Mason, San Francisco. By Ravit Jain, Founder and Host, The Ravit Show.

For 2 years, the enterprise AI conversation has been about intelligence. Which model reasons better. Which one scores higher.

Today at Glean:GO 2026, Glean changed the subject. It published a benchmark that is not about how smart the AI is. It is about what one completed task costs.

I saw this data before the keynote, at the VIP analyst and media session on day 1. I have sat through a lot of vendor benchmark slides. This one made the room lean in, because it is measuring something almost nobody publishes.

Here is the number, and then here is the part that actually matters.


The headline

Cost per task: Glean vs Claude Cowork

$0.58 per task on Glean. $2.98 per task on Claude Cowork.

That is an 81 percent reduction in token cost, or roughly 5.1 times cheaper per query.

And the quality claim runs in the same direction rather than against it.

Grader preference: Glean vs Claude Cowork

Graders picked the Glean response 78 percent of the time. Claude Cowork was preferred 22 percent of the time. That is about 3.6 times as often.

The usual tradeoff in this category is cheaper or better, pick 1. Glean is claiming both at once, and the interesting question is why that would even be possible.

Do the arithmetic before you react

Numbers per query feel small. They are not small.

Take an enterprise running 100,000 agentic tasks a month. That is not an aggressive number for a company with 10,000 employees once agents are actually in production.

At $2.98 per task, that is $298,000 a month, or roughly $3.6 million a year. At $0.58 per task, that is $58,000 a month, or roughly $700,000 a year.

The gap is close to $2.9 million a year on the same volume of work. That calculation is mine, not Glean's, and you should run it with your own task volume before you quote it. But the point holds. At any real scale, the unit economics stop being a procurement detail and start being the architecture decision.

This is why I keep telling data leaders in my community that cost per outcome is the metric to instrument in 2026. Not tokens. Not seat price. Cost per completed unit of work.

Why the gap exists: it is 2 variables, not 1

This is the part that separates people who read the headline from people who understand the system.

Cost per task decomposes cleanly:

tokens consumed x blended rate per million tokens

Glean's advantage shows up on both variables at the same time, which is why the gap compounds instead of just being a discount.

Variable 1: token volume

Glean's argument is that a pre built index and knowledge graph mean the agent does not have to rediscover the company on every single request.

Federated retrieval through generic connectors works differently. The agent searches across systems in real time, pulls back more than it needs, and carries that dead weight forward into every subsequent step. Each imprecise retrieval widens the context window, and a wide context window is a bill.

So the same task converges in fewer turns with tighter prompts. This is not the model being less verbose. It is retrieval precision showing up on the meter.

Variable 2: blended rate

Auto routing sits between the Model Hub and the AI Gateway. This is where the blended rate is decided.

The second variable is which model runs which step.

A single vendor stack tends to concentrate token volume on 1 model family, which puts you at 1 point on the capability and cost curve for everything you do. Summarizing a document and reasoning through a multi step account plan get billed at similar rates, even though only 1 of them needed the reasoning.

Glean's auto routing spreads work across a Model Hub of 40 plus open and frontier models. Commodity steps go to cheap capable models. Frontier models get reserved for the steps where the extra capability changes the outcome.

Here is the counterintuitive detail Glean's own engineers highlighted around this benchmark. The cheaper system used the expensive frontier model more often, not less. It just used it surgically. That only makes sense if you optimize at the task level rather than the token level.

That is a routing strategy, not a model strategy. And routing is an architecture choice, not a purchasing choice.

The context layer is an economic asset

The trade: build the index and graph once, then serve context to every surface.

The deeper argument Glean has been making across 3 benchmarks now is that an index is an economic instrument.

You pay a small, predictable storage cost up front. In exchange you reduce a large, variable, and rapidly rising compute cost, because the agent stops paying to rebuild context on every request.

David Lee at LegalZoom framed it well in Glean's materials. On other platforms you burn an enormous number of tokens recreating the same depth of context, while an already indexed environment delivers it immediately.

This is a real inversion. For most of the last decade, the index was a cost center you justified with search quality. In the agent era, the index is how you control your compute bill.

This is Glean's third benchmark, and the pattern matters

Do not read this in isolation. Glean has been building a case in public.

February 2026: Glean published a search evaluation showing its results preferred roughly 2 times more than ChatGPT and 1.6 times more than Claude, holding models constant to isolate for context quality.

May 2026: Glean published an MCP evaluation that is, in my view, the most methodologically clean of the 3. It held the harness constant, using Claude Cowork with Claude Sonnet 4.6, and swapped only the context layer behind MCP across roughly 175 real queries. Glean's MCP server against off the shelf servers for Drive, Gmail, Slack, Salesforce, GitHub, Calendar and others. Glean was preferred about 2.5 times as often and off the shelf tools consumed about 30 percent more tokens. The win rate widened as tasks got more complex, from 66 percent on simple tasks to 73 percent on multi step cross source work.

August 2026: today's benchmark, which measures the full stack including the harness and routing, not just the context layer.

That progression is deliberate. First prove context quality. Then prove context efficiency. Then prove the whole system economics. As a piece of analyst relations, it is genuinely well built.

Now the caveats, because you deserve them

I am covering this as a journalist, not as a reseller. So here is what I would put to Glean, and what I would put to any vendor publishing a comparative benchmark.

1. This is a vendor run benchmark on a competitor. Glean designed the tasks, ran both sides, and graded the results. That does not make it wrong. Glean has published methodology in more detail than most of its peers, which I credit. But a self reported comparison is evidence, not a verdict.

2. Preference scoring is a judgment metric. A 5 point preference scale across utility, correctness, completeness, and tool fidelity is a reasonable design. It is still humans deciding what they would use. Ask who the graders were and whether the evaluation was blind and randomized.

3. Configuration decides a lot. In any comparison like this, the losing side's setup matters enormously. Which connectors were enabled, how the competing system was configured, and whether it was tuned by people who use it daily are all fair questions.

4. Query mix drives the result. Glean's own earlier data showed the gap widening on multi step cross source work. If your actual workload is single source and simple, your gap will be smaller than the headline. If it is complex and cross system, it could be larger.

5. Anthropic has not responded publicly at the time of writing. I will cover their response if it comes.

None of this makes the benchmark unimportant. It makes it a starting point for your own evaluation rather than a substitute for one.

What I would actually do with this

If you are a CDO, CIO, or head of platform, here is my practical advice coming out of day 1 at Glean:GO.

Instrument cost per task now. Most enterprises cannot answer "what did that completed task cost us" for any agent in production. Until you can, every vendor conversation is theoretical.

Run the benchmark on your own workload. Take 50 real queries your teams actually run. Not demo queries. Run them across your candidate stacks. Score preference blind. Log tokens. It is 2 weeks of work and it will be the most useful 2 weeks of your year.

Separate the 2 variables. When you find a cost gap, decompose it. Is it token volume, which is a context and retrieval problem, or blended rate, which is a routing problem. They have completely different fixes and completely different vendors.

Do not let cheaper beat correct. A cheap wrong answer that a human has to redo is the most expensive output in the building. The reason this benchmark is interesting is that both metrics moved together. If you ever see them move apart, quality wins.

My take

The industry spent 2 years benchmarking models. We are now benchmarking systems. That is a sign of a category growing up.

The strategic point Glean is making, and I think it is correct regardless of whether you buy their product, is that models are becoming interchangeable and your company's context is not. Once models commoditize, the durable advantage sits in the layer that knows your business, enforces your permissions, and decides which model touches which step.

Whether Glean holds that layer, or Microsoft, or Anthropic, or someone not yet in the conversation, is the open question of the next 18 months.

But the argument itself has shifted from "whose AI is smarter" to "whose system costs less to run correctly." That is a much healthier question for enterprise buyers, and today is the clearest example of it I have seen.

I am here at Glean:GO through day 2, bringing you the conversations behind these numbers on The Ravit Show.

Sources and further reading