Open Models Are Reshaping AI Economics: Ollama's Founder Reveals the Truth About Enterprise Token Consumption

Ollama CEO reveals open models now drive 80-90% of enterprise token consumption, fueled by coding agents and Flash models.
Ollama co-founder and CEO Jeffrey Morgan, drawing on real token traffic data from 9 million developers, paints an AI enterprise landscape that diverges from mainstream narratives: open models (from both US and Chinese labs) are becoming the dominant force in enterprise token consumption, projected to reach 80–90%. Coding agents and AI collaboration tools are the primary drivers, with context windows leaping from 128K to over one million tokens — pushing Ollama Cloud's overall token consumption up 150x. Morgan foresees a "partner-associate" model split, where frontier closed-source models handle the hardest tasks while open and Flash models absorb the pipeline work, shifting value toward the orchestration and curation layer.
In the race to build AI infrastructure, a trend that's been flying under the radar is quietly transforming enterprise operations: open models are becoming the dominant driver of enterprise token consumption. Jeffrey Morgan, co-founder and CEO of Ollama, recently appeared on YC's podcast The Light Cone to share firsthand data from the front lines of token traffic.
Ollama currently serves 9 million developers, has 178,000 GitHub stars, and is used by 85% of Fortune 500 companies. Morgan has a unique vantage point — he doesn't just know which models score highest on benchmarks; he knows which ones developers actually download and keep using.
The Rise of Open Models: From Cost Pain Points to Customization Capabilities
The biggest shift Morgan has observed is an enterprise-scale migration toward open models — a mix of American and Chinese origin — driven primarily by coding agents and AI assistants in collaborative work settings.
When asked whether enterprise adoption of open models is purely cost-driven, Morgan offered a nuanced answer:
"Cost is the biggest pain point that open models can address. But every company's true North Star is gaining better control over AI and customizing it for their own business. Cost is something you can solve in the short term, but it then enables companies to customize these models for their unique use cases."
He cited a report from The Information about AT&T: the telecom giant has already shifted 40% of its token consumption to open models — primarily American and European models for now, but with Chinese models also under evaluation.

The Real Barriers to Open Model Adoption
You might not have noticed that the core obstacle blocking enterprise adoption of Chinese models isn't capability — it's security and compliance. Based on Ollama's conversations with Western clients, Morgan noted:
"If you can solve the security problem, adopting model labs of Chinese origin is essentially a completely viable option. That's very exciting for these companies."
This also creates a massive opportunity for startups focused on security and governance.
The Token Usage Explosion: The Inflection Point Triggered by Coding Agents
Morgan shared a graph of weekly token usage for individual developers on Ollama Cloud, showing two key inflection points:
- The first inflection appeared early in the year, driven by coding agents. With models like Kimi, GLM, and Minimax launching, open models could finally power coding agents effectively.
- The second inflection came in April, as the explosion of AI collaboration tools allowed non-developers — finance, support, marketing, and sales staff — to hand off complex tasks to open models for automated completion.
The numbers are striking: per-user token consumption jumped roughly 5x, and from the beginning of the year to now, overall token consumption on Ollama Cloud has grown 150x.
The technical enabler behind this is the leap in context windows from 128K to over one million tokens, allowing agents to autonomously decide what tools to use and what data to retrieve across extended timeframes.
Fine-Tuning Revival or Off-the-Shelf Dominance?
One interesting thread in the podcast was about the "cyclical" nature of the AI industry. Early on, everyone was eager to fine-tune custom models — then interest faded, because the results of fine-tuning could easily be wiped out by the next model release. Now that enthusiasm seems to be returning.
Morgan's take was characteristically balanced:
- On one hand, the pace of model releases is accelerating (DeepSeek Flash iterated three times in a single summer, compared to the previous six-month cycle), making custom training harder. The gap between frontier closed-source models and open models is narrowing — possibly to less than three months.
- On the other hand, the toolchain is improving, helping teams that want to fine-tune keep up with the pace.
The more important shift is this: early open models were mostly served as "custom models" (e.g., Cursor fine-tuning DeepSeek or Kimi for large-scale deployment), whereas off-the-shelf open models being served directly have only truly taken off recently.
Ollama's Role: The "Operating System" of the Open Model Era
Morgan used an apt analogy to define Ollama's positioning — it's like a classic operating system that needs to deeply integrate with all drivers and hardware, while ensuring all applications are well-optimized at the application layer.

He broke down a successful "Day Zero model launch" into three components that must all be packaged together:
1. Harness (Framework)
Whether off-the-shelf or SDK-based. There are now many excellent open-source harnesses available, such as the open-source Codex harness and OpenCode, which is popular among Ollama users.
2. The Model Itself
Ensuring the model is available and reliable. For cloud deployments, this means having sufficient capacity, since Day Zero is often the day with the highest traffic spike.
3. Hardware and Providers
Optimizing inference speed through partnerships with NVIDIA, Apple Silicon stacks, and others. "If the model is powerful but slow, that's not a good experience."
Morgan acknowledged this is a combinatorial explosion problem. Many pieces only come together in the final 24 hours before a model launch — "it's usually a fire drill." Ollama's value lies in providing a unified runtime that can match any harness to any model, creating a clearly understandable standard.
Frontier Models vs. Open Models: The Stable Enterprise Equilibrium
On the question of how enterprise spending will ultimately balance between closed-source frontier models and open models, Morgan offered a clear prediction:
"Within the enterprise, the vast majority of tokens will be open models — around 80% to 90%. But that doesn't mean 80% to 90% of the budget goes to open models. In reality, you might spend only 10% to 20% of the cost, but your token volume will be very high."
This creates an elegant organizational analogy: like a law firm where partners (frontier models) delegate work to associates (open models). The hardest tasks go to frontier labs staffed with top researchers, while the high-volume "pipeline" work is handled by open models — both working in concert.
This reinforces an increasingly clear trend: there won't be a single "god model" that rules everything. Instead, multiple smaller, more specialized, simpler models will be linked together through orchestration. Decomposing tasks across models makes results more reproducible, more reliable, and more cost-controllable.
The Local Model Renaissance: A Leap in Desktop Hardware
Ollama was born from running models locally (running open models on a MacBook), which gives Morgan a unique perspective on both worlds.
The latest generation of hardware has dramatically improved the experience of running models in the 20B–120B parameter range. Morgan offered a striking comparison: certain versions of GLM can match Opus-level performance on coding tasks, while running on "the second-cheapest MacBook you can buy from the store."
Even more noteworthy is NVIDIA's DGX Spark and DGX Station. The former offers 128GB of unified memory and can run 20–120B models; the latter, equipped with the GB300, can even be stacked like a mini data center rack to run 400B-scale models — and at a price not dramatically higher than a traditional workstation.
Morgan predicts a hybrid local-cloud model: simple tasks (like document processing) will run locally with lower latency and lower cost approaching zero, while hard tasks like coding agents will remain most effective on large cloud models for now. He predicts that as hardware catches up, local coding experiences (such as sub-100ms autocomplete) will eventually return to the desktop.
Looking at model origins, two contrasting charts are revealing: local model usage is roughly split evenly between American and Chinese models, while coding agents hosted in the cloud are almost 100% Chinese models — which highlights the market's urgent need for more American labs to release large open models (Nemotron Ultra is an early entrant in this wave).
Flash Models: Returning to the Age of "Unlimited Tokens"
For startups with limited budgets, Morgan recommends paying attention to a new class of ultra-low-cost, ultra-low-latency Flash models, such as DeepSeek Flash.
"We all remember the ChatGPT era — you didn't have to think about how many tokens you were using; you could just use it freely every day. Eventually we'll get back to that state."
These models are "good enough for 80% of tasks" and will become the workhorses that handle all the "heavy lifting." More importantly, through the orchestration layer mentioned earlier, multiple Flash models can be chained together to solve more complex problems — a massive opportunity for startups and inference providers alike. DeepSeek Flash is currently the fastest-growing model on Ollama Cloud.
From "Lost in the Wilderness" to Breakout: Ollama's Startup Journey

Ollama's own story is remarkably dramatic. Morgan and co-founder Michael previously built Docker Desktop at Docker, and deeply understood developer experience. They applied to YC with an idea for "Docker Desktop for Kubernetes," but eventually pivoted through multiple iterations — from Kubernetes security to desktop developer security — before landing on Ollama.
That "lost in the wilderness" period was agonizing. Leading a team of 10+ people without a true North Star, Morgan admitted: "Not being able to see your North Star is even more terrifying."
The turning point came from a "start from zero" brainstorm. The team gave themselves two weeks to ship the first version of Ollama, which coincided perfectly with the launch of Llama 2. Driven by a bias for action, they went from idea to launch in two weeks and quickly gained more users than all their previous products combined.
Worth noting: the name "Ollama" doesn't come from the Llama model — it stands for "Open models" + LLM, with a cute animal mascot to complete the brand.
Ollama's GitHub star growth rate far outpaced Docker and Kubernetes. It went from the "enthusiast hobbyist" communities on Reddit to adoption by 85% of Fortune 500 companies in just about 18 months — as the podcast put it: "It took PCs 10 years to go from the Homebrew Computer Club to the mainstream. This time it only took 12–18 months."
The deeper reason: open models are free, can run anywhere, and require no permission to access — making them equally appealing to Fortune 500 IT development teams and hobbyists alike. And since LLMs are stateless, enterprise migration is unusually frictionless.
Curation Is the New Scarcity
Morgan closed by articulating Ollama's core value proposition: when models, inference technology, cloud services, and harnesses form a fragmented universe, bringing all of it together into something that "actually works" is an enormously valuable curation problem.
"When models and providers become extremely abundant, integrating them into something usable becomes the new scarcity."
For developers who just want to build their own software, they shouldn't have to wrestle with cryptic errors from inference providers, undocumented parameters, or JSON formatting landmines. Ollama and services like OpenCode and OpenRouter exist precisely to let developers focus on building the next application — the next company.
Open models aren't just changing technical choices. They're reshaping the entire economics of AI. As tokens become abundant, value is moving up the stack — and how to orchestrate, how to curate, and how to create reliable experiences on top of abundance will be the most exciting battleground ahead.
Related articles

Insufficient Source Material to Generate a Valid Article
The provided source material is a single unrelated tweet with no AI or tech relevance — insufficient to support a complete, valid technical article.

Insufficient Source Material to Generate a Valid AI/Tech Article
This source material is a tweet about the ages of Underworld members — unrelated to AI or tech, and insufficient to support a full article.

Insufficient Material: Unable to Generate a Valid AI/Tech Article
The provided material is a condolence tweet about a San Diego mosque attack — unrelated to AI/tech and too limited to generate a valid technical article.