YC President Garry Tan: U.S. Open-Source AI Should Distill Frontier Models

YC's Garry Tan urges U.S. open-source AI labs to distill frontier models, arguing AI capabilities are a public good.
Y Combinator President Garry Tan has publicly called on U.S. open-weight AI labs to distill frontier closed-source models to accelerate open-source capability development. His core argument: since frontier models are trained on humanity's collective public knowledge, their capabilities should be treated as a public good — making distillation ethically justifiable. However, while distillation is technically mature, it sits in a legal gray area, as major closed-source vendors explicitly prohibit using their outputs to train competing products. Tan's stance carries clear industrial policy intent: keeping the open-source AI ecosystem competitive and giving resource-constrained startups affordable access to frontier-grade AI capabilities.
Y Combinator President Garry Tan has put forward a controversial argument: U.S. open-weight AI labs should follow existing industry practice and "distill" frontier models to accelerate capability parity within the open-source ecosystem. This stance strikes at one of the most sensitive nerves in today's AI industry — whether the capabilities of frontier models are private assets, or whether they should be treated as a kind of "public good."

Tan's Core Argument: AI Capability as a "Public Good"
Tan's reasoning is built on a foundational premise: frontier models are themselves trained on humanity's collective public knowledge. Whether it's the GPT series or other large models, the vast majority of their training data — web text, books, code, forum discussions — represents publicly created intellectual wealth produced collectively by humanity.
From this premise, Tan draws the conclusion that since these powerful AIs are "built on the shoulders of humanity's shared public knowledge," access to their capabilities should "become a form of public good."
In other words, if a model's intelligence is derived from the knowledge contributions of all humanity, then locking its capabilities behind a closed-source paywall — available only to paying or authorized users — strikes Tan as having a legitimacy gap. Open-source labs that obtain and release these capabilities through distillation are, in a sense, "returning" value that originated in the public domain.
What Is Model Distillation, and Why Is It Controversial?
Model distillation is an established machine learning technique: a more capable "teacher model" guides the training of a smaller, more efficient "student model," enabling the latter to achieve performance close to the former at a fraction of the cost of training from scratch.
The technique itself is not new, but things get complicated when the "teacher model" is a competitor's closed-source frontier model. The industry has already seen multiple distillation-related controversies — leading vendors almost universally prohibit users from using their model outputs to train competing products in their terms of service. Tan's public call for U.S. open-source labs to "distill too" is therefore effectively encouraging a capability acquisition path that operates in a gray area.
The word "too" is telling. It implies a reality: some labs — including open-source forces abroad — may already be doing this, and Tan believes the U.S. open-source camp should not be hamstrung in this race, but should adopt the same strategy to remain competitive.
A New Front in the Open vs. Closed Source War
At its core, Tan's position extends the long-running open-source vs. closed-source debate into the domain of "how capabilities are acquired." For years, this debate has centered on whether model weights should be made public. The distillation question pushes the conflict even further upstream: how can the open-source camp rapidly close the gap with frontier closed-source models when it lacks comparable resources and compute?
For open-source advocates, distillation offers a pragmatic shortcut. Training a frontier-grade model from scratch requires hundreds of millions of dollars in investment — well beyond what most open-source teams can afford. Through distillation, they can obtain usable capabilities with limited resources, thereby sustaining the vitality of the open-source ecosystem and preserving user choice.
Critics, however, will point out that this approach raises intellectual property and terms-of-service concerns. Leading vendors invest enormous sums in research and alignment; if their results can be easily "extracted," it undermines their incentive to continue investing. Tan's "public good" framework is essentially an attempt to establish ethical justification for that extraction.
What It Means When a Top Accelerator's President Takes This Stance
Garry Tan's position lends considerable weight to these remarks. As president of Y Combinator, he runs Silicon Valley's most influential startup accelerator, which incubates a large number of AI-focused startups each year. His public stance, to some extent, reflects the startup ecosystem's demand for "capability democratization."
For early-stage startups, low-cost access to powerful AI capabilities is directly tied to their survival. If frontier capabilities are monopolized by a handful of giants, entrepreneurs are left building application-layer wrappers within boundaries set by those giants. Open-source models and the distillation pathway could, in theory, give startups greater autonomy and a lower cost structure.
Tan's call is therefore not just a technical or ethical argument — it carries clear industrial policy implications. He wants the U.S. open-source AI ecosystem to remain competitive and to provide the broader startup ecosystem with more open foundational infrastructure.
Questions That Remain Open
It's worth noting that the publicly available information is primarily Tan's core argument, with limited detail on specific operational pathways, legal risk mitigation, or how he would respond to pushback from closed-source vendors. The "public good" argument has moral appeal, but how it translates into legal and commercial reality remains an unresolved question.
This debate around distillation is likely to become a major topic in upcoming open-source AI policy discussions. It touches on technical pathways, intellectual property, competitive fairness, and the accessibility of AI capabilities — each dimension rich enough to fuel extended debate.
Related articles

Vercel AI SDK Releases Vue 3.0.282 Patch Update
Vercel AI SDK releases @ai-sdk/vue@3.0.282 patch update, syncing with core package ai@6.0.282. Learn about the changes, release cadence, and upgrade recommendations.

Vercel AI SDK Sandbox Component Receives Patch Update
Vercel AI SDK releases sandbox-vercel@1.0.109 patch update, syncing the harness dependency to the same version. A look at this maintenance release and what it means for AI app developers.

Vercel AI SDK Vue 4.0.99 Released: Dependency Update Overview
The @ai-sdk/vue 4.0.99 patch release syncs the underlying ai@7.0.99 dependency. Learn what this means for Vue developers building AI apps with Vercel AI SDK.