AI Alignment: But Aligned to Whom?

AI alignment is fundamentally a power question: whoever defines the values shapes the AI's worldview.
The article *Aligned to whom?* challenges a long-obscured contradiction in AI alignment discourse: "making AI conform to human values" in practice often means conforming to the specific value system endorsed by developers. This is a power distribution problem, not a purely technical one. Since human values are inherently plural and incommensurable, building a single universal alignment target is logically untenable. The core deficit isn't whether to align, but who decides the direction — and whether that process is transparent and accountable.
A Question Repeatedly Raised, Rarely Interrogated
AI Alignment has become one of the most frequently invoked terms in tech ethics discussions in recent years. The prevailing industry understanding frames it as making AI systems behave in accordance with human intentions and values. But an article titled Aligned to whom? — which earned 158 upvotes and sparked 95 comments on Hacker News — poses a far sharper question: when we talk about "alignment," whose values are we actually aligning to?
The question sounds simple, yet it cuts to the heart of a core contradiction in AI governance. "Human values" is not a unified, directly encodable objective function. Fundamental differences — and outright conflicts — exist across cultures, communities, and stakeholder groups. When a company declares that its model is "aligned," it has in practice embedded a particular set of value judgments into the system.

The Power Structure Hidden Inside "Alignment"
The article's central argument is this: alignment has never been a purely technical problem — it is a problem of power distribution. Who gets to define what the "correct" values are? The answer almost always points to the institutions that train and deploy these models: large tech companies, their internal safety teams, and the commercial and regulatory forces behind them.
This means that "AI alignment" in practice frequently amounts to "alignment with the value system endorsed by the model's developers." When ordinary users interact with these systems, they implicitly accept a ruleset they likely had no hand in creating. When a model refuses to answer certain questions, or responds to particular topics with a consistent slant, that reflects value choices made by developers — not some objective, neutral "human consensus."
This implicit power structure deserves scrutiny. It transfers ethical judgments that should belong in the public sphere into the hands of a small group of technical elites and corporate decision-makers — a process that typically lacks both transparency and accountability.
The Fundamental Challenge of Value Pluralism
The difficulty of alignment ultimately stems from the plurality and incommensurability of human values. An output deemed "safe" in a Western context may be perceived as biased in another cultural setting. An answer that serves one user community well may actively harm another.
Attempting to construct a "universal alignment" target that satisfies everyone runs into a logical impasse. This is not a technical puzzle solvable through algorithmic optimization — it is a longstanding philosophical and political problem: how do we handle value conflicts in a pluralistic society? Delegating such questions to a single unified model to "resolve" is itself a form of simplification, even distortion.
A more realistic path may be to acknowledge this plurality — to allow different alignment targets to coexist rather than chasing a single, top-down standard answer. But this runs directly against the commercial logic of today's dominant large models: one model, serving the entire world.
Where the Community Agrees and Disagrees
The Hacker News discussion revealed a noticeably divided range of views. A portion of commenters agreed with the article's critique, arguing that current alignment practices do indeed suffer from "delivering value outputs under the guise of neutrality." They called for more transparent disclosure of values and a greater diversity of model choices.
Other voices were more reserved, arguing that even if alignment carries a value orientation, an AI system with clearly defined boundaries is still preferable to one that is entirely unconstrained. Their point: refusing to align doesn't produce genuine neutrality — it just produces a different form of bias or chaos.
The crux of the debate ultimately circles back to the original question — not "should we align or not," but "who gets to decide the direction of alignment, and is that decision-making process open and accountable?"
Implications for the Industry
This conversation surfaces several directions worth serious reflection for the AI industry. Transparency should be a prerequisite for alignment: model developers have a responsibility to articulate which value judgments they have embedded in their systems, rather than packaging those choices as technical inevitabilities.
Choiceability matters just as much. Rather than pursuing a one-size-fits-all alignment standard, the industry should create space for users and societies to choose among different value configurations. This echoes what the open-source model community has long emphasized: autonomy and self-determination.
Most fundamentally, alignment needs to be pulled out of closed-door engineering decisions and returned to the realm of public discourse. AI systems are increasingly mediating access to information and supporting consequential decisions across society. The values they carry should not be determined behind closed doors by a handful of people.
"Aligned to whom?" — this question itself may be closer to the heart of the problem than any technical solution.
Related articles

Waymo AI Team to Host AMA: Focusing on Foundation Models and Autonomous Driving Simulation
Waymo's AI technical leads are hosting an AMA on Reddit's r/MachineLearning, covering foundation models, large-scale simulation, multimodality, and end-to-end autonomous driving architectures.

Docket: Building Per-Commit Evidence Trails for AI Agent-Generated Code
Docket builds per-commit evidence trails for AI agent-generated code, making every AI commit traceable, auditable, and verifiable — a pragmatic step in AI coding governance.

Reverse-Engineering Claude Web's Sandbox: Uncovering Anthropic's Hidden MicroVM
A reverse-engineering analysis of Claude Web's code sandbox reveals Anthropic's likely MicroVM architecture and internal "Antspace" environment, with insights for AI product security.