[KongchangAI]
· 2 min read· 1,119 words

I Gave Four AI Coding Agents a $100 Budget Each: Who Could Build a PDF Editor?

I Gave Four AI Coding Agents a $100 Budget Each: Who Could Build a PDF Editor?

Four AI coding agents each got $100 to build a PDF editor — a real-budget test of autonomous end-to-end delivery.

A developer designed a practically grounded experiment: giving four AI coding agents a $100 budget each to independently build a complete PDF editor. Spanning file parsing, rendering, and UI interaction, the task sits at the boundary between "AI-assisted" and "AI-autonomous" development — making it ideal for testing planning, tool use, and self-correction. The hard budget cap is the experiment's most distinctive feature, forcing agents to trade off exploration against efficiency rather than retrying endlessly. The experiment mirrors a broader shift in AI coding benchmarks: from code completion accuracy toward end-to-end project delivery, and from "can it run a demo" to "can it ship a usable product."

A Real-Budget Experiment in AI Coding Capability

As AI coding agents mature rapidly, a critical question has emerged: when we hand these tools real budgets and genuine autonomy, can they independently deliver a software project of actual value? One developer ran a widely discussed experiment — giving four different coding agents a $100 budget each and asking them to build a PDF editor from scratch.

The core value of this experiment isn't about who wins in the end. It's about stress-testing today's AI coding tools against a complex, multi-stage task in a way that closely mirrors real development scenarios. A PDF editor is no simple "Hello World" project — it involves file parsing, rendering, UI interaction, format compatibility, and more, putting each agent's planning ability, tool-calling capability, and error-correction skills to a serious test.

hackernews source: I gave four coding agents $100 budget to build a PDF editor

Why a PDF Editor?

A PDF editor is a particularly representative test case. Its technical complexity is high enough to be meaningful: the PDF format itself has an intricate structure involving text layers, image layers, vector graphics, and embedded fonts. At the same time, it has clear, verifiable functional goals — a user should be able to open, view, edit, and save a PDF file.

This type of task poses several layers of challenge for AI agents:

Task Planning and Decomposition

A complete PDF editor needs to be broken down into sub-tasks: choosing the right libraries, implementing core rendering logic, building editing features, and constructing the user interface. Whether an agent can map out a sensible development path directly determines how efficiently it uses its budget.

The fundamental difference between an AI coding agent and a traditional code completion tool (like early GitHub Copilot) is that an agent can autonomously call external tools, execute terminal commands, read and write the filesystem, and advance a task through multi-step loops. A typical coding agent first generates a "plan," then calls tools step by step (running code, installing dependencies, searching documentation), and reflects and adjusts course when it hits errors. This "perceive–decide–act" feedback loop makes it fundamentally closer to an autonomous software engineer than a passive code generator. Representative products currently on the market include Anthropic's Claude with tool use, OpenAI's Codex/Operator ecosystem, Devin (Cognition AI), and open-source options like SWE-agent.

Toolchain and Cost Control

The $100 budget introduces a real-world constraint — token consumption and API calls both cost money. The agent doesn't just need to get the features working; it needs to do so within limited resources. This tests decision-making efficiency, not the ability to retry indefinitely.

The running cost of an AI coding agent comes primarily from two sources: LLM API call fees (billed by token) and the number of tool invocations. A complex task often requires dozens to hundreds of "think–act" cycles, and each cycle feeds the model the full current context (code, error logs, conversation history), causing token consumption to grow exponentially as the task progresses. Using GPT-4o or Claude 3.5 Sonnet pricing as a reference, a single run on a mid-sized codebase costing a few dollars is not uncommon. If an agent gets stuck in a loop of repeated failed attempts, a $100 budget can evaporate within hours with nothing to show for it. The budget ceiling, therefore, isn't just a financial control mechanism — it's systemic pressure that forces the agent to make careful decisions.

Debugging and Self-Correction

In real development, broken code is the norm, not the exception. The ability to recognize errors, pinpoint the issue, and self-repair is the key dividing line between a "toy-grade" and a "production-grade" coding agent.

Why the Budget Constraint Matters

Setting a hard $100 budget for AI agents is the most interesting design choice in this experiment. In many AI coding demos, cost is deliberately overlooked — people only ask "did it eventually produce something?" But in real engineering environments, cost is an unavoidable variable.

When budget becomes a constraint, agent behavior changes: it can't retry failed approaches indefinitely and must make trade-offs between exploration and exploitation. This setup more closely mirrors the real calculus that enterprises face when evaluating AI coding tools — not "can it do this?" but "how much does it cost to do this, and how far does that get us?"

Reflections on AI Coding Tool Adoption

This type of comparative experiment reflects a broader trend in AI coding evaluation: benchmarks are shifting from pure "code completion accuracy" toward "end-to-end project delivery capability." Whether a single function is written correctly is no longer the most critical measure. What the industry actually cares about is whether an agent can independently produce a cohesive, runnable, complete project.

Feedback from developer communities suggests these experiments are also sparking discussions about the practical utility of AI agents. Building something of moderate complexity like a PDF editor sits right at the boundary between "AI can assist" and "AI can do it autonomously" — which is exactly why it's such a good tool for observing real capability gaps between different products.

For teams considering adopting AI coding tools, experiments like this offer a valuable reference point: rather than trusting vendor marketing claims, pay attention to delivery quality under real budgets and real tasks. There is often a vast gulf between "getting a demo to run" and "shipping a usable product."

Benchmarking end-to-end project delivery has produced several standardized evaluations, the most influential being SWE-bench — which extracts tasks from real GitHub issues that require modifying a codebase to resolve, asking agents to submit patches that pass the test suite without human intervention. Pass rates on the SWE-bench Verified subset have become a metric that vendors compete fiercely on: in early 2024, top models passed fewer than 10% of tasks; by 2025, some systems have surpassed 50%. A "build a complete app from scratch" task like a PDF editor is more open-ended than SWE-bench and lacks a unified scoring standard, but it's also closer to real engineering for that reason — correctness, usability, code quality, and cost efficiency all need to be weighed simultaneously.

Conclusion

Giving four coding agents $100 each to build a PDF editor may look like an entertaining stunt, but it touches on the core question of AI coding tools' commercial viability — just how far can autonomous coding agents go within a controlled cost envelope? As competition in this space intensifies, real-world comparative evaluations like this will become increasingly important. They reveal what a tool is truly capable of far more honestly than any marketing material ever could.

(Note: This article is based on a discussion thread on Hacker News. Due to limitations in the source material, the specific outcome data from the experiment is pending full disclosure by the original author.)

Share:

Related articles