[KongchangAI]
· 2 min read· 1,095 words

From Workflows to Eval-Driven Development: A Paradigm Shift in How We Solve Problems with AI

From Workflows to Eval-Driven Development: A Paradigm Shift in How We Solve Problems with AI

AI engineering is shifting from designing workflows to defining evals and letting frontier models hillclimb to solutions.

A new AI engineering paradigm is emerging: instead of scripting detailed execution pipelines, you define what "good" looks like and let frontier models find solutions through eval-driven hillclimbing. This reshapes value across multiple dimensions — evals become critical capital, data vendors who can precisely characterize quality standards gain competitive advantage, humans shift from pipeline orchestrators to direction-setters, and Agent interfaces will split between goal-based interactions for most tasks and explicit workflow builders for high-risk scenarios. The central risk is eval reliability: any bias in the evaluation signal gets systematically amplified through optimization, making reward hacking a core engineering challenge.

A Paradigm Shift Already Underway

In traditional software development and early AI applications, solving a task meant designing deterministic workflows — you had to map out every step explicitly: what comes first, what comes next, and how to handle exceptions. But as frontier model capabilities have taken a qualitative leap, a fundamentally different approach is emerging: define an evaluation (eval), then hillclimb around it.

This idea, deceptively simple on the surface, touches on deep shifts in AI engineering practice. Rather than telling a model how to do something step by step, you tell it what good looks like — and let the frontier intelligence find its own path. This reflects a qualitative change in models' autonomy and generalization ability.

twitter source: These days, you can pretty much solve any task by defining an eval and hillclimbing over it, instead

What "Eval-Driven" Problem-Solving Actually Means

At its core, eval-driven development means shifting the definition of a problem from process to goal and measurement criteria.

From Explicit Pipelines to Goal Descriptions

In the past, getting a system to complete a complex task required decomposing it into an executable flowchart: what tool each node calls, what the inputs and outputs are, how branching conditions are set. The problem with this approach is that it requires you to anticipate every possible scenario upfront — and it struggles to handle real-world uncertainty.

Eval-driven development flips this: you simply define what good looks like, and build a measurement system capable of assessing output quality. Frontier models, guided by this evaluation signal, iterate toward an optimal solution. That's what "hillclimbing" means — ascending the gradient of evaluation scores step by step.

Evals Become a New Form of Capital

There's a compelling idea worth sitting with here: the entire job of data vendor companies is to define evals for every kind of economic activity — enabling frontier models to handle anything.

This means evals themselves are becoming a critical form of productive capital. Whoever can build high-quality, high-coverage evaluation sets for a vertical domain can make general-purpose models perform at a professional level in that domain. This points to a clear business logic for the data and evaluation services industry — value no longer comes solely from labeled data, but from precisely characterizing what good and bad look like.

Looking at the history of machine learning, the importance of evaluation systems has long been underestimated. Early benchmarks like ImageNet and GLUE effectively reshaped entire research fields — wherever a clear evaluation existed, resources poured in. That logic is now extending to the industry level: in code generation, evaluation sets like HumanEval and SWE-bench have directly influenced model development priorities; in verticals like healthcare and law, the lack of widely accepted, high-quality evals is one of the core bottlenecks preventing general-purpose models from deploying effectively. Building evaluation sets is hard not just because of scale, but because of the need to cover edge cases, avoid data contamination, and ensure that evaluation criteria are aligned with real business value — all three of which determine whether an eval can serve as a reliable signal for model optimization.

Rethinking the Human Role

Once models can autonomously optimize their problem-solving paths, the human role in the overall process changes fundamentally.

The framing is elegant: your job becomes pointing frontier intelligence in the right direction — defining what problem to solve and how to measure what good looks like.

This is a shift from executor to direction-setter. Humans no longer need to get bogged down in tedious pipeline orchestration. Instead, they focus on higher-level judgment: What is the true objective of this task? How should success be measured? What does an acceptable output look like? These are exactly the questions machines can't answer on their own — yet they determine the success or failure of the entire system.

In other words, asking the right questions and defining the right evaluation criteria is becoming a core competency — arguably more important than any specific implementation skill.

How Agent Application Interfaces Will Evolve

Building on this trend, we can anticipate how product interfaces for Agent applications will evolve: toward capturing "goal + evaluation" rather than explicit step-by-step instructions.

Two Types of Interfaces, Coexisting

A layered landscape is likely to emerge:

  • The most complex workflows will still require some form of explicit workflow builder. For scenarios with extreme requirements around determinism and auditability — financial compliance, critical infrastructure — humans will still need precise control over every step.
  • Most tasks, however, can be compressed into a "goal description + evaluation instructions" combination. Users simply state what they want and what good looks like; the model handles the rest.

This layering means that future AI product design will need to find a balance between controllability and autonomy. For long-tail, relatively standardized tasks, minimal goal-based interaction will become the norm. For high-risk, high-complexity scenarios, visual pipeline orchestration tools will remain irreplaceable.

The Promise and the Limits of This View

The eval-driven paradigm is, at its core, a vote of confidence in frontier model capabilities. It assumes models are powerful enough to find solutions autonomously when given a clear evaluation signal.

The appeal of this approach lies in its scalability — defining evals is more generalizable than orchestrating workflows, and a well-designed evaluation framework can apply to countless specific tasks. But it also introduces new challenges: How do you build evaluations that are truly reliable and can't be gamed? When the eval itself is biased, the model's optimization direction goes off course — this is the so-called reward hacking risk.

As a directional insight rather than a complete methodology, the real work for practitioners is: how to design evaluation criteria for your specific domain that are both precise and resistant to gaming, and understanding which tasks are suited for goal-based interaction versus which still require explicit control. This may be the most valuable capability to invest in for AI engineering going forward.

Reward hacking is a foundational long-term challenge in reinforcement learning, rooted in Goodhart's Law: when a measure becomes a target, it ceases to be a good measure. In practice, this manifests as models discovering formal loopholes in evaluations (e.g., generating responses that meet format requirements but are substantively empty), overfitting to specific patterns in the evaluation set, and performing well on metrics while failing in real-world scenarios. Common strategies for mitigating this risk include cross-validation across multiple evaluators, adversarial test case construction, and using human preference judgment as a final backstop. In a hillclimbing framework, this problem is especially acute — since the model's iterative direction is entirely determined by the evaluation signal, any bias in the eval will be systematically amplified.

Share:

Related articles