The Missing Layer in Enterprise AI Agents: Why Harness Design Matters More Than Model Upgrades
Harvey's experiment showed that improving the system around an AI agent, not upgrading the model, drove task performance from 40.8% to 87.7% across complex legal work. The lesson: harness > model.
Introduction
Most enterprise AI deployments follow a familiar pattern: give the model a prompt, get a response, have a human check the work. It’s useful, but it’s fundamentally a faster version of what people already do. Harvey, the legal AI platform used by major law firms and in-house teams worldwide, just demonstrated something different: agents that improve their own performance through structured self-experimentation, without updating the underlying model.[1]
The implications extend well beyond law. What Harvey has shown is an early blueprint for how enterprises can deploy AI agents that get meaningfully better at domain-specific work: automatically, measurably, and at scale. And it points to a strategic reframe that too few executive teams have internalized: the constraint on agent performance is increasingly not the model itself, but the system around it.[1][2]
The Experiment: Agents That Coach Themselves
In a paper published by Niko Grupen, Head of Applied Research at Harvey, the team described an experiment combining two techniques: autoresearch, where an agent runs its own experimentation loop, and harness engineering, where an agent’s capabilities are shaped by its environment and feedback loops rather than by changes to model weights.[1]
The setup: twelve complex legal tasks (commercial lease review, complaint drafting, tax memos, due-diligence questionnaire responses), each with source documents, instructions, and a detailed grading rubric. The agent attempts a task, gets scored by an LLM judge against the rubric, receives structured feedback, then a coding agent reads that feedback, clusters the failures, hypothesizes what harness improvements would help, builds them, and reruns the task.[1]
The results were striking. Baseline agents with generic harnesses started five of the twelve tasks between 2–7% success. After optimization, the average score across all tasks moved from 40.8% to 87.7%. Seven of twelve tasks finished above 90%. One reached 100%.[1]
Why This Matters for Enterprise AI
The standard enterprise AI playbook (fine-tune a model, deploy it, monitor drift) treats AI as a static capability. Harvey’s experiment points to a different architecture: agents that develop domain-specific skills through structured practice, much like a junior employee who improves through feedback and repetition.[1][3]
Three things make this approach distinctive. First, the model itself doesn’t change. The improvements come from the agent’s harness: its tools, prompts, review playbooks, and output pipelines. That means enterprises don’t need to retrain or fine-tune anything. Second, the learning is auditable. Each improvement is a discrete, inspectable change (a new cross-document review playbook, a validation hook, a file-conversion pipeline), not an opaque weight update. Third, the quality bar is set by humans. The rubric is what drives the agent’s improvement cycle. As Grupen puts it: “Humans steer. Agents execute.”[1]
The Organizational Shift: From Throughput to Judgment
Harvey’s president, Gabe Pereyra, recently described the broader organizational implications in a company blog post. His argument: as agents take on more execution, the bottleneck shifts from throughput to coordination and judgment. Organizations built as information-routing hierarchies, where managers aggregate context and route decisions, must rethink their structure when agents can carry context, trigger work, and surface decisions independently.[2]
For law firms, this means the pyramid model, where junior associates handle high-volume, lower-complexity work, faces pressure. When agents can draft, review, and iterate on complex deliverables, every lawyer becomes valued primarily for judgment, not output. Firms must rethink staffing, apprenticeship, pricing, and practice-area structure.[2]
For in-house legal teams, the shift is twofold: they must transform their own work and govern how the rest of the organization deploys agents. Legal will increasingly define the boundaries of what autonomous systems can do: where accountability sits, what risks are acceptable, and what governance is required.[2]
This raises a set of management questions that are distinct from model selection. Who sets the rubric? Who signs off on acceptable error rates? Who owns exception handling? Who updates the workflow when business rules change? Who decides whether an agent’s output is merely useful or operationally safe? These are organizational design questions, and they will determine whether agents scale or stall.[2][3]
The Pattern Beyond Law
Harvey’s experiment is legal-specific, but the pattern (structured rubrics, automated evaluation, iterative harness improvement) transfers to any domain with clear quality criteria.[1][3]
Consider insurance claims adjudication, where an agent could be scored against rubrics for coverage determination accuracy, documentation completeness, and regulatory compliance, then iteratively improve its review playbooks. Or financial services compliance, where a regulatory change triggers an agent to re-evaluate a portfolio of client accounts against updated rules, with structured feedback loops catching misclassifications. Or procurement, where agents evaluating supplier bids could develop domain-specific scoring frameworks through rubric-based iteration, learning to weight delivery reliability, cost structure, and compliance history the way a seasoned sourcing manager would.[1][3]
The critical enabler is the rubric itself. Harvey’s team found that when rubrics are high quality, agents can improve dramatically. When they’re vague or incomplete, the feedback loop stalls. For enterprise leaders, this reframes the AI investment question: the constraint isn’t the model; it’s whether your organization can articulate, in sufficient detail, what “good” looks like for a given task.[1]
Executive Guidance
Invest in rubrics, not just models. The organizations that will benefit most from self-improving agents are those that can codify quality standards into machine-readable evaluation criteria. Start with your highest-volume, expert-dependent workflows and build grading rubrics for them now.[1][3]
Think harnesses, not prompts. Prompt engineering gets you a better single response. Harness engineering (the tools, validation hooks, review playbooks, and output pipelines surrounding the agent) gets you a better system. Allocate engineering resources to the scaffolding, not just the instructions.[1][3]
Move from “humans in the loop” to “humans on the loop.” Inspecting and correcting individual outputs does not scale. The higher-leverage design is to have humans improve the harness that produced the output: refining rubrics, updating exception taxonomies, and tuning validation gates. That shifts effort from one-off correction to system improvement.[1][2]
Build proprietary advantage through institutional context. Frontier models are diffusing quickly. What is harder to replicate is a company’s task library, rubric set, exception taxonomy, workflow memory, and domain-specific improvement loop. That is where enterprise learning compounds, and where competitive moats will form.[1][3]
Prepare for the judgment bottleneck. As agents handle more execution, your organization’s scarcest resource becomes expert review and decision-making capacity. Redesign workflows to concentrate human attention on the decisions that matter most, and build trust frameworks that let agents operate with appropriate autonomy.[2]
Govern early. Self-improving agents raise new governance questions: who approves the harness changes, how are improvements validated, and what oversight is required? Establish these policies before scaling, not after.[2][3]
The first generation of enterprise AI was about making individuals faster. The next is about building systems that get better on their own, within boundaries set by human expertise. Harvey’s experiment is small-scale, but it demonstrates a principle that will reshape how enterprises deploy AI: the agents that learn are the agents that last.
Sources
[1] Auto-Research for Legal Agents. Published on X, Grupen, N., 2026
[2] How Autonomous Agents Will Transform Legal, Pereyra, G., Harvey Blog, 2026
[3] Harvey Drives Legal Agent Learning Via ‘Harness Engineering,’ Artificial Lawyer, 2026)
Join the AI Realized Community
If you are an executive adopting AI, you are invited to join the AI Realized Community and meet peers, attend events, and enjoy content curated for the leaders of Enterprise AI.


