O
OOMeta
← Back to Insights

September 2026 · 7 min read

1,100-person firm, 50+ agents: ROI is a discipline

1,100-person firm, 50+ agents: ROI is a discipline

Key Definitions

Efficiency ratio (agent) In ABC Legal's accounting, the value an agent delivers measured against what it costs to run, reported in hours and dollars per run. The metric the company tracks for every agent — the same discipline a portfolio manager applies to positions.

Agent J-curve The observed pattern where a new agent runs underwater — costing more than it returns while it uses large models and lacks evals — then flips positive as the team writes evaluations, moves to cheaper and faster models, and trims tokens. Agents graduate, they do not arrive profitable.

Eval gate The threshold an agent must cross to move from human-in-the-loop recommendation to autonomous action: it must prove it is as good as or better than humans on that specific task, using the labeled dataset built from human accept/reject responses. After graduation the agent stays inside the same measurement framework.

The headline numbers of ABC Legal’s agent fleet are the least interesting part of its story. A roughly 1,100-person legal document delivery company with 50+ agents in production and a ~50% cost cut on covered tasks is a vendor-reported case study — useful, but not transferable. What is transferable is the accounting: a company that treats agents like a portfolio, gives every agent an efficiency ratio in hours and dollars, and only lets an agent graduate to autonomy after it passes an eval gate. That is the missing discipline most enterprises never build.

What ABC Legal did: agents with names, owners, and one job each

ABC Legal is a US legal document delivery company (service of process, eFiling, appearance counsel operations). After rolling out Claude Enterprise to its 1,100 employees, teams across the company started building automations on their own. The company then deployed Claude Managed Agents: one common deployment structure, shared workspaces, a single audit and billing surface, always-on agents in the cloud. As of July 2026, the company reports: 50+ agents built in production; up to ~50% reduction in the cost of the human tasks some agents cover, before heavy optimization; and ~310 employees across every department using Claude for daily work (source: Claude.com case study, vendor-reported — https://claude.com/blog/how-abc-legal-turned-every-employee-into-a-builder-with-claude-managed-agents ).

The agents are narrow by design: each has a name, an owner, and a single job. Examples from the case study: the eFiling Rejection Diagnoser fires when a court rejects a filing, reads the job details, checks the court’s rules, and posts a diagnosis to Slack in about a minute — work that used to consume hours. A job-verification agent checks every incoming job against court websites. An AR-remittance agent parses a remittance email, builds the NetSuite payment file, posts it for one-click approval, then imports it. A review agent called Charvis checks completed service jobs and now agrees with the compliance team about 98% of the time.

The ladder: recommendation → labeled data → eval gate → automation

The operational pattern matters more than any single agent. Most agents start with a human in the loop: the agent looks at the job or ticket and makes a recommendation for a person to review before anything is acted on. The recommendation is either stored in the job and surfaced in a banner, or posted to a Slack channel where people reply in the thread. Those responses build a labeled dataset of good and bad calls, which feeds a harvester-and-tuner loop and lets the team write evals and benchmark agents across frontier models (source: same Claude.com case study — https://claude.com/blog/how-abc-legal-turned-every-employee-into-a-builder-with-claude-managed-agents ).

Only after an agent proves it is as good as or better than the humans on that specific task does it shift into automation mode and act on its own — and it stays inside the same measurement framework afterward to watch for changes in performance. This is a graduation ladder, not a launch. The human review gate is doing double duty: it prevents bad actions today, and it produces the labeled dataset that justifies autonomy tomorrow.

The J-curve: agents start underwater and flip positive

The most revealing detail is the accounting. ABC Legal tracks an efficiency ratio for every agent: the value delivered against what it costs to run, reported back to a data warehouse in hours and dollars on each run. The company observes that agents follow a J-curve — they often start underwater while new, running larger models, then flip positive as the team writes evals, moves them to cheaper and faster models, and trims tokens (source: same Claude.com case study — https://claude.com/blog/how-abc-legal-turned-every-employee-into-a-builder-with-claude-managed-agents ).

Read that sentence twice, because it contradicts two common enterprise assumptions. Assumption one: an agent that is not profitable after two weeks is a failed project. Assumption two: cost optimization is a project-phase activity. ABC Legal’s observed pattern says both are wrong: an agent is a position that moves through an underwater phase by design, and the lever that flips it positive — evals, cheaper models, token trimming — is an ongoing operational practice. If you do not have per-agent cost and value numbers, you cannot even see the J-curve, let alone manage it.

Why this is the transferable artifact, not the numbers

Every number in the case study is vendor-reported: the case study lives on Claude.com, Anthropic’s own blog, and there is no independent audit of the ~50% cost reduction or the 310 daily users. Treating those figures as enterprise benchmarks would be a category error. But the accounting method is not vendor-owned. The ladder — recommend, review, label, eval, automate — and the per-agent efficiency ratio are patterns any company can implement with any vendor’s stack, and they are the patterns that separate the firms that reach 50 production agents from the firms that run 50 pilots.

Our judgment: this case is the best recent illustration of the measurement-discipline thesis — the gap between pilot and production in enterprise AI is not a model gap, it is a measurement gap. The company did not wait for a perfect model; it built a review loop that generates the labeled data, an eval gate that decides graduation, and a per-agent accounting line that makes the J-curve visible. That is the same conclusion OOMeta reaches from its own agent infrastructure work: the differentiating capability in agent deployment is not the agent, it is the measurement layer around it.

Our judgment

First, agent ROI is a measurement discipline, not a case study — a static “ROI 3.2x” number is a lagging artifact of a process that must keep running. Second, the recommendation → labeled data → eval gate → automate ladder is the operational core of scaling agents safely; automation is a graduation, not a launch. Third, the J-curve is normal and manageable only with per-agent cost and value accounting; without it, teams either kill agents too early or keep them underwater indefinitely. Fourth, buyer due diligence should ask for the accounting, not the showcase: how does the vendor measure value per agent, how is the labeled dataset built, and what does the eval gate require before autonomy?

A buyer’s action list

First, write the value metric before you write the agent. For each candidate, define what “value” means in hours or dollars per run — if you cannot, you will not know whether it is underwater or profitable. Second, put every agent behind a human review gate first. The gate is not a compliance checkbox; it is the labeled-data factory that makes the eval gate possible. Third, build the eval gate explicitly. An agent graduates to autonomy only when it beats humans on the specific task, measured on the dataset your review loop produced. Fourth, account per agent, not per program. Efficiency ratio per run, reported automatically, is what makes the J-curve visible and the portfolio manageable.

The decision question left for you: if every agent in your organization had an efficiency ratio reported per run, how many would you keep running — and how many would you kill today?

OOMeta AI

OOMeta builds the measurement layer for agent fleets: per-agent efficiency accounting, eval-gate design, and the review loops that turn human feedback into labeled data. We help enterprises run agents as a portfolio, not as a bet.

Schedule a Diagnostic

References: ABC Legal case study, Claude.com (vendor-reported, 2026) — https://claude.com/blog/how-abc-legal-turned-every-employee-into-a-builder-with-claude-managed-agents

FAQ

What did ABC Legal actually deploy?+

ABC Legal, a roughly 1,100-employee US legal document delivery company, rolled out Claude Enterprise and then Claude Managed Agents. As of July 2026 the company reports 50+ agents in production, ~310 employees across every department using Claude daily, and up to ~50% reduction in the cost of the human tasks some agents cover, before heavy optimization.

Are the numbers independently verified?+

No. The case study is published on Claude.com, Anthropic's own blog, so the figures are vendor-reported. That is why the transferable lesson is the accounting method — the ladder and the J-curve — not the specific percentages.

What is the graduation ladder?+

Most agents start with a human in the loop: the agent makes a recommendation, a person accepts or rejects it. Those responses build a labeled dataset that feeds an evaluation and tuning loop. Once an agent proves as good as or better than humans on that task, it shifts into automation mode — and stays inside the same measurement framework afterward.

Why do agents start underwater?+

Because a new agent runs larger models, has no evaluation suite yet, and burns tokens learning the task. The J-curve flips positive as the team writes evals, moves the agent to cheaper and faster models, and trims tokens. ABC Legal tracks an efficiency ratio for every agent in hours and dollars.

What is the decision for a mid-size company?+

Do not start from 'deploy agents.' Start from the accounting: define each candidate agent's value metric, put it behind a human review gate, build the labeled dataset, and only automate after an eval gate says it beats humans. The agent fleet is a portfolio, not a project.