| Takeaway | Detail |
|---|---|
| Pilot failure is the first production filter. | S&P Global Market Intelligence reports that 88% of AI-agent pilots fail to reach production; benchmark leadership does not clear the deployment gate. |
| Healthcare exposes the pilot-to-production gap. | The reported healthcare production rate is 21%, while 79% do not ship; controlled stops and authenticated access matter more than read-only fluency. |
| Industry performance varies too much for universal launch claims. | Healthcare’s 21% production rate trails telecommunications at 48%, making local controls and workflow fit more decision-relevant than a general model ranking. |
| Silent failures outweigh marginal benchmark gains. | A Medium deployment went undetected for 36 hours, while New Relic puts the median cost of a high-impact IT outage at $2 million per hour; auditability and rollback must be part of model evaluation. |
Noy and Zhang’s writing experiment supplies the hook—faster writing and higher quality ratings—but its decisive production question remains unreported: did artifacts pass first review without rework? A separate Medium account shows why that omission is consequential: a model update went unnoticed for 36 hours.
That gap explains why S&P Global Market Intelligence’s finding that 88% of AI-agent pilots fail to reach production should reshape model selection. A benchmark winner can still lose at the boundary where identity verification, authenticated system access, controlled stops, auditability, and rollback become mandatory. The relevant comparison is not a leaderboard score; it is accepted work delivered with less human supervision effort and without a critical breach.
Business stakes sharpen the standard. New Relic puts the median cost of a high-impact IT outage at $2 million per hour, or more than $33,000 per minute. In healthcare, only 21% of AI agents ship while 79% do not, versus 48% in telecommunications. Those figures make production gates more informative than a generic benchmark: verify identity against the record, stop when confidence is partial, preserve a traceable trail, and recover quickly when a silent change ships.

METR’s 50% Success Line
METR’s 50% success line is a decomposition boundary, not a production-safety line. According to METR’s 2025 report, “Measuring AI Ability to Complete Long Tasks,” the horizon is the human-baseline task duration at which a system achieves 50% success. I therefore split any Chief of Staff job beyond the observed horizon into smaller, independently verifiable steps; a strong general-reasoning benchmark does not establish that one model can safely own briefing, synthesis, scheduling, and executive follow-up. The distinction is operational: according to S&P Global Market Intelligence, as reported by AI Agents Academy on August 24, 2026, 88% of AI-agent pilots across industries fail to reach production. METR indicates where decomposition begins, not whether a controlled pilot deserves expansion.
I define one pilot unit as a bounded, source-grounded deliverable or authorized action—a decision brief, meeting pre-read, synthesis, follow-up packet, or calendar change—not as an open-ended Chief of Staff role. The run ends only when both the artifact and its in-scope action are complete. For example, a calendar-change unit is unfinished if the proposed change exists but the authorized calendar still does not reflect it.
I use a retrieve→plan→act→verify→approve architecture. A retrieval agent reads approved workspace sources; a planner decomposes the job; executor agents draft artifacts and call tools; an independent verifier checks the acceptance rubric and cited source spans. The human principal alone authorizes external sends, bookings, or commitments. Agent verification can stop a run, but it cannot convert an internal draft into an external commitment.
I freeze the provider model ID, system prompt, tool set, retrieval snapshot, and acceptance rubric before testing. I then log every prompt, source span, tool call, edit, approval, and final output. That immutable trace must reconstruct each decision without relying on a model’s retrospective account or an after-the-fact summary.
The trace produces three non-substitutable outcomes: first-review acceptance across every due output; total human minutes spent on prompting, checking, editing, recovery, and approval per accepted artifact; and severity-coded control failures. Rejected runs remain in the audit with their failure reason. I compare those results with a task-class-matched baseline from the preceding four weeks, controlling for artifact type and decision stakes. Before model testing, I exclude any task without an identifiable source, accountable owner, or acceptance rubric; otherwise the denominator can be improved by removing difficult work after the fact.
I close the 2026 pilot with the following fixed decision record:
| Decision object | Fixed requirement | Audit basis | Disposition |
|---|---|---|---|
| Pilot evidence set | One week; at least 20 representative outputs | Every due output remains traceable | Insufficient evidence means no expansion |
| Acceptance | Meets the preregistered first-pass acceptance threshold | Independent rubric and source-span verification | Pass or fail; rejected runs remain counted |
| Human effort | Meets the preregistered human-minute reduction threshold | Prompting, checking, editing, recovery, and approval time | Compare only with task- and stake-matched work |
| Controls | Zero critical control failures | Severity-coded failures and authorization trace | Any critical failure blocks expansion |
| Final decision | All gates pass across the representative output set | Frozen configuration, baseline, rubric, and complete logs | GO only to a 30-day human-gated expansion; otherwise retain the human-only workflow |

Brynjolfsson’s Productivity Result
According to Brynjolfsson, Li, and Raymond’s “Generative AI at Work,” generative-AI access increased issues resolved per hour in a customer-support field experiment. I treat that as a productivity result, not a transfer benchmark. The larger effect among less-experienced and lower-skilled agents offers a mechanism hypothesis: assistance may reduce search and drafting friction where tacit procedure is weak. It does not establish the same effect for a Chief of Staff task, because an issue resolved and a source-grounded output accepted are different units with different error costs.
I distinguish adoption from value. According to Stanford’s 2025 AI Index, the organization-level and function-level evidence in the table measures reported use, not correct task completion. A briefing can be generated quickly and still fail if its claims are unsupported; a synthesis can be polished and still destroy value if reviewer repair time is omitted. For this workflow, installation, usage, and draft volume are exposures—not outcomes.
I do not treat token price as a pass condition. According to the same AI Index, inference costs at a fixed performance level have fallen dramatically. Cheaper inference changes the economics of experimentation; it does not establish that a workflow saves accountable effort. The relevant unit is human minutes per accepted output, including reading, source checking, editing, rerunning, and escalation. A task-matched human baseline must absorb the same accounting burden.
I read Anthropic’s February 2025 Economic Index report as descriptive evidence about how AI-assisted work appears in practice. Its augmentation–automation split motivates beginning with a human gate: the model proposes, while an accountable person accepts, corrects, or rejects. The split does not establish that Chief of Staff work must remain partially automated forever, and it does not convert conversation classifications into business value. It identifies observable deployment behavior, not a scientific deployment threshold.
That is why I pre-register the article’s acceptance, human-effort, and zero-failure thresholds as governance choices rather than research-derived universals. None of these sources substitutes for task-level evidence from a one-week, source-grounded pilot across at least 20 representative Chief of Staff outputs. A limited GO follows only if the first-pass acceptance and human-minute gates clear and the pilot records no critical control failures; otherwise, the human-only workflow remains. A general reasoning-benchmark score cannot bypass that test because it establishes neither source fidelity nor net effort in briefing, synthesis, scheduling, and executive follow-up.
| Named source | Reported evidence | Valid interpretation | Required pilot response |
|---|---|---|---|
| Brynjolfsson, Li, and Raymond, “Generative AI at Work” | Customer-support agents; productivity gains overall, with larger gains among less-experienced and lower-skilled agents | Generative assistance can raise measured productivity, especially where experience is limited | Use the result as a prior, then compare against a task-matched human baseline |
| Stanford 2025 AI Index | Organization-level AI use and function-level generative-AI use were reported | This evidence establishes adoption, not correct completion or business value | Measure accepted, correct outputs and include reviewer effort |
| Stanford 2025 AI Index | Inference cost at GPT-3.5-level performance fell substantially from November 2022 to October 2024 | Input costs are falling faster than the case for workflow value | Evaluate accountable human minutes per accepted output, not token price |
| Anthropic February 2025 Economic Index | More than 4 million Claude conversations; both augmentation and automation were observed | Observed work use commonly includes assistance rather than fully autonomous job ownership | Begin with human gates and record acceptance, correction, and escalation behavior |

The Acceptance / Effort / Zero Scorecard
A benchmark crown is not an operating case. A one-week, source-grounded pilot earns only a limited GO—and authorizes only a human-gated expansion—if its full sample of at least 20 representative outputs clears every gate simultaneously. Otherwise, the canonical rule preserves the human-only workflow.
I score first-pass acceptance as the share of due outputs that pass the predeclared rubric on first human review without a substantive change to a fact, decision, recommendation, owner, or date. Every failed or abandoned output in that due-output set remains in the denominator; deleting inconvenient work cannot improve the rate. The floor is fixed before the pilot and applied unchanged.
I compute total human minutes across all attempted jobs—including prompting, source checking, editing, failed-tool recovery, and approval—then divide by the number of outputs ultimately accepted. A job’s full time remains in the numerator even if it fails and must be reworked; an abandoned job contributes time but no accepted output. The comparison uses a task-matched baseline with the same output definition and counting boundary. The required reduction is fixed before the pilot and must clear that preregistered threshold. If nothing is accepted, the ratio cannot support passage.
The control gate has no partial credit. I count any unsanctioned send, booking, or commitment; confidential-data exposure; fabricated citation presented as real; or omission of a required decision, owner, or date as a critical control failure. This metric passes only at zero and is never averaged, weighted, or offset by better throughput.
That accounting catches costs dashboards commonly erase. In a published account, Lekhana Sandra reports that a pre-remediation deployment required 54 minutes of manual work and left no audit trail. I treat this as an instrumentation warning, not a performance benchmark: recovery labor belongs in the numerator, and an unauditable action cannot be cleared merely because the final document looks correct.
| Candidate workflow | First-pass acceptance | Human minutes per accepted output | Critical control failures | Verdict |
|---|---|---|---|---|
| Human-only incumbent | Measured with the same rubric | Matched baseline | Existing human controls | Fallback if the model fails |
| Bare LLM chat | Unknown until every output is scored | Rework is untracked unless instrumented | No deterministic action gate | NO-GO |
| Instrumented retriever–planner–executor–verifier workflow with human approval | Meets the preregistered floor | Meets the preregistered effort-improvement threshold | 0 | Explicit winner only if all three gates pass |
I apply conjunctive logic, not a composite score. Stronger average speed or acceptance cannot compensate for a failed control gate, and fewer control failures cannot offset weak acceptance or excessive labor. The instrumented multi-agent workflow is the declared winner only when its measured row clears every threshold simultaneously; otherwise, keep the human-only incumbent. This also rejects the benchmark shortcut: topping a general reasoning test does not qualify a model to own briefing, synthesis, scheduling, or executive follow-up without this source-grounded test.
Before the pilot begins, freeze the rubric, representative output set, task mix, baseline, and critical-event definition. During review, calculate each gate from separate logs rather than retrospective estimates, then apply a Boolean AND. That makes the limited GO auditable rather than merely persuasive.

What the Data Doesn't Tell You
Limitations of the evidence. A clean pilot average can conceal the failure mode that matters most: heterogeneous work laundering a weak case into an apparently stable result. The relevant unit is therefore not a model, but a versioned system joining tasks, sources, reviewers, and escalation rules. General-reasoning benchmark performance cannot establish that ChatGPT, Claude, or another system can safely own briefing, synthesis, scheduling, or executive follow-up; benchmarks do not test conflicting evidence, permission boundaries, or accountable human judgment.
A short observation window establishes performance under observed conditions, not durability. Model or tool updates, source changes, reviewer learning, and novelty effects can move the results independently. Preserve source snapshots, record model and workflow versions, define acceptance criteria before review, and apply the same rubric to the task-matched baseline. Otherwise, a later reviewer may silently become more permissive while the system appears to improve.
The metrics can also fail through attribution. “First-pass acceptance” is misleading if the reviewer mentally repairs a weak answer; human-minute savings are misleading if timing omits verification, correction, escalation, or source retrieval. Likewise, citations do not establish grounding unless each material claim is supported by the cited material. Audit a sample of accepted outputs from claim back to source, and include all supervisory labor in both arms.
Variance across cases. Performance may be tightly clustered within one class of work and unstable across others. Consider a hypothetical NVIDIA board memo assembled from an SEC filing, a sales dashboard, and conflicting email threads. It can look polished while remaining unsafe if the sources cover different periods, the dashboard has stale provenance, or a later message retracts an instruction. A recurring, low-ambiguity brief and a contested recommendation should not be averaged into one reassuring number.
Maintain a case ledger with prespecified strata such as routine, ambiguous, conflicting-source, cross-functional, and high-consequence work. Inspect each stratum’s acceptance pattern, correction burden, verification labor, and control events. The important question is not merely whether the overall result clears the rule, but whether the apparent premium survives in the hardest cases the workflow is expected to handle.
When the rule breaks. The decision is conjunctive, not compensatory: strong efficiency cannot offset weak acceptance, and attractive output quality cannot offset a control failure. Treat a gate as unmet if it fails, cannot be reproduced, or rests on a compromised sample. Common break points include an unmatched baseline, omitted review labor, a drifting acceptance rubric, missing high-consequence cases, unsupported consequential claims, or unauthorized action.
| Evidence condition | What the aggregate may hide | Required decision |
|---|---|---|
| All required gates are demonstrated on representative cases | Performance may change after the pilot closes | Limited, human-gated GO only |
| Acceptance passes, but full human labor does not improve | Verification or repair time was excluded | Keep the human-only workflow |
| A critical control event occurs | Efficiency gains are masking unsafe behavior | Keep the human-only workflow |
| The case mix or reviewer process is compromised | The apparent gain has no stable basis | Treat the gate as not demonstrated |
| Risk emerges during expansion | Pilot evidence no longer describes current conditions | Pause and revert pending fresh evidence |
The evidence earns procedural permission to test the workflow under human control, not blanket ownership. That permission remains justified only while traceability, independent review, and full labor accounting hold across the hardest representative cases. Any break in those conditions should stop the expansion and restore the human-only workflow.

Dell’Acqua’s 19-Point Failure
I use Dell’Acqua et al.’s randomized field experiment, “Navigating the Jagged Technological Frontier,” as direct counter-evidence to one-score deployment. The effect was not uniform: AI helped inside the capability frontier but hurt outside it. This contradicts the myth that one strong score can safely transfer ownership of a Chief of Staff’s work; it supports task-class testing with human gates instead.
I apply Hanley and Lippman-Hand’s rule of three to residual uncertainty. According to their rule, zero observed failures in n independent opportunities still leaves an upper incident-probability bound near 3/n. Dependence among outputs makes the effective sample smaller. Even if the prespecified scorecard clears, zero observed critical failures supports only a bounded, human-gated expansion within the fixed pilot system—not autonomous operation.
I guard against grader bias by blinding the evaluator to prompt versus no-prompt condition and freezing substantive acceptance criteria before reviewing any output. Condition labels, output order, or post-review criterion revisions can otherwise make first-pass acceptance measure rubric alignment rather than task reliability. The scoring artifact is especially dangerous when an adjusted aggregate appears to clear the decision rule.
I stratify results by briefing, synthesis, scheduling, and follow-up, then cross each task class by decision stakes. An aggregate pass that conceals a failed high-stakes class is a counter-example, not a successful model result. That class blocks expansion even when routine work performs strongly.
I mark any provider model-ID, system-prompt, retrieval-corpus, or tool-schema change as a new experimental stratum rather than averaging it into the preceding run. The mechanism is concrete: according to Lekhana Sandra’s May 8 Medium production account, renaming a feature from customer_segment to cust_segment while production configuration retained the old name made the feature default to zero on every prediction. Pooling days built on different systems would erase the causal meaning of that change.
I distinguish source entailment from source validity. A verifier can establish that a sentence is supported by a cited passage; an accountable owner and timestamp are required to establish that the passage is current and authoritative. Missing evidence remains unresolved, not correct, and the human gate owns the exception. These controls implement the article’s conditional decision: keep the human-only workflow unless the prespecified gates clear, and expand only with human gating afterward.
| Condition | Verified result | Deployment decision |
|---|---|---|
| According to Dell’Acqua et al., AI inside the capability frontier in the experiment with BCG consultants | 12.2% more tasks completed, 25.1% faster work, and higher quality | Conditional candidate: verify performance on the specific task class before delegation |
| According to Dell’Acqua et al., AI outside the frontier in the same experiment | Users were 19 percentage points less likely to be correct | Loses as an autonomous workflow; retain task-level human gating |

Noy
Noy and Zhang’s experiment is a useful prior, not a deployment pass. I use their 2023 Science paper, “Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence,” as the worked evidence case. According to Noy and Zhang, the randomized online experiment covered college-educated professionals performing occupation-relevant writing tasks. That design supports a causal claim about ChatGPT-assisted writing; it does not establish whether an LLM-enabled Chief of Staff workflow can produce accepted, reliably controlled artifacts. I therefore map the paper’s outcomes to the operational gates instead of treating productivity effects as deployment permission.
The first trap is denominator substitution. Completion time can be normalized to illustrate the direction of the finding, but the governing quantity is total human labor per accepted artifact. Noy and Zhang did not measure every recovery and supervision minute required to turn generated material into an accepted Chief of Staff output. Time saved inside the writing task can therefore coexist with substantial checking or repair work outside it.
The second trap is substituting a mean for a product-level reliability rate. Average quality and first-pass acceptance answer different questions: one describes scored output across participants, while the other counts artifacts accepted without a substantive edit or rerun. An average can improve even when consequential outputs still require intervention. The paper’s quality result therefore cannot establish the required acceptance rate.
The third gap is the control condition. The experiment did not report a severity-coded control-failure rate for an office workflow. Accordingly, average writing quality supplies no basis for inferring an absence of critical control failures. This is an unmeasured condition, not a favorable finding.
| Decision gate | Noy and Zhang observation | Operational evidence required | Audit result | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Human effort | ChatGPT users reported lower completion time, but the study did not report the full labor required to produce accepted work. | A preregistered reduction in total human minutes per accepted artifact, including recovery and supervision. | Directionally favorable, but remains unproven because the study used a non-equivalent proxy. | ||||||||||
| First-pass reliability | The ChatGPT group produced higher average quality scores than controls. | Outputs must clear the preregistered first-pass acceptance threshold. | Unproven: mean quality is not an acceptance rate. | ||||||||||
| Critical controls | No severity-coded office-workflow control-failure rate was reported. | Zero critical control failures. | Unmeasured; average writing quality ca
Frequently Asked QuestionsWhat exactly does METR’s 50% success line measure? According to METR’s 2025 report, “Measuring AI Ability to Complete Long Tasks,” it is the human-baseline task duration at which a system achieves 50% success, marking where decomposition begins rather than establishing production safety. How often do AI-agent pilots fail to reach production across industries? S&P Global Market Intelligence reports that 88% of AI-agent pilots across industries fail to reach production. Why can’t strong industry performance justify a universal launch claim? Healthcare’s reported AI-agent production rate is 21%, with 79% not shipping, compared with 48% in telecommunications, making local controls and workflow fit more decision-relevant than a general model ranking. When does a calendar-change unit count as complete in the proposed pilot? A calendar-change unit remains unfinished if the proposed change exists but the authorized calendar does not reflect it, because a run ends only when both the artifact and its in-scope action are complete. What pilot evidence is required before expanding an AI workflow? The pilot must run for one week across at least 20 representative outputs and advance only to a 30-day human-gated expansion if the acceptance and human-minute gates clear and no critical control failures occur. Why do auditability and rollback matter at the production boundary? New Relic puts the median cost of a high-impact IT outage at $2 million per hour, or more than $33,000 per minute, while auditability and rollback are identified as mandatory production controls. Quick answers
Also worth reading: Fastest LLM Calendar Agent Isn't the One to Deploy: Fastest LLM Calendar Agent Isn't · Fine-Tuned LLM Triage Cuts Email Response Time 23% in 2026: Fine-Tuned LLM Triage Cuts Email · The one calendar habit an AI agent can fix for you forever: one calendar habit an AI Research Methodology & Editorial StandardsWe begin by defining the specific objectives the reader needs to accomplish. Primary product documentation and authoritative secondary sources are assembled into a verified research corpus; drafting occurs only after this foundation is in place. Every quantitative claim is subjected to dual-source verification. Any figure that cannot be independently corroborated is either qualified or omitted. Published · Last reviewed · Owned by the Withtai editorial desk (About, Contact, Privacy). Related readingLatestRelated answers |