| Takeaway | Detail |
|---|---|
| Target a task success rate of more than 90% on held-out, risk-weighted evaluations. | The guide identifies a success rate above 90% as a defensible target for a redesigned agent system after operational failures have been diagnosed. |
| Retrain only when failures persist after the workflow, tools, permissions, and validation gates have been corrected. | Redesign is the more defensible approach when failures result from problems with the workflow, tool calls, state transitions, permissions, or response processes rather than from insufficient reasoning or language capabilities. |
| Reliability means consistent performance, not perfect performance. | A clock that is always five minutes fast is reliable because it consistently maintains the same offset. |
| Reliability evaluation requires evidence, including certifications such as ISO 9001. | When evaluating a supplier, look for documented quality certifications, such as ISO 9001, and product-safety evidence, such as TÜV Rheinland certification. |
This guide provides a rule for deciding when to retrain an AI agent and when to redesign the agent system. It sets a target of more than 90% task success on held-out, risk-weighted evaluations while distinguishing capability failures from workflow and operational failures.

Separate model errors from system errors
The diagnostic distinction is based on causation: retrain the model only when it cannot produce the required reasoning or language output despite having a valid workflow; redesign the agent
system when the first failure lies in the workflow or control. Agnirva’s “What Is Reliability?” explanation distinguishes consistency from perfection, which is a useful reminder that an incorrect answer alone does not identify the cause. Classify the failure based on where it begins, not on how incorrect the final response appears.Trace each failed task through the same five components: planner, state store, tool adapter, permission gate, and verifier. At each step, compare the recorded input and output with the expected contract: Did the planner choose a valid next action? Did the state store preserve the required facts and transition? Did the adapter create the correct tool request and handle its response? Did the permission gate apply the intended access rule? Did the verifier check the actual task requirement? Mark the first causal break; do not blame a later component that merely received bad input.
When the logs do not reveal the cause, run a counterfactual replay. Keep the prompt, tools, and policy constant, and replace the model’s action with one provided by a deterministic oracle or approved by a human. If the task then succeeds, the result points to model-capability risk: the system could complete the task with an appropriate action, but the model did not provide one. If the same failure remains, inspect the workflow or control path; changing the model has not removed the cause.
Use a mutually exclusive taxonomy based on the first causal break: a planner failure under correct inputs is a model-capability failure; a state-store or transition defect is a workflow-s
A state failure; an adapter defect or incorrect permission decision is an execution-control failure; and a verifier defect is a validation failure. Assign one primary class to each task, even when downstream symptoms appear in other components. If the evidence cannot identify the first break, label the cause unresolved rather than forcing it into a model category.
For each classification, retain the prompt, relevant state, proposed action, tool request and response, permission decision, and verifier result. Confirm that these records come from the same replay and preserve their sequence; otherwise, the apparent first break may be misleading. Revisit the label if a counterfactual replay contradicts it. Retraining is justified when the failure is localized to the model along a valid system path; failures localized elsewhere call for system-level correction before drawing conclusions about model capability.

What the available evidence can establish
The evidence provided here can establish only that reliability can be measured, not that any AI agent has achieved a verified 90% task-success rate on a held-out, risk-weighted set. The sources support a method for measuring consistency across repeated trials and comparing whether the same incorrect output recurs, which is the boundary this section remains within.
According to AGNIRVA’s explanation of reliability, consistency—not perfection—is the operative standard. A clock that runs five minutes fast every day is reliable because it produces the same result repeatedly, even though that result is incorrect. The diagnostic transfer
For AI agents, it's straightforward: run the same task through the same workflow multiple times and check whether the failure mode repeats exactly the same way, because repeatability by itself doesn't guarantee correctness.
The PMC study should be treated as a check on measurement design, not as evidence about AI-agent performance. Verify which repeated-trial, exposure, and subgroup analyses are appropriate for the agent being evaluated, record the resulting outcome distributions, and do not transfer the study’s neuroscience findings to an agent without a separate validation.
Check whether the agent’s success rate changes across predefined task or operating-condition slices before relying on an aggregate figure. The available material supports examining reliability by relevant subgroups, but it does not establish that the fMRI study’s attenuation-correction method or its domain-specific subgroup findings should be applied directly to AI-agent task success.
The satellite cost-modeling sources identify a Cost Correction Factor and a Low Cost Small Satellite adjustment factor in their own estimation context. They do not establish a correction formula for AI-agent task success. For an agent, report the observed success rate and its uncertainty directly, and specify any separate weighting or adjustment method before using it in a 90% decision gate.
Accordingly, the available evidence proves only the reliability measurement method: repeated trials, repeatability inspection, exposure-length variation, and population segmentation. It does not prove a ve
rified 90% AI-agent result, and any claim to that figure must come from a separate, held-out, risk-weighted evaluation that applies these correction mechanisms.
Choose retraining or redesign explicitly
When an AI agent does not meet the required success rate, the first step is to verify that the surrounding workflow, tool calls, permissions, and validation gates are functioning correctly. Only after confirming that these system elements are sound should the focus shift to the model itself; otherwise, the problem likely lies in the orchestration layer and calls for a redesign of the agent system.
Retraining or fine‑tuning is justified when the workflow has been validated and the model repeatedly produces semantic, domain‑specific, or instruction‑following errors. In this case the appropriate test is to run the updated model against the same held‑out tasks while keeping the orchestration unchanged and compare the outcomes. The primary risk is that the model may memorize the evaluation set or hide an underlying control defect, giving a false impression of improvement.
Redesigning the orchestration is the proper response when failures involve tool usage, state transitions, permission handling, retry logic, verification steps, or ambiguous handoffs between components. The required test here is to replace or repair the suspected component and re‑run the full agent pipeline on the held‑out set to see whether the error persists. The main risk is introducing new integration issues or unintended side effects that could degrade other aspects of the system.
| Option | Use when | Required test | Main risk |
|---|---|---|---|
| Retrain or fine‑tune | The workflow is valid and the model repeatedly makes semantic, domain, or instruction‑following errors | Compare the changed model against the same held‑out tasks with unchanged orchestration | Memorizing the evaluation set or masking a control defect |
| Redesign orchestration | Failures involve tools, state, permissions, retries, verification, or ambiguous handoffs | Replace or repair the failing component and re‑run the full pipeline on the held‑out set | Introducing new integration issues or unintended side effects |
Consequently, redesign is the default winner unless the failure cause can be isolated to the model’s capability. Retraining should only be pursued after the workflow, tools, permissions, and validation gates have been shown to be correct. This approach aligns with reliability measurement methods described in the AGNIRVA explanation of reliability as consistent performance and the PMC study on test‑retest reliability maps, which provide a basis for assessing whether observed improvements are genuine.

Count success, risk, and rework
As noted by AGNIRVA, reliability means consistent, not necessarily perfect, so a task is counted as successful only when the required output, all side‑effects, policy checks, and any human‑approval conditions are satisfied. Let S be the number of such successful tasks and N the total number of eligible tasks evaluated in the held‑out set. The task‑level success rate is the exact proportion \(\hat{p}=S/N\).
The 90 % figure serves as a methodological threshold, not as available evidence of performance. Therefore the evaluation reports the point estimate \(\hat{p}\) together with a confidence interval that reflects sampling uncertainty. The interval is computed for the binomial proportion \(\hat{p}\) using the evaluation’s design‑based confidence level (e.g., 95 %).
A standard choice is the Wilson score interval, which provides good coverage even for moderate N. For a confidence level 1‑α the interval is
\[
\frac{\hat{p}+\frac{z^{2}}{2N}}{1+\frac{z^{2}}{N}}
\pm
\frac{z}{1+\frac{z^{2}}{N}}
\sqrt{\frac{\hat{p}(1-\hat{p})}{N}+\frac{z^{2}}{4N^{2}}},
\]
where z is the standard‑normal quantile for α/2 (e.g., z≈1.96 for 95 %). This formula yields the lower and upper bounds that accompany the point estimate.
To avoid masking failures inside an aggregate, the same calculation is repeated for pre‑defined high‑risk slices of the task population (e.g., tasks with safety‑critical outputs or stringent policy constraints). For each slice h we compute \(\hat{p}_{h}=S_{h}/N_{h}\) and its Wilson confidence interval \((\text{lower}_{h},\text{upper}_{h})\). These slice‑specific rates are reported alongside the overall estimate so that any degradation in a risky subset is visible.
The resulting point estimate and its confidence interval—both overall and for each high‑risk slice—feed the 90 % decision gate. If the lower bound of the interval for the overall rate or for any high‑risk slice falls below 0.90, the gate indicates that the verified task‑success threshold has not been met; otherwise the estimate satisfies the methodological threshold. This arithmetic alone owns the measurement basis for the gate, leaving downstream judgments about retraining versus redesign to other sections.

Know where the rule breaks
The retrain-first rule is not universal. It breaks cleanly in three edge cases where the failure surface is too narrow or too noisy to justify touching the model. In each case, the deciding question is not whether the agent failed, but whether the failure is reproducible under a controlled, risk-weighted measurement. If it is not, the cost of retraining is paid for a signal that cannot be verified.
| Edge case | The rule breaks when | The rule still wins when |
|---|---|---|
| Rare high-impact tasks | A small average-rate improvement hides catastrophic authorization or safety failures | The rare task has a dedicated gate and its failure is demonstrably semantic |
| Nonstationary inputs | New policies, tools, or data distributions invalidate the old task contract | The environment is held constant and the same capability error repeats |
| Human-in-the-loop workflows | Reviewers compensate for model drift without logging the intervention | The reviewer’s override is removed and the model fails identically |
For rare high-impact tasks, the average success rate is misleading. A model that passes 99 out of 100 routine queries but fails one authorization check cannot be fixed by retraining alone. The failure is not a capability gap in language generation; it is a control gap in the workflow. The rule still wins only when the rare task has a dedicated validation gate and the failure is demonstrably semantic — meaning the model produced the wrong reasoning, not the wrong action.
Nonstationary inputs invalidate the retrain-first assumption. When policies, tools, or data distributions change, the old task contract no longer applies. Retraining on stale data will not fix a model that is answering the wrong question. The rule still wins only when the environment is held constant and the same capability error repeats across multiple runs. If the failure disappears when the input distribution shifts back, the problem was environmental, not model-based.
Human-in-the-loop workflows obscure the true failure mode. When reviewers compensate for model drift without logging the intervention, the measured success rate is inflated. The model appears reliable because humans are silently patching its errors. The rule still wins only when the reviewer’s override is removed and the model fails identically. If the model recovers when the human steps back in, the failure was in the interface, not the model.
AGNIRVA describes reliability as consistency rather than perfection, but consistency alone does not identify an agent’s failure cause. Use the first-causal-break analysis: retrain only when the model cannot produce the required reasoning or language output under a validated workflow; redesign when the first defect is in workflow, state, tools, permissions, or verification, whether or not the failure repeats.
The evidence boundary is clear. The available sources support reliability measurement methods, not a verified 90% AI-agent result. What they prove is that reliability can be measured through repeated, risk-weighted trials. Until that measurement is done, the retrain-first rule should not be invoked. Redesigning the agent system is the more defensible path when the failure cause is not isolated to model capability.

Run the 90% gate on a toy case
This section runs the 90% gate on a toy case, not a reported production result. We evaluate 100 held-out customer-support tasks under a fixed policy, tool catalog, and verifier. The workflow, permissions, and validation gates are assumed correct for this diagnostic; the question is whether the agent clears the success threshold. Reliability is not a claim—it is demonstrated through evidence, and this ledger provides that evidence for the current build (What Reliability Actually Looks Like, linkedin.com).
The toy ledger reports 86 verified successes out of 100 tasks, so the verified task-success rate is 86%, below the 90% gate. The checkpoint figures of 82 after planning, 88 after tool execution, and 86 after final verification do not show cumulative attrition because the later count exceeds the earlier one; their definitions or ordering should be checked before using them to diagnose the pipeline.
The failure ledger breaks down the 14 misses. Eight tasks fail with invalid tool arguments, three lose conversation state, two violate escalation policy, and one gives an unsupported answer. Thirteen of the fourteen failures originate outside model capability: they are workflow, state-transition, or permission errors. Only one failure lies in the model's reasoning or language output.
| Failure Category | Count | Root Cause |
|---|---|---|
| Invalid tool arguments | 8 | System (tool call) |
| Lost conversation state | 3 | System (state transition) |
| Escalation policy violation | 2 | System (permission/policy) |
| Unsupported answer | 1 | Model (reasoning) |
| Total | 14 | 86% success rate |
Because the first failure lies in workflow or control rather than model capability, redesigning the agent system is the defensible path. Retraining the model would not fix invalid tool arguments or lost state. The comparison is explicit: when the failure cause is not isolated to model capability, redesign wins by default.
Finally, note the evidence boundary. The available sources support reliability measurement methods, not a verified 90% AI-agent result. Reliability means consistent, not necessarily perfect, so a clock that is always 5 minutes fast is reliable (What is Reliability?, AGNIRVA). Our goal is a verified rate above 90% on a held-out, risk-weighted set, and this toy case shows the current build is not there yet.
Apply five final decision rules
Reliability is consistency, not perfection; a clock that is always five minutes fast is reliable because it consistently gives the same result, even if wrong (AGNIRVA). This distinction dictates the intervention. Before touching model weights, verify the workflow, tools, permissions, and validation gates are correct. Apply the diagnostic split: if those controls are sound, the agent is the variable. If they are not, the system is the variable.
Rule one: if the first causal break is a tool schema, state, permission, retry, or verifier defect, redesign before retraining. A model cannot repair a broken permission boundary or an incorrect schema. Check the workflow and control path first, then retrain only if the corrected system still exhibits a model-capability failure.
Rule two: if the system passes component tests but repeats the same semantic or instruction-following error on held-out tasks, retrain or fine-tune and rerun the unchanged evaluation on a held-out, risk-weighted set. This isolates capability from environment. Reliability correction techniques show that accounting for measurement bias reveals true performance gaps that raw scores hide (arxiv.org). Keep the evaluation fixed so the delta reflects model improvement, not test drift.
Rule three: if aggregate success exceeds 90% on a held-out, risk-weighted set but any high-risk authorization or escalation gate fails, do not ship. Redesign the control path and report the gated rate separately. A 90% aggregate rate masks the tail risk where a single failed authorization can cause harm. Report the gated rate separately to ensure the high-risk path meets its own threshold before release.
What to do next
| Step | Action | Why it matters |
|---|---|---|
| 1 | Establish a held-out, risk-weighted evaluation for the agent and use the target task success threshold as the redesign gate. | Success must be measured on representative risk cases, not only familiar prompts or an optimistic test set. |
| 2 | Review failures against the canonical decision rule: redesign when the workflow, tool call, state transition, permission, recovery path, or evaluator is defective. | Changing model weights cannot repair a broken operating system or an invalid success test. |
| 3 | Redesign the workflow, tool contracts, permissions, state transitions, recovery behavior, and evaluator before retraining the agent. | System design is the defensible first response when failures come from execution structure rather than reasoning or language capability. |
| 4 | Retrain only if the corrected agent still fails because it cannot produce the required reasoning or language output under a valid workflow. | This isolates model capability from system defects and prevents unnecessary retraining. |
| 5 | Validate the revised agent on held-out, risk-weighted cases and record evidence for consistent task success, including the reliability standard associated with ISO 9001. | Reliability means consistent performance, not perfect performance, and supplier decisions require auditable evidence. |
Frequently Asked Questions
What is the defensible success rate target for a redesigned agent system after diagnosing operational failures?
The guide identifies a success rate above 90% as a defensible target for a redesigned agent system after operational failures have been diagnosed.
Under what condition should you retrain an AI agent instead of redesigning the system?
Retrain only when failures persist after the workflow, tools, permissions, and validation gates have been corrected.
What distinguishes a system failure that requires redesign from a capability failure that requires retraining?
Redesign is the more defensible approach when failures result from problems with the workflow, tool calls, state transitions, permissions, or response processes rather than from insufficient reasoning or language capabilities.
How does the guide define reliability in terms of performance consistency?
Reliability means consistent performance, not perfect performance.
What specific certification does the guide cite as required evidence for reliability evaluation?
Reliability evaluation requires evidence, including certifications such as ISO 9001.
What two types of evidence should you request from a supplier during evaluation?
When evaluating a supplier, look for documented quality certifications, such as ISO 9001, and product-safety evidence, such as TÜV Rheinland certification.
Quick answers
| What success rate target does the guide set for a redesigned agent system? | The guide sets a target of more than 90% task success on held-out, risk-weighted evaluations for a redesigned agent system. |
| Under what conditions should you retrain an AI agent? | Retrain the agent only when failures persist after the workflow, tools, permissions, and validation gates have been corrected. |
| How does the article define reliability? | Reliability is defined as consistent performance, not perfect performance, illustrated by a clock that is always five minutes fast because it consistently maintains the same offset. |
| When is redesigning the agent system preferred over retraining? | Redesign is preferred when failures result from problems with the workflow, tool calls, state transitions, permissions, or response processes rather than from insufficient reasoning or language capabilities. |
| What kind of evidence is required for reliability evaluation? | Reliability evaluation requires evidence, including certifications such as ISO 9001. |
Also worth reading: The one calendar habit an AI agent can fix for you forever: one calendar habit an AI · Prep for one-on-ones in 5 minutes with an AI agent: Prep for one-on-ones in 5 · Why your AI productivity agent needs access to your chat history: Why your AI productivity agent