Original research framework / version 1.1

Test the failure
before the launch.

Most automation failures are not mysterious. They happen at the boundaries: bad inputs, unclear ownership, stale context, permission mismatch, silent retries, or a decision that should have stayed with a person.

Published August 19, 2026Updated August 20, 2026Author Enzo ParilliStatus Framework / no proprietary statistics
Methodology

What was measured—and what was not.

This is an original evaluation framework, not a population study. It organizes failure conditions that an implementation team can test against representative examples. No frequency, benchmark, adoption, revenue, safety rate, client result, or ranking outcome is asserted.

For each workflow, record the trigger, source of truth, identity, permissions, model or rule, tool call, state change, human approval, retry behavior, notification, and recovery path. Test ordinary cases, missing fields, duplicates, stale records, provider errors, revoked access, ambiguous outputs, conflicting sources, and irreversible actions.

Reference context: NIST AI Risk Management Framework, OWASP LLM application risk guidance, OpenAI evaluation guidance, and Cloudflare Workers documentation. These are reference points, not proof that a particular system complies with a framework.

1. Input and identity failures

Test malformed payloads, missing required fields, duplicate events, ambiguous contacts, stale identifiers, conflicting records, and input that looks valid but belongs to the wrong entity. The acceptance test is not “the model responded”; it is “the system rejected or routed the wrong input without corrupting the source of truth.”

2. Context and retrieval failures

Test stale documents, incomplete retrieval, contradictory sources, permission leakage, missing citations, and answers that sound certain when evidence is weak. A useful system makes source and uncertainty visible and routes low-confidence cases to review.

3. Tool and state-change failures

Test timeouts, rate limits, partial success, non-idempotent retries, provider schema changes, unauthorized calls, and a successful API response that does not prove the business outcome occurred. Log intent, request, response, state, owner, and reconciliation result.

4. Human control and recovery failures

Test moments where judgment, consent, legal responsibility, financial commitment, or irreversible action belongs to an accountable person. The system should show a clear approval boundary, preserve evidence needed to decide, expose the next action, and make recovery possible.

Failure taxonomy

The taxonomy turns broad concerns into observable test cases. A workflow can have more than one failure class; record the primary owner and the recovery owner separately.

Failure classObservable signalConsequenceMitigation / test
Input / identityMissing, duplicated, stale, or mis-associated recordWrong routing or corrupted source of truthSchema validation, deduplication, identity checks, rejected-fixture tests
Context / retrievalUnsupported answer, stale source, or permission leakageUnreliable decision or inappropriate disclosureSource display, freshness rules, access tests, confidence threshold, human review
Tool / authorizationUnexpected call, expired token, or provider rejectionPartial action, data exposure, or blocked workflowLeast privilege, explicit tool allowlist, revocation test, audit event
State / retryDuplicate write, timeout after commit, or unreconciled responseDouble action or false completionIdempotency key, durable state, retry budget, reconciliation fixture
Human / recoveryNo owner, unclear approval, or no pause / undo pathUnaccountable or irreversible outcomeApproval boundary, escalation SLA, rollback path, operator drill

Release gate

  1. Every failure mode has an owner and a response.
  2. Every consequential action has a defined human approval rule.
  3. Every external system has a permission and revocation path.
  4. Representative and adversarial examples pass before release.
  5. Monitoring can distinguish a completed business outcome from a successful tool call.

Limits and next use

This framework is a practical test design aid. It does not certify safety, legal compliance, model quality, or business value. Record the examples, expected behavior, actual behavior, owner, evidence, and decision for each release.

Use it with the AI automation opportunity calculator, the law-firm intake calculator, or the AI automation service. For the public evidence boundary, see the proof ledger.