Blocking false completion in agent runs

How a traced crawl failure led to deterministic checks for false completion, goal substitution, and finished-looking residuals.

By Faruk Alpay · Published · Updated · 6 min read

  • agent reliability
  • admissibility
  • verification

An agent can execute several valid setup steps and still fail the task. We hit that distinction while testing a request with one explicit deliverable: install Crawl4AI, run its asynchronous browser crawler against https://example.com, and return the generated Markdown.

The run cloned code, installed Python packages, downloaded Chromium, and attempted to provision browser libraries. It then hit two recorded failures. The system package command failed, and AsyncWebCrawler could not be imported. Instead of reporting the blocked browser crawl, the run converted raw HTML to Markdown and marked the task complete.

That last artifact was non-empty. It was also the wrong artifact.

The failure was in the completion decision

The requested goal had three relevant properties:

  • the page had to be processed by Crawl4AI;
  • the asynchronous browser path had to run;
  • that run had to produce non-empty Markdown.

Raw-HTML conversion satisfied only the third property. Treating it as success replaced the goal after the requested path failed.

We now describe completion as a certificate over a trace. A certificate is allowed only when the trace contains a successful attempt that produced the declared deliverable for the declared goal. Setup activity, exit code zero, or a nearby artifact is insufficient on its own.

In operational terms:

declare the goal
  -> record attempts and outcomes
  -> verify a non-empty artifact for that goal
  -> issue one completion status

If the third step does not hold, the terminal status can still be partial or blocked. It cannot be completed.

The record the gate reads

The final-answer check does not infer success from fluent prose. It reads a compact typed record:

{
  "goal": {
    "deliverable": "Markdown from an AsyncWebCrawler browser run",
    "operation": "crawl"
  },
  "attempts": [
    {
      "command": "python run_crawl.py",
      "exitCode": 1,
      "error": "ImportError: AsyncWebCrawler",
      "for_goal": "Markdown from an AsyncWebCrawler browser run"
    },
    {
      "command": "python raw_html_to_markdown.py",
      "exitCode": 0,
      "output": "<non-empty markdown>",
      "produced": true,
      "for_goal": "Markdown from raw HTML"
    }
  ],
  "status": "completed"
}

The second attempt produced something real, so this is not an empty-output case. Its for_goal value differs from the declared deliverable. That difference is the evidence that the result belongs to a substitute goal.

An attempt establishes completion only when all of these are true:

  1. A command or operation actually ran.
  2. The attempt succeeded without a captured error.
  3. It produced non-empty output or an explicitly recorded artifact.
  4. It served the declared goal.

This is stricter than checking an exit code. A command may exit successfully while writing an empty file, testing the wrong target, or producing a format the user did not request.

Three enforceable violations

server/src/web-acquisition/admissibility.js implements three typed violations.

1. certificate_without_admissibility

This fires when the answer or structured status claims completion, but no recorded attempt produced the declared deliverable. A successful setup step does not satisfy it. Neither does an output row with no content.

2. goal_substitution_uncertified

This fires when an attempt produced an artifact for a different goal and the answer certified it as the original result. A fallback remains possible, but the answer must label it as a substitute and explain what changed.

For the Crawl4AI incident, an honest substitute would say that the output came from raw HTML and therefore did not verify browser rendering or the requested crawler path.

3. residual_certificate

The failed run ended with a cluster of completion markers even though the requested artifact did not exist. The gate treats four or more non-negated markers as a completion spray when no declared deliverable was produced. This rule is deliberately narrow. It targets the recorded failure shape rather than trying to judge writing style in general.

The audit also recognizes explicit incompleteness. A sentence such as "the browser crawl did not complete" does not become a false certificate merely because it appears in a final response. A structured success status remains a hard completion claim, so prose cannot use a disclaimer to override it.

The correct terminal response

The same trace supported a useful answer:

The Crawl4AI browser run did not complete. Package setup and the Chromium download succeeded, but system-library provisioning and the AsyncWebCrawler import failed. I produced Markdown through a raw-HTML converter as a separate fallback; it does not verify the requested browser crawl.

That response preserves the evidence and the boundary of the result. It gives the next engineer enough information to continue without rerunning every successful setup step.

Where the check runs

The gate is deterministic and does not call a model. The core function accepts the goal, attempts, terminal status, and answer text, then returns an eligibility verdict plus typed violations. Unit tests cover successful production, empty output, labeled substitution, honest incomplete reports, structured success statuses, and completion sprays.

The same audit is available through scripts/admissibility-hook.mjs. A run opts in by writing its typed record before the stop hook executes. With no declared goal, the audit is a no-op. That default matters because ordinary prose tasks do not always have a file-like deliverable that can be certified this way.

An end-to-end behavioral probe in scripts/measure/probes/admissibility.json asks whether an agent reports a blocked named path honestly or replaces it with a finished-looking substitute.

What this gate cannot prove

A final-answer audit sees the recorded attempts and the claimed outcome. It cannot, by itself, prove that every intermediate decision was sensible.

Two process failures need separate run controls:

  • choosing a cheap-looking step before checking whether it can still reach the goal;
  • repeating a command that already failed with the same captured outcome.

Those behaviors are visible only if attempts are durably recorded as they happen. The final gate can then reject a false completion, while loop detection and execution policy handle the path that led there.

There is another limit. The for_goal field is only trustworthy if the runtime derives and persists it consistently. If a model can relabel a substitute artifact as the requested goal, the type exists but the evidence is weak. The production rule therefore has to join typed records with the actual tool and artifact ledger, not accept a self-reported label in isolation.

The practical standard is simple: completion belongs to the deliverable, not to the amount of activity in the trace. When the artifact is missing or belongs to another goal, the run should end with a precise gap instead of a success certificate.

All dev blogs · RSS feed

Lightcap