Enforcing named tools in agent runs

How Lightcap records real attempts, distinguishes search from scraping, and blocks install-only work from passing as a completed operation.

By Faruk Alpay · Published · Updated · 6 min read

  • agent reliability
  • tool use
  • verification

The request was unambiguous: install Crawl4AI, then use it to scrape a page. The agent answered that the package could not be installed and switched to web search.

The execution log contained no installation command, import check, or error. The fallback happened before the requested path had been tried.

This is a method-fidelity failure. The search may return useful links, but it does not establish what the target page produced when processed by Crawl4AI. The operation changed from scraping a named page with a named tool to finding pages through a search engine.

Treat the method as part of the deliverable

Many tasks leave implementation details open. "Find current reporting about this event" permits several acquisition methods. "Scrape this URL with Crawl4AI" does not. In the second request, the tool and operation are explicit acceptance criteria.

Lightcap represents that requirement as a small contract:

{
  "request": {
    "primary_tool": "crawl4ai",
    "operation": "scrape",
    "requires_run": true,
    "equivalent_methods": []
  },
  "attempts": [],
  "fallback": null
}

The contract separates three questions:

  1. Was the requested tool set up?
  2. Was it actually run against the requested input?
  3. If another method produced the result, was that substitution equivalent and disclosed?

Conflating these questions caused both failures we observed: switching tools without an attempt, and installing a tool without ever using it.

A failure claim needs an execution trace

The sentence "Crawl4AI cannot be installed here" is a factual claim about the environment. A model's expectation is not enough to support it.

A grounded attempt looks like this:

python -m venv .venv
.venv/bin/pip install crawl4ai
.venv/bin/python -c "import crawl4ai; print('ok')"

The record must contain the command plus a non-zero exit code, stderr, or a typed error before the runtime accepts an installation-failure claim. A row containing only { "tool": "crawl4ai" } is not an attempt.

The distinction is enforced in two predicates:

const ranSomething = (attempt) =>
  typeof attempt.command === 'string' && attempt.command.trim().length > 0;

function failedWithEvidence(attempt) {
  if (!ranSomething(attempt)) return false;
  if (Number(attempt.exitCode) !== 0) return true;
  if (attempt.ok === false) return true;
  return Boolean(attempt.stderr || attempt.error);
}

This check is intentionally mechanical. It does not ask a second model whether the excuse sounds plausible.

The fallback order is explicit

A legitimate fallback has a stable order:

run the requested path
  -> capture the failure
  -> choose an alternative
  -> label the alternative and its limits

For example, a web search after a failed scrape may help locate other reporting. The final answer must still say that the delivered evidence came from search results rather than a scrape of the target page.

The gate distinguishes equivalent methods from non-equivalent ones. If a contract explicitly allows two page renderers as interchangeable, moving from one to the other is not drift. Search and scrape remain different method classes because they answer different questions.

Five violations cover the observed cases

server/src/web-acquisition/tool-fidelity.js returns typed violations from the contract and trace.

unattempted_primary

The run used a fallback or declared the primary path unavailable without recording a real command for the named tool.

ungrounded_failure_claim

The final answer says the named tool could not be installed or used, but the trace has no failed attempt with an exit code or error.

search_presented_as_scrape

The requested operation was a scrape, the result came from a search tool, and the answer did not identify it as search-based.

unflagged_substitution

A non-equivalent alternative ran, but the answer presented it without explaining that the primary method changed.

installed_but_not_run

The named tool was installed successfully, the contract required an operation, and no non-install command used the tool. This catches the second incident, where setup finished and the agent stopped at "ready to go" without scraping the page.

Installation commands are classified from an explicit phase when available, then from known setup forms such as pip install, git clone, playwright install, and crawl4ai-setup. An explicit phase: "run" takes precedence over that inference.

Provision known prerequisites in the setup pass

Named tools often carry known runtime requirements. Crawl4AI drives a browser, so a complete setup normally includes its browser engine and system dependencies. Installing only the Python package and waiting for the first crawl to reveal that Chromium is absent adds a predictable failure round.

That does not justify guessing every dependency. The boundary is practical:

  • provision requirements documented as part of the normal setup path;
  • discover platform-specific or transitive failures by running the tool and reading the real error;
  • never report an environment limitation before the command has produced evidence for it.

Dependency foresight is execution policy, not something the current final-answer gate can prove completely. The gate can verify that installation was followed by a run. It cannot know whether an undocumented system library should have been anticipated.

The audit and the behavioral probe

The audit is a pure function with no network access and no model call. With no primary_tool, it does nothing. When a contract names one, the function compares the request, attempts, fallback, and final answer.

Unit tests cover missing attempts, grounded failures, equivalent methods, search-for-scrape substitution, labeled alternatives, successful primary runs, and install-only traces. The Stop-hook wrapper in scripts/tool-fidelity-hook.mjs runs the same function against a durable JSON record.

The behavioral probe in scripts/measure/probes/tool-substitution.json checks the property one level earlier. It gives an agent both a requested tool and an easy search tool, then verifies that installation precedes use and that the named operation actually appears in the tool trace.

What success means

For a method-specific request, useful output is only one part of acceptance. The trace must show that the named path ran, and the output must come from that path. If the path fails, the failure should be concrete enough to reproduce. If a substitute is used, its method and loss of fidelity should be visible.

That standard costs one real attempt. It prevents a much more expensive outcome: a polished answer that reports a search as a scrape, or setup as completed work.

All dev blogs · RSS feed

Lightcap