Why Lightcap replaced nine personas with two compute tiers

How the product moved from character selection to Small and Medium, with two historical one-run measurements and strict limits on what they prove.

By Faruk Alpay · Published · Updated · 6 min read

  • product architecture
  • model routing
  • benchmarks

Lightcap once asked users to choose among nine named personas. The names suggested specialties such as frontend work, architecture, algorithms, and long-horizon builds. They also made the first product decision harder than it needed to be: users had to guess which personality mapped to the compute their task required.

The current interface exposes two choices instead:

  • Lightcap Small is the fast default for everyday builds, code, research, and answers.
  • Lightcap Medium provides a deeper reasoning tier for harder algorithms, larger builds, and repair work that needs more compute.

The server keeps provider and engine identifiers private. The client receives product-level capability, not a vendor catalog.

This was an interface and routing change

The nine-persona roster mixed three different concepts:

  1. presentation, through a character name and visual identity;
  2. task specialization, through labels such as frontend or architecture;
  3. compute policy, through the model and reasoning budget behind the character.

Those layers drifted. A character description could remain stable while the underlying route changed. Two characters could eventually use similar engines. Users saw a wider choice even when the runtime difference was narrow.

Small and Medium make the compute decision explicit. A user can select the faster or deeper tier without learning an internal cast. The orchestration layer can also assign bounded pieces of a larger job to the appropriate tier.

This does not prove that two is the ideal number of choices. We did not run a controlled usability study of the old and new selectors. Reduced decision friction is a product hypothesis supported by the simpler information architecture, not a measured conversion claim.

What happened to old links and runs

The public roster now contains only small and medium. Lightcap Small is the default. Legacy slugs and unknown character slugs resolve to that default so stored runs and old client links do not fail with a missing agent.

That compatibility behavior is intentionally one-way. Old identifiers remain readable, but the current interface does not recreate the retired character menu.

The migration also separates product naming from infrastructure. A provider outage or engine change can be handled in the server-side route without renaming the user-facing tier.

Two historical measurements, kept in context

The original article combined several benchmark numbers and presented them as evidence for the new pair. They came from different workloads and legacy engines. They should not be read as a Small-versus-Medium comparison.

Reasoning-policy sample

A June 2026 script ran the same 25-horses puzzle under two reasoning policies.

MetricVerbose baselineBudgeted reasoning
Latency91.6 s30.1 s
Reasoning tokens3,027478
Recorded cost$0.004877$0.001297
OutcomeNo final answer within budgetCorrect answer: 7

This was one run. The baseline exhausted its output budget, so the percentage differences mix efficiency with a failed terminal outcome. The result motivated tighter reasoning budgets; it is not a general performance estimate.

Tool-loop sample

A separate Rust build drove a legacy fast engine through 29 model steps and 27 tool calls. It took about 228 seconds, consumed 4,371,865 input tokens, and cost about $0.86. Four tests passed, but the benchmark printed 0 ms for every timing.

That trace showed why raw decoding speed does not determine agent-loop speed. Repeated context input, reasoning before tool calls, sequential round trips, and external build commands all contribute to wall time.

Again, this is not a comparison between the current tiers. It is a historical workload trace that influenced the routing design.

The current routing contract

Small and Medium have different roles in the runtime:

  • Small starts the normal low-latency path.
  • Medium is available as the deeper tier.
  • Crew execution can keep ordinary pieces on the base tier and escalate failed or structurally insufficient pieces.
  • Engine identifiers and provider fallbacks stay on the server.

This contract is easier to test than persona prose. Tests can assert which public slugs exist, what a legacy slug resolves to, whether a free or paid route admits a provider, and whether escalation preserves the requested tier.

The safety boundary is independent of the tier labels. Changing the reasoning budget or provider order must not expand file, command, network, or credential authority.

What the simplification improves

The smaller interface has concrete engineering benefits even without a conversion experiment.

First, product copy and server behavior share a stable vocabulary. "Small" and "Medium" describe relative compute rather than a fictional job title.

Second, telemetry can be grouped by a durable tier. A character rename no longer fragments latency, cost, or completion measurements.

Third, fallbacks can change behind the tier boundary. Users choose the service level while the server handles funded and emergency routes.

Fourth, the orchestration layer can assign work by dependency and verification need instead of pretending each batch requires a distinct personality.

What still needs measurement

The next product evaluation should compare the old and new selection experiences directly. Useful measures include time to first prompt, selector abandonment, manual tier changes, task completion, user retries, and cost per verified result.

Model evaluation should remain workload-specific. Small and Medium need a shared suite with identical acceptance criteria for research, coding, multi-file changes, and recovery after tool failure. Report pass rates and uncertainty before throughput or token savings.

Nine personas made the system feel broad, but breadth in a selector is not the same as capability in a run. Two compute tiers give the runtime a clearer contract. The benchmarks above explain parts of the design history; they do not certify the current product by themselves.

All dev blogs · RSS feed

Lightcap