Skip to content

Factories > Management & observability

Benchmarking factory agent configurations

Open in ChatGPT ↗
Ask ChatGPT about this page
Open in Claude ↗
Ask Claude about this page
Copied!

Benchmarks compare a factory agent's model and runner configurations on fixed tasks. Use the results to choose a production configuration.

Benchmarks compare model and runner configurations for one factory agent on the same fixed tasks. Use a benchmark to test a change on representative work before you apply it to your factory. This video shows how to compare benchmark results for quality and cost before changing a coding agent’s configuration.

A benchmark helps you choose an agent configuration based on repeatable evidence. Run representative tasks with different configurations, then compare their results before you change a live factory.

A benchmark suite is a reusable collection of tasks that evaluates one factory agent. Each task has a prompt and “Correctness criteria,” which tell the built-in Correctness Scorer what a successful trial must do. A trial is a single run of a task under one configuration. Repetitions create additional trials.

When you launch a suite, choose the configurations, Scorers, and repetitions to compare. Warp runs the trials, then scores the completed ones. The run keeps those inputs, so later changes to the suite do not change its past results.

The Runs tab for a benchmark suite, with task, configuration, run time, and cost columns.

The Runs tab for a benchmark suite.

Use separate suites for focused questions:

  • Factory default - Compare models and runners to choose the default configuration for a coding agent.
  • Frontend changes - Compare configurations on representative frontend tasks before applying one to that workflow.

To use Benchmarks, you need a factory with an agent to evaluate. Use a completed run from that agent when you want a task to reproduce real work, or write a task yourself.

This video shows how to turn your team’s coding tasks into a reusable benchmark suite.

  1. In the Warp Factories web app, open your factory, click Benchmarks, then click New.

    The Benchmarks page showing saved benchmark suites, a search field, and the New button.

    The Benchmarks page for a factory.

  2. Enter a name and optional description, then choose the agent to evaluate. The suite runs every task as that agent.

    The benchmark editor showing a documentation-link review suite, a selected triage agent, and the Add task action.

    The benchmark editor with a selected agent.

  3. Click Add task. Warp saves the benchmark, then opens task setup.

  4. Select a completed run, then click Add task. To write a task instead, click Start from scratch instead.

    The Add task pane with a run search field, the list of completed runs, and the Start from scratch instead option.

    The task source picker.

  5. Review the task prompt and enter “Correctness criteria” for the task.

    The task editor with fields for a title, prompt, correctness criteria, and pinned repositories.

    The task editor for a benchmark suite.

  6. Click Run. In the launch dialog, choose the model and runner for each configuration. Optionally mark one configuration as the baseline. Third-party harness comparisons are not available yet.

  7. Add configurations, select Scorers, and set “Repetitions.” The dialog shows the number of trials created. More trials and Scorers increase the run’s cost.

    The New benchmark run dialog with model and runner configurations, a baseline option, selected Scorers, a repetition count, and projected trial count.

    The launch configuration for a benchmark run.

  8. Click Run benchmark. The benchmark page shows its status and scored trials. You can cancel a running or scoring benchmark.

After the run completes, review the result as a comparison, not as a universal model ranking:

A completed benchmark result showing the overall recommendation, additional recommendations, total cost, and a correctness versus average cost chart.

A completed benchmark result and comparison chart.

  • Overall recommendation - Identifies the highest-quality configuration when at least two configurations have comparable results.
  • Additional recommendations - Highlight the most efficient and lowest-cost configurations when the result supports those comparisons.
  • Comparison chart - Compare the selected result dimensions across configurations.
  • Overall table - Compare each Scorer’s average and the combined Overall value. Expand a configuration, task, and repetition to inspect its individual trials.
  • Scorer grids - Show each task’s results across configurations for a selected Scorer.

Expanded benchmark results showing trial run duration, cost, and scores for a configuration.

Expanded trial results for a configuration.

The run’s “Total cost” includes model usage for trials and Scorer usage. It estimates those costs from credits at your team’s current rate, so it is not a billed amount. A failed or cancelled benchmark shows only results that finished scoring before the run stopped.

Change one configuration at a time. If the evidence supports a candidate, update the agent’s model or runner in the factory dashboard, or submit the change through your factory definition. Keep the relevant Scorers active, then compare later production runs with the baseline you recorded before the change.

For version-controlled factories, define reusable suites in benchmarks/<suite-slug>/suite.yaml and their tasks in benchmarks/<suite-slug>/tasks/<task-slug>.yaml. See benchmark suite files.