Skip to main content
Gauge Agents recreates realistic developer tasks, runs coding agents in isolated workspaces, and records the evidence behind their results. You can compare models and tool configurations, then test changes to the documentation or skills agents use.

What Gauge measures

Agent Preference

Measure which products an agent recommends, mentions, installs, or uses for an open-ended task.

Agent Experience

Define a task and explicit criteria. Gauge judges completed runs against those criteria and records verdicts and evidence.

Set up what to test

An eval contains a task, pass criteria, and the settings for running it. Choose the repository, models, and tools on the eval itself. To test a different repository or tool setup, copy the eval and change those settings. A preference prompt asks the same question in one or more scenarios. Each scenario combines a repository with a persona—the kind of developer you want to simulate. Choose models and tools once on the prompt; every scenario uses them. Each eval or preference prompt has its own schedule. Run it daily, weekly, monthly, or only when you choose Run now. See Configure runs for examples and run-count calculations.

From measurement to result

1

Define the measurement

Write a preference question or an eval task. For an eval, add observable pass criteria that describe the outcome you want to test.
2

Configure the runs

Choose agents, models, tools, credentials, and samples. Set the eval’s repository and persona, or add repository/persona scenarios to a preference prompt.
3

Launch

Choose Run now or set a schedule. Gauge runs each selected model as many times as you requested. For preference prompts, it does this for each scenario. Running now does not change the schedule.
4

Run the agent

Gauge prepares an isolated workspace and runs the coding agent. Limits on time, turns, and spend keep sessions from running indefinitely.
5

Inspect evidence

Open the session trace, final response, working-tree diff, timing, and exit reason. A research task may produce no code changes.
6

Judge and improve

Gauge analyzes preference or judges eval criteria. Use the evidence to identify a specific change, test it, and measure again.
A session can finish without errors and still fail the task. Gauge checks its work against your criteria. Runs without a finished judgment are left out of the pass rate.

Test improvements

Click New optimization on an eval or preference prompt. You can test changes to a documentation page or skill and see whether agents do better. Review the results and the proposed change before publishing it. See Improve agent results for the steps. Legacy content experiments are retired. Older experiments stay readable in the web app.

Where execution and credentials live

The web app, CLI, and API configure measurements and launch work. Coding-agent sessions execute in Gauge’s isolated environments, not on the machine running the CLI. Service credentials are available through the selected run configuration or an authorized MCP server. Agents use placeholders; Gauge releases the secret only to approved service hosts. See Credentials for setup and access controls.