What Gauge measures
Agent Preference
Measure which products an agent recommends, mentions, installs, or uses for an open-ended task.
Agent Experience
Define a task and explicit criteria. Gauge judges completed runs against those criteria and records verdicts and evidence.
Set up what to test
An eval contains a task, pass criteria, and the settings for running it. Choose the repository, models, and tools on the eval itself. To test a different repository or tool setup, copy the eval and change those settings. A preference prompt asks the same question in one or more scenarios. Each scenario combines a repository with a persona—the kind of developer you want to simulate. Choose models and tools once on the prompt; every scenario uses them. Each eval or preference prompt has its own schedule. Run it daily, weekly, monthly, or only when you choose Run now. See Configure runs for examples and run-count calculations.From measurement to result
1
Define the measurement
Write a preference question or an eval task. For an eval, add observable pass criteria that describe the outcome you want to test.
2
Configure the runs
Choose agents, models, tools, credentials, and samples. Set the eval’s repository and persona, or add repository/persona scenarios to a preference prompt.
3
Launch
Choose Run now or set a schedule. Gauge runs each selected model as many times as you requested. For preference prompts, it does this for each scenario. Running now does not change the schedule.
4
Run the agent
Gauge prepares an isolated workspace and runs the coding agent. Limits on time, turns, and spend keep sessions from running indefinitely.
5
Inspect evidence
Open the session trace, final response, working-tree diff, timing, and exit reason. A research task may produce no code changes.
6
Judge and improve
Gauge analyzes preference or judges eval criteria. Use the evidence to identify a specific change, test it, and measure again.