When an agent uses you, can it succeed?This is different from Agent Preference, which measures what an agent chooses when several products could solve the task.
What Agent Experience reveals
An Agent Experience eval can show whether an agent:- Finds the correct setup path
- Uses current APIs and configuration
- Produces working code instead of a plausible stub
- Handles required security and error cases
- Recovers from documentation or tool friction
- Reaches the same outcome across agents, models, and repositories
The eval model
Each eval combines three things:Task prompt
Write a task with a concrete outcome. Give the agent the context a developer would provide, but do not prescribe every implementation step unless following that step is what you want to test.Pass criteria
Gauge uses an LLM judge to check the agent’s work against your pass criteria. Write each criterion as one specific outcome the judge can verify. For the password-reset task above, use separate checks:- The code sends password-reset emails through Resend.
- The code reads the API key from
RESEND_API_KEY. - The code does not contain a hardcoded API key.
When to use separate evals
Separate outcomes when one depends on the other. For example, these two criteria make it hard to tell what went wrong:- My tool is selected for implementation.
- My tool is implemented properly.
The first eval tests whether the agent chooses your tool. The second gives it the tool and tests whether it can use it. This makes each result easier to understand and act on.
Run configuration
Set up each eval separately. Choose its repository, agent, model, and tools. Then choose how often it runs and how many times to repeat the task. Choose multiple agents or models to compare their results in the same eval. To compare a different repository or tool setup, duplicate the eval in the app and edit the copy. Keep the task and criteria the same, and change one setting at a time. See Configure runs.How an eval works
1
Define the task
Create an eval set with the coding task you want agents to complete.
2
Write the pass criteria
Describe exactly what the final result must contain or do. Gauge never judges an eval without a criterion.
3
Configure the eval
Choose the repository, persona, models, tools, and credentials the agent will have. Choose how many times to run the task and how often.
4
Run the coding sessions
Run the eval now or set it to repeat. Each run gets its own workspace. Gauge records what the agent does and which files it changes.
5
Judge each criterion
Gauge checks the agent’s code changes, activity, and final response. It marks each criterion as passed or failed and explains why.
6
Review and act
Gauge calculates pass rates, compares agents and models, and groups recurring failures into actions you can investigate.
How judging works
Gauge checks the finished work. An agent can make mistakes along the way and still pass if it fixes them. Those mistakes only affect the result if your criteria say they should. For each criterion, the judge returns:- Pass or fail
- An explanation that points to the agent’s code or activity
- What was missing or wrong if it failed
If Gauge checks a criterion but cannot find evidence that it was met, it marks that criterion as failed.
Read the results
The criteria summary shows how many checks passed. For example, 3 of 4 passed means one check failed. The run as a whole fails because it must pass every check.
Checks that have not been judged yet are left out of the summary. Runs without a finished judgment are left out of the pass rate. Check how many runs were judged before drawing conclusions from the percentage.
Choose Agent Experience when
Use an eval when you want to measure:- Installation and onboarding success
- API integration correctness
- Migration or upgrade instructions
- Framework-specific setup
- Troubleshooting and recovery
- The effect of a skill, MCP server, or documentation change
Use both measurements together
Preference and experience can move independently:Agent Preference
Measure which products agents choose before you evaluate whether the implementation succeeds.