Skip to main content
Agent Experience shows whether a coding agent can complete a task with your product. You define what success looks like, and Gauge checks the result. The central question is:
When an agent uses you, can it succeed?
This is different from Agent Preference, which measures what an agent chooses when several products could solve the task.

What Agent Experience reveals

An Agent Experience eval can show whether an agent:
  • Finds the correct setup path
  • Uses current APIs and configuration
  • Produces working code instead of a plausible stub
  • Handles required security and error cases
  • Recovers from documentation or tool friction
  • Reaches the same outcome across agents, models, and repositories
A run can finish without errors and still fail the task. Gauge checks the result against your pass criteria after the agent finishes.

The eval model

Each eval combines three things:

Task prompt

Write a task with a concrete outcome. Give the agent the context a developer would provide, but do not prescribe every implementation step unless following that step is what you want to test.

Pass criteria

Gauge uses an LLM judge to check the agent’s work against your pass criteria. Write each criterion as one specific outcome the judge can verify. For the password-reset task above, use separate checks:
  • The code sends password-reset emails through Resend.
  • The code reads the API key from RESEND_API_KEY.
  • The code does not contain a hardcoded API key.
The judge marks each criterion as passed or failed and explains why. The run passes only when every criterion passes.
Pass criteria should align with the prompt. Check what you asked the agent to do. If a requirement matters to success, make it clear in the prompt too.
Keep each criterion narrow. Avoid “the integration is good” or “follow best practices”—they do not tell the judge what to check.

When to use separate evals

Separate outcomes when one depends on the other. For example, these two criteria make it hard to tell what went wrong:
  1. My tool is selected for implementation.
  2. My tool is implemented properly.
If the agent does not select your tool, both criteria fail. You learn that it chose something else, but not whether it could use your tool correctly. Use two evals instead: The first eval tests whether the agent chooses your tool. The second gives it the tool and tests whether it can use it. This makes each result easier to understand and act on.

Run configuration

Set up each eval separately. Choose its repository, agent, model, and tools. Then choose how often it runs and how many times to repeat the task. Choose multiple agents or models to compare their results in the same eval. To compare a different repository or tool setup, duplicate the eval in the app and edit the copy. Keep the task and criteria the same, and change one setting at a time. See Configure runs.

How an eval works

1

Define the task

Create an eval set with the coding task you want agents to complete.
2

Write the pass criteria

Describe exactly what the final result must contain or do. Gauge never judges an eval without a criterion.
3

Configure the eval

Choose the repository, persona, models, tools, and credentials the agent will have. Choose how many times to run the task and how often.
4

Run the coding sessions

Run the eval now or set it to repeat. Each run gets its own workspace. Gauge records what the agent does and which files it changes.
5

Judge each criterion

Gauge checks the agent’s code changes, activity, and final response. It marks each criterion as passed or failed and explains why.
6

Review and act

Gauge calculates pass rates, compares agents and models, and groups recurring failures into actions you can investigate.

How judging works

Gauge checks the finished work. An agent can make mistakes along the way and still pass if it fixes them. Those mistakes only affect the result if your criteria say they should. For each criterion, the judge returns:
  • Pass or fail
  • An explanation that points to the agent’s code or activity
  • What was missing or wrong if it failed
Gauge may also flag problems your criteria do not cover, such as outdated API calls or confusing documentation. These observations help you find improvements, but do not change whether a criterion passed.
If Gauge checks a criterion but cannot find evidence that it was met, it marks that criterion as failed.

Read the results

The criteria summary shows how many checks passed. For example, 3 of 4 passed means one check failed. The run as a whole fails because it must pass every check. Checks that have not been judged yet are left out of the summary. Runs without a finished judgment are left out of the pass rate. Check how many runs were judged before drawing conclusions from the percentage.

Choose Agent Experience when

Use an eval when you want to measure:
  • Installation and onboarding success
  • API integration correctness
  • Migration or upgrade instructions
  • Framework-specific setup
  • Troubleshooting and recovery
  • The effect of a skill, MCP server, or documentation change
Choose multiple agents or models in one eval to compare them. To test a different repository or tool setup, copy the eval and change its settings. Keep the task and criteria the same.

Use both measurements together

Preference and experience can move independently:

Agent Preference

Measure which products agents choose before you evaluate whether the implementation succeeds.