Skip to main content
Gauge Agents tests whether coding agents can complete a task with your latest CLI, package, docs, or Skills before you merge. The GitHub App reads the configuration and eval cases from your PR, runs coding sessions with the candidate inputs, and reports the results in a Gauge evaluations check and PR comment.
Start with your coding agent. We recommend the gauge-agents Skill and Gauge CLI. The Skill’s dedicated PR setup instructions help your agent inspect the repository, reuse existing CI or previews, and verify the configuration. Follow the copy-paste setup below.

Before you begin.

Use Gauge CLI 0.15.0 or later for local setup verification. It requires Node.js 22.12 or later; install it with npm install -g @withgauge/cli, or update an existing installation with gauge update. Automatic PR evaluations launch real coding sessions. Start with one case, one agent, and one sample while getting set up. Installing the App or verifying configuration alone does not launch sessions. The examples below use placeholder names and paths. Replace them with yours. Screenshots use example data.

Set up with your coding agent

With Node.js 22.12 or later installed, run this in your repository to install the CLI and Skill, sign in, and list your Gauge organizations:
Then start a new coding-agent session in the repository so it can discover the installed Skill, and paste:
The agent can prepare the repository files; complete the GitHub App installation and repository settings below before opening the PR. The install, sign-in, and verification commands do not launch coding sessions.

Set up your first pull request eval

1

Install the Gauge GitHub App

Open Settings → GitHub and select Install GitHub App.
Gauge Agents GitHub settings showing Connect your repositories and the Install GitHub App button.

Connect GitHub from your organization's settings.

Screenshot text: The disconnected GitHub settings page shows Connect your repositories and an Install GitHub App button. Select that button to begin installation.On GitHub, choose your account and select the repositories whose CLI, docs, or Skills you want to test. Return to Gauge to see the connected account and its repositories.If GitHub requests an account owner’s approval, finish that approval before continuing. See repository access and troubleshooting.
2

Configure pull request checks for each repository

On the same Settings → GitHub page, find Pull request checks. Enable checks for each repository you want to test.Use Edit beside Target branches to choose which branches should receive checks. For example, select main to test PRs that merge into main. These are the PR’s destination branches.Repositories without custom settings default to checks enabled for main. Branch names are exact names, such as main or release/next, rather than glob patterns. Changing these settings does not itself launch an evaluation; use a new PR update or rerun an existing check after setup.
Connected GitHub settings with PR checks enabled for acme/cli, acme/docs, and acme/skills, and target branches configured for each.

Enable checks and choose target branches independently for each repository.

Screenshot text: The Connected panel has an Add or remove repositories button. Under Pull request checks, each repository has an Enabled toggle and an Edit action beside Target branches. In this example, acme/cli and acme/skills target main; acme/docs targets both main and staging.
3

Create gauge.json

Create gauge.json at the repository root. Copy an input configuration below and set org to your Gauge organization slug from the app URL. That organization must own the App installation connected to this repository.Use the configuration schema for accepted fields, input types, and enum values. See schemas and enums for validation guidance.This file tells Gauge what to make available to the agent. Packages and native executables come from your existing CI, docs pages come from a public preview, and Skills or files can come directly from the PR’s commit. Gauge consumes these inputs; it does not trigger builds. Custom prepare/install scripts are not supported in gauge.json.gauge-evals/**/*.md is the default case location. The examples declare it explicitly. No github.enabled field is needed; repository access and target branches are controlled in Settings → GitHub.
4

Write your first eval case

Create gauge-evals/onboarding.md. Start with this template and replace the task and success criterion before opening the PR.
gauge-evals/onboarding.md
The Markdown body is the agent’s task. The criterion tells Gauge how to judge the result and is not shown to the agent. For example, an onboarding task might ask the agent to set up your product and run its first command; the criterion could require a successful response from that command.The eval-case schema describes the YAML frontmatter, including the exact agent enum values. Use CLAUDE_CODE, CODEX_CLI, PI, or OPENCODE; lowercase CLI aliases such as claude-code are not valid here.By default, the case uses every input in gauge.json. To choose specific inputs, add inputs: [cli], inputs: [docs], or inputs: [product-skill] under config, using the names from your configuration.If the task needs an existing application, add repoUrl and repoRef under config, such as https://github.com/your-org/example-app and main. Otherwise, the agent starts in an empty workspace. Each case owns its agent and sample settings. Start with one agent and one sample.The consumer application is separate from the repository supplying the PR’s inputs. Use a pinned commit in repoRef when you need a repeatable starting point. Omitting models uses the harness default; run gauge models list -o json after signing in to choose and pin available model IDs under each agent’s models array.
5

Verify the setup locally

From the repository, check your working files before committing:
Verification works offline without login or a PR. It inspects configuration, case input references, Git input paths in the working tree, and recognizable workflow declarations. yes means a local check passed, no means a definite problem, and unknown means that declaration could not be resolved statically. Exit code 1 indicates a definite problem; exit code 0 can still include unknowns.Fix the no results and inspect any unknown results. Verification does not execute CI, inspect built artifacts, prove installation, check App activation, or start sessions. A repository with no cases is reported as information, so also confirm that it found the case you just wrote.After committing, optionally sign in and preview the committed setup:
Plan shows the session count and input readiness without launching sessions. Unlike verify, it reads committed HEAD and ignores uncommitted edits.
6

Open a pull request and review the result

Commit gauge.json, your eval case, and the changes you want to test. Push a branch in the connected repository and open a PR targeting one of its configured branches.Gauge adds a Gauge evaluations check to the latest commit. It reads the committed evals, waits for any required package or docs preview, then starts the sessions. You can include the configuration in your first PR; it does not need to be merged first. Pull requests from forks do not launch evals.The PR comment updates through input preparation, execution, and judging. When judging finishes, the check shows whether the evals passed. Open its session links to review messages, commands, code changes, and criterion explanations. Use that evidence to improve your product, then push a new commit to test again.

Configuration schemas and enums

Use the published JSON Schemas when writing configuration yourself or asking a coding agent to generate it: Enum values are case-sensitive. Two common examples: Model IDs are not a fixed enum in the schema. Use gauge models list -o json after signing in to discover available models for each agent. Associate the schemas through your editor’s schema settings or use an external validator that supports YAML frontmatter. Do not add $schema to gauge.json or case frontmatter; both reject unknown fields. Run gauge evals verify -o json to check the files and their references together.

Choose your inputs

Choose the configuration that matches what you’re changing. Each example is a complete gauge.json. You can combine the named entries under inputs to test more than one together.
Gauge takes the Skill from the PR’s commit and installs it where the coding agent discovers Skills. The agent attempts your task with the revised instructions, and Gauge judges whether it completes the outcome you defined.
gauge.json
Set path to the directory containing your SKILL.md. Use "." if the Skill is at the repository root. This configuration reads the committed files directly and does not require a build workflow.Skills can also come from an Actions artifact. Select a directory, or a ZIP, .tar.gz, or .tgz file whose root contains SKILL.md.

Previews from PR comments

Use github-pr-comment when the preview provider already posts a comment that identifies both a ready URL and the candidate commit. Match the exact author login, an identifying text marker, and a readiness pattern with named url and sha captures. For example, this complete configuration expects a comment from github-actions[bot] in this format:
gauge.json
Adapt these fields to your provider’s existing comment. The pattern uses JavaScript regular expressions with multiline/global matching. The captured SHA must identify the candidate commit; comments with only a preview URL or timestamp are insufficient. Gauge uses the newest matching author/marker comment, so a newer pending or stale comment cannot silently reuse an older ready URL. This source needs a unique open PR in the same repository at the candidate head, including when launched from the CLI.

Choose when checks run

The App handles PRs opened, updated with new commits, reopened, or retargeted to another base branch. The PR must be open, come from the same repository, target an enabled branch, and contain a valid gauge.json with the connected organization’s org. Fork PRs do not launch evaluations. To skip unrelated file changes, add a top-level checks entry to gauge.json:
Use repository-relative, case-sensitive globs; matching any pattern is enough. Changes to gauge.json or configured case paths always count. Without checks.paths, every eligible update can run. After a passing evaluation, Gauge compares subsequent changes against that evaluated commit. A failed or unfinished evaluation, changed target branch, or unavailable comparison causes another evaluation rather than hiding the result behind a skip.

Run an eval only for relevant file changes

Use Gauge CLI 0.16.0 or later to validate, plan, or run cases with per-eval path filters. To give an individual eval its own file filter, add checks.paths under config in its Markdown frontmatter. For example, this case runs when documentation or its preview workflow changes:
gauge-evals/docs-onboarding.md
Use the input name, task, and repository paths that match your project. Case patterns are relative to the repository root, just like the top-level filter. A match on any pattern selects the case. Cases without config.checks.paths run whenever the suite runs. If you also set top-level checks.paths in gauge.json, that remains the gate for the whole suite. Include every path that should trigger any of its cases, or omit the top-level filter and let each case decide. A case’s config.inputs selects the inputs it receives; it does not determine when it runs. Changes to gauge.json select every case. Editing a case selects that case even if its watched paths did not change. Failed or unfinished evaluations, target-branch changes, and unavailable comparisons run conservatively instead of skipping cases. Explicit GitHub reruns bypass both levels of path filtering. If no cases match, the check is skipped without launching sessions or reserving credits. These conditions apply to automatic GitHub PR checks. Local gauge evals verify validates all discovered cases; manual gauge evals plan and gauge evals run use your case selection without filtering by changed files. A repository with no cases gets No evals configured, a skipped check with no sessions or credit reservation. This lets you commit input configuration before adding your first case. Use GitHub’s rerun action on the current Gauge evaluations check to evaluate the same commit again. Explicit reruns bypass both path filters and launch new sessions, which can consume credits. A rerun reuses a complete captured input set when its inputs are unchanged. If previously skipped cases need additional inputs, Gauge resolves a new set from existing CI or previews; it does not trigger builds. Push a new commit to evaluate a new candidate.

Read the results

Gauge maintains a summary comment for the current PR head. During preparation it shows planned cases and input waits; during execution it shows queued, running, and judging states. Pending judgments are not failures. The table links each session and shows the case, agent/model, outcome, criteria passed, duration, and token usage. For large suites it summarizes cases and limits the individual session table. Once results are settled, each failed criterion gets a separate follow-up comment with the expected behavior, the judge’s explanation, the case, agent/model, commit, and a session link. Execution and judging errors receive their own callouts so you can distinguish a failed product outcome from an incomplete evaluation. Open the linked session to inspect the agent’s work.
  • Criteria failed: the session completed, but at least one judged outcome failed.
  • Execution failed: the session failed or timed out; inspect its failure reason and trace.
  • Judging failed: Gauge could not produce a complete judgment; this is separate from a failed product outcome.
  • Input preparation failed: an artifact, preview, or configuration prevented sessions from starting.
The check completes after sessions and required judging finish. Each new head or explicit rerun updates the summary. Older sessions remain available from their links.

Troubleshoot setup

Discovery waits up to five minutes for a missing build or preview. A known queued or running build/deployment can continue waiting, with abandoned requests expiring after seven days. Waiting for inputs creates no coding sessions or session debit. A failed workflow or absent/expired artifact produces an input error.

Get support

Reach out in Slack or email [email protected] if you need help getting set up or understanding a result.