MCP server for agents
The Meticulous MCP server provides tools which enable agents to interface with Meticulous: get test run diffs, replay details, coverage, and more. It's a hosted Model Context Protocol endpoint, hosted at https://app.meticulous.ai/api/mcp. We furthermore provide agent skills which compose these tools into higher-level workflows, like a skill to review a PR test run.
- What it's for
- Connecting a client
- Authenticating with a token instead
- Available tools
- What this connector can access
- Troubleshooting
- Support and policies
What it's for
Meticulous records real user sessions in your app and replays them against every commit to catch visual regressions before they ship. This server lets an agent — reviewing a PR, debugging a failing check, or investigating coverage — pull that data directly: which screenshots changed and why, whether a diff is a real regression or noise, which lines of a change are covered by a test, and (for CI/build agents) trigger a new run against a build.
Connecting a client
The server is an OAuth 2.1 protected resource — it supports dynamic client registration and standard discovery (/.well-known/oauth-protected-resource, /.well-known/oauth-authorization-server), so clients connect with just the endpoint URL above. There is no client id, secret, scope, or authorization-server URL to configure by hand. Login happens in your browser on first connect and tokens refresh automatically.
Claude Code
Run the following in your terminal:
claude mcp add --transport http Meticulous https://app.meticulous.ai/api/mcp
Then, in Claude Code, type /mcp and choose "Authenticate" for the Meticulous MCP.
Cursor
Add to ~/.cursor/mcp.json (global) or .cursor/mcp.json (per project):
{
"mcpServers": {
"Meticulous": { "url": "https://app.meticulous.ai/api/mcp" }
}
}
Codex/ChatGPT
Add a server with name "Meticulous" and URL https://app.meticulous.ai/api/mcp, then click "Authenticate".
Calls are scoped to your default project. If you only have access to a single project, that one is used automatically; otherwise set a default in your user settings in the web app, or via the CLI: meticulous auth set-project.
Authenticating with a token instead
Any MCP client can also authenticate with a static bearer token instead of the browser OAuth flow — useful in CI, or wherever an interactive login isn't possible. The token is a Meticulous OAuth token or a project API token (see Setup > Org-wide setup for how to obtain one). A project API token scopes every call to that one project.
{
"url": "https://app.meticulous.ai/api/mcp",
"headers": { "Authorization": "Bearer <token>" }
}
Available tools
Every tool takes broadly the same arguments as its matching CLI command and returns the same data as that command's --json output — with minor differences inherent to a hosted endpoint rather than a local CLI (for example, no git-inferred commitSha, since the server has no checkout to infer it from). "Access" below reflects each tool's MCP annotations, which determine whether a client can call it without per-call confirmation.
Identity and project selection
| Tool | Access | What it does |
|---|---|---|
whoami | Read | Show the identity the connection is authenticated as, and the project it resolves to. |
list_projects | Read | List the projects the authenticated user, or API token, can access. |
get_project | Read | Show the project project-scoped tools use when not given a project argument, and where that came from. |
set_project | Write | Change the default project for the user account — every session and machine, not just this connection. |
Test runs and diffs
| Tool | Access | What it does |
|---|---|---|
get_test_run_for_commit | Read | Look up the latest test run for a commit or pull request; returns its ID and status. |
get_test_runs | Read | List a project's pull request test runs, or its base test runs, newest first, optionally for one pull request. |
get_test_run_diffs | Read | Get the (curated, full, or Similar-group) list of screenshot diffs for a test run. |
get_test_run_diffs_counts | Read | Get aggregate diff counts for a test run, including the six-way review-decision breakdown. |
get_image_urls | Read | Get signed URLs for a screenshot diff's before/after/diff images. |
get_images | Read | Get a screenshot diff's before, after, and diff images as native MCP image blocks. |
get_dom_diff | Read | Get the structural DOM diff for one screenshot diff, as unified-diff-style hunks. |
get_timeline_diff | Read | Get the list of timeline event differences (e.g. network requests, DOM mutations) for a replay diff. |
get_test_run_check | Read | Get the Markdown report for a builtin or custom non-visual check. |
get_test_run_check_available_ids | Read | List the check IDs available for a test run, for the checks that have reported results so far. |
Results that are not ready yet
Every tool that reads a test run's or replay's results — get_test_run_diffs, get_test_run_diffs_counts, get_test_run_check, the four test-run coverage tools, and the replay-level get_replay_js_coverage and get_replay_diff_js_coverage_diff — returns { status: 'processing' } (the diffs list also { status: 'pending' }) while the test run or replay is still running, or while its result is still being computed after it finished, rather than an empty or partial result that would read as a finished one. Callers should poll the same tool every 10s until it returns the result, for up to 10 minutes. A failed result (with a reason) means nothing is still computing and retrying reproduces it.
For get_test_run_diffs, a poll can itself start the underlying compute workflow; that lazy compute does not make the tool a write. A completed get_test_run_check response is { status: 'complete', text }, or { status: 'complete', text, url } when the report is too large to return inline — text is then a short notice and the full report is downloadable from url. With checkType: 'custom', an error saying the run is not expecting custom check results can be transient shortly after the run completes, since your CI registers its checks separately from the run itself — this is usually resolved by retrying for a minute or so. Use get_test_run_check_available_ids to find a valid checkId instead of guessing one.
Reviewing diffs
| Tool | Access | What it does |
|---|---|---|
get_diff_comments | Read | Get the review comments (with replies) for a screenshot diff, oldest first. |
approve_diff | Write | Agent-approve a screenshot diff, optionally commenting why, if the project allows it. |
reject_diff | Write | Agent-reject a screenshot diff and comment why. |
ignore_diff | Write | Agent-ignore a screenshot diff as unrelated to the change under review, and comment why; a real ignore only if the project allows it. |
create_diff_comment | Write | Start a review comment thread at approximate image coordinates. |
reply_to_diff_comment | Write | Reply to an existing review comment thread. |
Reviewing non-visual checks
| Tool | Access | What it does |
|---|---|---|
get_check_comments | Read | Get the reasons recorded with agent decisions on a non-visual check of a test run, oldest first. |
approve_check | Write | Agent-approve a failing non-visual check, optionally with a reason, if the project allows it. |
reject_check | Write | Agent-reject a failing non-visual check, with a reason. |
ignore_check | Write | Agent-ignore a failing non-visual check as unrelated to the change under review, with a reason, if the project allows it. |
JS coverage
| Tool | Access | What it does |
|---|---|---|
get_test_run_js_coverage | Read | Get per-file JavaScript coverage for a test run. |
get_project_js_coverage | Read | Get per-file JavaScript coverage for a project's latest successful test run. |
get_test_run_js_coverage_summary | Read | Get aggregate JavaScript coverage totals for a test run. |
get_project_js_coverage_summary | Read | Get aggregate JavaScript coverage totals for a project's latest successful test run. |
get_replay_js_coverage | Read | Get per-file JavaScript coverage for a single replay, or one screenshot of it. |
get_test_run_js_coverage_diff | Read | Get per-file JavaScript coverage differences for a test run against its own base run. |
get_test_run_js_coverage_diff_summary | Read | Get the aggregate JavaScript coverage difference for a test run against its own base run. |
get_replay_diff_js_coverage_diff | Read | Get per-file JavaScript coverage differences (base vs. head) for a replay diff. |
Sessions
| Tool | Access | What it does |
|---|---|---|
get_sessions | Read | List a project's recorded sessions, newest first by default, optionally narrowed to the selected set. |
get_session_data | Read | Get the recorded user-flow and network summary for a session — useful for understanding what a replay exercises. |
Reporting statistics
| Tool | Access | What it does |
|---|---|---|
get_test_run_stats | Read | Export reporting statistics for a project's test runs. |
get_project_daily_stats | Read | Export daily project reporting statistics matching the Metrics dashboard. |
get_test_run_event_stats | Read | Export a project's test-run reporting events such as views, reviews, and comments. |
Triggering a test run
| Tool | Access | What it does |
|---|---|---|
request_asset_upload | Write | Request a signed URL to upload a zipped static-asset build. |
register_asset_build | Write | Register an uploaded zipped asset build as an ephemeral deployment. Keep the returned deploymentId; local uploads are not discoverable later by commit SHA. |
request_container_upload | Write | Request registry credentials to push a Docker container build. |
register_container_build | Write | Register a pushed container build as an ephemeral deployment. Keep the returned deploymentId; local uploads are not discoverable later by commit SHA. |
trigger_test_run | Write | Trigger a test run against a registered deploymentId, or a persistent CI deployment identified by commit SHA. |
complete_base_run | Write | Replay the selected sessions a base run has not run yet. |
promote_sessions | Write | Add sessions a pinned-session test run replayed to the project's selected set, without waiting for session selection. |
Feedback
| Tool | Access | What it does |
|---|---|---|
submit_feedback | Write | Send free-form feedback about Meticulous to the Meticulous team. |
What this connector can access
- Reads: your identity and the projects you can access; test runs and their status; reporting statistics and engagement events; screenshot diffs, outcomes, and images; DOM and timeline diffs; JavaScript coverage; and recorded session data (the user flow and network requests a session exercises).
- Writes (only when the corresponding tool is called): uploads a build's assets or container image, registers it as a deployment, triggers a new test run against it, changes your account's default project, and submits feedback to the Meticulous team.
- Never accesses: your repository's source code, or PR titles/descriptions — Meticulous's diffs and coverage are computed from screenshots, DOM snapshots, and instrumented JS execution, not from reading your code.
- Scope: every call is scoped to the projects the authenticated user (or, for a project API token, the single owning project) has access to.
Troubleshooting
- 401/403 on every call: your token has expired or was revoked — reconnect via
/mcp(Claude Code) or your client's equivalent to re-authenticate. - Claude Code: "Issuer mismatch in authorization response (RFC 9207)": Claude Code is reusing an authorization-server URL it cached from an earlier login. In
/mcp, pick the Meticulous server and choose Clear authentication (not Re-authenticate), then Authenticate again. - "No default project" errors: call
set_project(list_projectsshows the options), set one from your user settings in the web app, or runmeticulous auth set-project— or, for a one-off call, pass the optionalprojectargument most tools accept (id,org/proj, or simplyproj; with an API token, only projects the token has access to —list_projectsshows them). - A lookup returns nothing when you expected results: it may have run against a different project. The default project is stored against your user account, so it is shared with the CLI and every other session, and changing it anywhere changes it everywhere. Call
get_projectto see which project is in play — an empty test-run or session result already names it for you. - SSO-only organizations: if your organization enforces a specific identity provider, a token issued outside that flow is rejected; log in again through your organization's SSO entry point.
- Anything else: contact us (see below).
Support and policies
- Support: support@meticulous.ai
- Privacy policy: meticulous.ai/privacy-policy
- Terms of service: meticulous.ai/terms-conditions
- Security and compliance: security.meticulous.ai