Scorecard for AI agents

Scorecard is an evaluation platform for AI systems, letting teams manage test sets, test cases, evaluation runs, and scores to measure model and agent quality. Connect it once through Arc0, and your agent, or Claude, ChatGPT and Cursor, can use it through one MCP endpoint, limited to what each user approved.

AI and modelsOAuth 2.0 (MCP)MCP + RESTscorecard.io
AUDIT LOG · SCORECARDPOLICY: acme-support
09:41:07 · claude · u_8f2read
scorecard.search_docs
Search docs✓ allowed · 212ms
09:41:09 · claude · u_8f2write
scorecard.execute
Execute✓ approved · approved by user
EVERY SCORECARD CALL, ON THE RECORD
01 · USE CASES

What agents do in Scorecard.

01

Run an evaluation on a test set

Executes a scoring run against a defined test set to measure how a system performs.

02

Look up scores for a run

Reads recorded scores and annotations from a past evaluation run before deciding on changes.

03

Search evaluation documentation

Looks up how a metric or scoring method works before configuring a new evaluation project.

02 · ACTIONS

2 Scorecard actions, graded by risk.

Every Scorecard action is tagged read, write or destructive, so one policy covers the whole app and new actions inherit the right default.

read

1

Look things up. Allowed by default.

  • scorecard.search_docs
    Search docs

write

1

Create and change things. Allow, or ask the user first.

  • scorecard.execute
    Execute

destructive

0

Delete, cancel or archive. Ask first, or deny outright.

  • No destructive actions.
03 · HOW IT WORKS

Scorecard in three steps.

  1. 01Your users connect ScorecardThey sign in to Scorecard on Arc0 Connect, under your brand, and approve the access you ask for.
  2. 02You set the rulesReads run, and writes like “execute” can wait for the user to approve.
  3. 03Any agent can actYour agent calls Scorecard through the Arc0 SDK or MCP, and so can Claude, ChatGPT and Cursor. Every call lands on the audit log.
POLICY.TS
await arc0.policies.set('scorecard', {
  read: 'allow',
  write: 'ask',        // execute
  destructive: 'deny',  
})

# Claude Code: the same connection, one URL
$ claude mcp add --transport http arc0 \
    https://mcp.arc0.ai/u/u_8f2
04 · AUTH AND DATA

How Scorecard connects.

Scorecard runs its own MCP server behind OAuth. Arc0 registers the client, users approve access on Arc0 Connect, and your agent reaches Scorecard through the same endpoint, policies and audit log as every other app.

The same Scorecard connection serves your agent over MCP and your own backend over REST and the proxy, so a user connects once. How Arc0 handles credentials →

AUTH
OAuth 2.0 (MCP)
CREDENTIALS
Per-tenant encrypted vault
MODEL SEES
Results only, never credentials
AUDIT LOG
Every call, on every plan
05 · WORKS WITH

Use Scorecard from any agent.

Claude
ChatGPT
Cursor
Codex
VS Code
OpenAI Agents SDK
Claude Agent SDK
Vercel AI SDK
Mastra
LangGraph
07 · FAQ

Scorecard and Arc0, answered.

Q01

Can I use Scorecard with Claude, ChatGPT or Cursor?

Yes. Connect Scorecard to Arc0 once, then add your Arc0 MCP URL to Claude, ChatGPT, Cursor, Claude Code or any other remote-MCP client. Each assistant only gets the Scorecard actions you allow.

Q02

How do users connect Scorecard?

Scorecard runs its own MCP server behind OAuth. Arc0 registers the client, users approve access on Arc0 Connect, and your agent reaches Scorecard through the same endpoint, policies and audit log as every other app.

Q03

Which Scorecard actions can my agent take?

2 in total: 1 read, 1 write and 0 destructive, such as “execute”. Your policies decide which of them each agent may call.

Q04

Can I make my agent read-only in Scorecard?

Yes. Allow read actions and deny writes in the Scorecard policy. Your agent can still look things up, and any write it attempts is blocked and logged.

Q05

Can my own backend call Scorecard too?

Yes. The same Scorecard connection is available over REST and through the Arc0 proxy, so your product and your agent share one connection per user.

Get started

Plug Scorecard into your agent.

Your users connect Scorecard once, under your brand. Your agent gets 2 actions behind your policies, with every call on the record.

Free to build · MCP + REST · Audit log on every plan