Braintrust for AI agents

Braintrust is an AI engineering platform that teams use to manage prompts, datasets, and experiments, and to evaluate and monitor AI applications in production. Connect it once through Arc0, and your agent, or Claude, ChatGPT and Cursor, can use it through one MCP endpoint, limited to what each user approved.

AI and modelsAPI keyMCP + RESTbraintrust.dev
AUDIT LOG · BRAINTRUSTPOLICY: acme-support
09:41:07 · claude · u_8f2read
braintrust.list_resources
List Braintrust Resources✓ allowed · 212ms
09:41:08 · claude · u_8f2read
braintrust.query_data
Query Braintrust Data✓ allowed · 164ms
09:41:09 · claude · u_8f2write
braintrust.create_prompt
Create Prompt✓ approved · approved by user
09:41:12 · claude · u_8f2destructive
braintrust.delete_project
Delete Project✕ blocked · policy: deny
EVERY BRAINTRUST CALL, ON THE RECORD
01 · USE CASES

What agents do in Braintrust.

01

Fetch experiment results

Read-only pull of events from an experiment or dataset to compare model evaluation runs.

02

Log a new evaluation run

Create an experiment and insert its events after the user confirms which dataset and prompt version to use.

03

Require approval before deleting a project

Treat deleting a project, experiment, or dataset as destructive since it removes evaluation history.

02 · ACTIONS

17 Braintrust actions, graded by risk.

Every Braintrust action is tagged read, write or destructive, so one policy covers the whole app and new actions inherit the right default.

read

5

Look things up. Allowed by default.

  • braintrust.query_data
    Query Braintrust Data
  • braintrust.get_resource
    Get Braintrust Resource
  • braintrust.list_resources
    List Braintrust Resources
  • braintrust.fetch_events
    Fetch Events
  • braintrust.summarize_resource
    Summarize Experiment or Dataset

write

10

Create and change things. Allow, or ask the user first.

  • braintrust.create_prompt
    Create Prompt
  • braintrust.update_prompt
    Update Prompt
  • braintrust.create_dataset
    Create Dataset
  • braintrust.create_project
    Create Project
  • braintrust.update_dataset
    Update Dataset
  • braintrust.update_project
    Update Project
  • braintrust.create_experiment
    Create Experiment
  • braintrust.update_experiment
    Update Experiment
  • braintrust.insert_dataset_events
    Insert Dataset Events
  • braintrust.insert_experiment_events
    Insert Experiment Events

destructive

2

Delete, cancel or archive. Ask first, or deny outright.

  • braintrust.delete_project
    Delete Project
  • braintrust.delete_resource
    Delete Experiment, Dataset, or Prompt
03 · HOW IT WORKS

Braintrust in three steps.

  1. 01Your users connect BraintrustThey add their Braintrust api key on Arc0 Connect, under your brand. It goes straight into the vault.
  2. 02You set the rulesReads run, writes like “create Prompt” can wait for the user, and “delete Project” can be denied outright.
  3. 03Any agent can actYour agent calls Braintrust through the Arc0 SDK or MCP, and so can Claude, ChatGPT and Cursor. Every call lands on the audit log.
POLICY.TS
await arc0.policies.set('braintrust', {
  read: 'allow',
  write: 'ask',        // create_prompt
  destructive: 'deny',  // delete_project
})

# Claude Code: the same connection, one URL
$ claude mcp add --transport http arc0 \
    https://mcp.arc0.ai/u/u_8f2
04 · AUTH AND DATA

How Braintrust connects.

Users add their Braintrust api key on Arc0 Connect. It is encrypted in the vault, never shown to the model, and each user can rotate or revoke it at any time.

The same Braintrust connection serves your agent over MCP and your own backend over REST and the proxy, so a user connects once. How Arc0 handles credentials →

AUTH
API key
CREDENTIALS
Per-tenant encrypted vault
MODEL SEES
Results only, never credentials
AUDIT LOG
Every call, on every plan
05 · WORKS WITH

Use Braintrust from any agent.

Claude
ChatGPT
Cursor
Codex
VS Code
OpenAI Agents SDK
Claude Agent SDK
Vercel AI SDK
Mastra
LangGraph
07 · FAQ

Braintrust and Arc0, answered.

Q01

Can I use Braintrust with Claude, ChatGPT or Cursor?

Yes. Connect Braintrust to Arc0 once, then add your Arc0 MCP URL to Claude, ChatGPT, Cursor, Claude Code or any other remote-MCP client. Each assistant only gets the Braintrust actions you allow.

Q02

How do users connect Braintrust?

Users add their Braintrust api key on Arc0 Connect. It is encrypted in the vault, never shown to the model, and each user can rotate or revoke it at any time.

Q03

Which Braintrust actions can my agent take?

17 in total: 5 read, 10 write and 2 destructive, such as “create Prompt”. Your policies decide which of them each agent may call.

Q04

Can I stop my agent from deleting things in Braintrust?

Yes. Actions like “delete Project” are graded destructive. Set destructive actions to deny, or to ask so the user approves each one, and blocked calls still show up on the audit log.

Q05

Can my own backend call Braintrust too?

Yes. The same Braintrust connection is available over REST and through the Arc0 proxy, so your product and your agent share one connection per user.

Get started

Plug Braintrust into your agent.

Your users connect Braintrust once, under your brand. Your agent gets 17 actions behind your policies, with every call on the record.

Free to build · MCP + REST · Audit log on every plan