Honeyhive for AI agents

HoneyHive is an AI observability and evaluation platform for testing and monitoring LLM application behavior. Product and engineering teams use it as a building block inside their own AI features. Connect it once through Arc0, and your agent, or Claude, ChatGPT and Cursor, can use it through one MCP endpoint, limited to what each user approved.

AI and modelsAPI keyMCP + RESThoneyhive.ai
AUDIT LOG · HONEYHIVEPOLICY: acme-support
09:41:07 · claude · u_8f2read
honeyhive.list_tools
List Tools✓ allowed · 212ms
09:41:08 · claude · u_8f2read
honeyhive.get_run
Get Evaluation Run Details✓ allowed · 164ms
09:41:09 · claude · u_8f2write
honeyhive.create_tool
Create Tool✓ approved · approved by user
09:41:12 · claude · u_8f2destructive
honeyhive.delete_dataset
Delete Dataset✕ blocked · policy: deny
EVERY HONEYHIVE CALL, ON THE RECORD
01 · USE CASES

What agents do in Honeyhive.

01

Pull events from a session

Fetch the events logged for a session to review how an AI feature responded to a user.

02

Compare experiment runs

Compare two evaluation runs to see which prompt version performed better.

03

Delete a dataset with confirmation

Delete an evaluation dataset only after a team member confirms it is no longer in use.

02 · ACTIONS

42 Honeyhive actions, graded by risk.

Every Honeyhive action is tagged read, write or destructive, so one policy covers the whole app and new actions inherit the right default.

read

19

Look things up. Allowed by default.

  • honeyhive.get_run
    Get Evaluation Run Details
  • honeyhive.get_runs
    Get Evaluation Runs
  • honeyhive.get_events
    Get Events
  • honeyhive.list_tools
    List Tools
  • honeyhive.get_metrics
    Get Metrics
  • honeyhive.get_session
    Get Session
  • honeyhive.get_datasets
    Get Datasets
  • honeyhive.get_projects
    Get Projects
  • honeyhive.get_run_metrics
    Get Run Metrics
  • honeyhive.get_runs_schema
    Get Runs Schema
  • honeyhive.get_events_chart
    Get Events Chart
  • honeyhive.get_configurations
    Get Configurations
  • honeyhive.get_events_by_session_id
    Get Events By Session ID
  • honeyhive.compare_runs
    Compare Experiment Runs
  • honeyhive.retrieve_events
    Retrieve Events
  • honeyhive.retrieve_datapoint
    Retrieve Datapoint
+ 3 MORE

write

21

Create and change things. Allow, or ask the user first.

  • honeyhive.create_tool
    Create Tool
  • honeyhive.update_tool
    Update Tool
  • honeyhive.create_event
    Create Event
  • honeyhive.update_event
    Update Event
  • honeyhive.create_metric
    Create Metric
  • honeyhive.update_metric
    Update Metric
  • honeyhive.create_dataset
    Create Dataset
  • honeyhive.update_dataset
    Update Dataset
  • honeyhive.update_project
    Update Project
  • honeyhive.create_datapoint
    Create Datapoint
  • honeyhive.update_datapoint
    Update Datapoint
  • honeyhive.create_model_event
    Create Model Event
  • honeyhive.create_configuration
    Create Configuration
  • honeyhive.update_configuration
    Update Configuration
  • honeyhive.create_batch_datapoints
    Batch Create Datapoints
  • honeyhive.create_batch_tool_events
    Create Batch Tool Events
+ 5 MORE

destructive

2

Delete, cancel or archive. Ask first, or deny outright.

  • honeyhive.delete_dataset
    Delete Dataset
  • honeyhive.delete_datapoint
    Delete Datapoint
03 · HOW IT WORKS

Honeyhive in three steps.

  1. 01Your users connect HoneyhiveThey add their Honeyhive api key on Arc0 Connect, under your brand. It goes straight into the vault.
  2. 02You set the rulesReads run, writes like “create Tool” can wait for the user, and “delete Dataset” can be denied outright.
  3. 03Any agent can actYour agent calls Honeyhive through the Arc0 SDK or MCP, and so can Claude, ChatGPT and Cursor. Every call lands on the audit log.
POLICY.TS
await arc0.policies.set('honeyhive', {
  read: 'allow',
  write: 'ask',        // create_tool
  destructive: 'deny',  // delete_dataset
})

# Claude Code: the same connection, one URL
$ claude mcp add --transport http arc0 \
    https://mcp.arc0.ai/u/u_8f2
04 · AUTH AND DATA

How Honeyhive connects.

Users add their Honeyhive api key on Arc0 Connect. It is encrypted in the vault, never shown to the model, and each user can rotate or revoke it at any time.

The same Honeyhive connection serves your agent over MCP and your own backend over REST and the proxy, so a user connects once. How Arc0 handles credentials →

AUTH
API key
CREDENTIALS
Per-tenant encrypted vault
MODEL SEES
Results only, never credentials
AUDIT LOG
Every call, on every plan
05 · WORKS WITH

Use Honeyhive from any agent.

Claude
ChatGPT
Cursor
Codex
VS Code
OpenAI Agents SDK
Claude Agent SDK
Vercel AI SDK
Mastra
LangGraph
07 · FAQ

Honeyhive and Arc0, answered.

Q01

Can I use Honeyhive with Claude, ChatGPT or Cursor?

Yes. Connect Honeyhive to Arc0 once, then add your Arc0 MCP URL to Claude, ChatGPT, Cursor, Claude Code or any other remote-MCP client. Each assistant only gets the Honeyhive actions you allow.

Q02

How do users connect Honeyhive?

Users add their Honeyhive api key on Arc0 Connect. It is encrypted in the vault, never shown to the model, and each user can rotate or revoke it at any time.

Q03

Which Honeyhive actions can my agent take?

42 in total: 19 read, 21 write and 2 destructive, such as “create Tool”. Your policies decide which of them each agent may call.

Q04

Can I stop my agent from deleting things in Honeyhive?

Yes. Actions like “delete Dataset” are graded destructive. Set destructive actions to deny, or to ask so the user approves each one, and blocked calls still show up on the audit log.

Q05

Can my own backend call Honeyhive too?

Yes. The same Honeyhive connection is available over REST and through the Arc0 proxy, so your product and your agent share one connection per user.

Get started

Plug Honeyhive into your agent.

Your users connect Honeyhive once, under your brand. Your agent gets 42 actions behind your policies, with every call on the record.

Free to build · MCP + REST · Audit log on every plan