Moderation API for AI agents

Moderation API detects harmful content and routes it through human review queues for platforms that need content moderation at scale. Connect it once through Arc0, and your agent, or Claude, ChatGPT and Cursor, can use it through one MCP endpoint, limited to what each user approved.

AI and modelsAPI keyMCP + RESTmoderationapi.com
AUDIT LOG · MODERATION APIPOLICY: acme-support
09:41:07 · claude · u_8f2read
moderation_api.list_actions
List Moderation Actions✓ allowed · 212ms
09:41:08 · claude · u_8f2read
moderation_api.get_account
Get Account and Quota✓ allowed · 164ms
09:41:09 · claude · u_8f2write
moderation_api.execute_action
Execute Moderation Action✓ approved · approved by user
EVERY MODERATION API CALL, ON THE RECORD
01 · USE CASES

What agents do in Moderation API.

01

Evaluate content before publishing

Evaluate a piece of user content against moderation rules without storing it, a read-only check.

02

Check the review queue

List items waiting in the review queue to prioritize what a moderator looks at next.

03

Approve before taking a moderation action

Require approval before executing a moderation action, since it can remove or restrict user content.

02 · ACTIONS

11 Moderation API actions, graded by risk.

Every Moderation API action is tagged read, write or destructive, so one policy covers the whole app and new actions inherit the right default.

read

6

Look things up. Allowed by default.

  • moderation_api.get_account
    Get Account and Quota
  • moderation_api.get_wordlist
    Get Wordlist
  • moderation_api.list_actions
    List Moderation Actions
  • moderation_api.list_wordlists
    List Wordlists
  • moderation_api.get_review_queue
    Get Review Queue
  • moderation_api.list_review_queue_items
    List Review Queue Items

write

5

Create and change things. Allow, or ask the user first.

  • moderation_api.execute_action
    Execute Moderation Action
  • moderation_api.evaluate_content
    Evaluate Content Without Storage
  • moderation_api.modify_wordlist_words
    Modify Wordlist Words
  • moderation_api.set_review_queue_item_status
    Set Review Queue Item Status
  • moderation_api.submit_content_for_moderation
    Submit Content for Moderation

destructive

0

Delete, cancel or archive. Ask first, or deny outright.

  • No destructive actions.
03 · HOW IT WORKS

Moderation API in three steps.

  1. 01Your users connect Moderation APIThey add their Moderation API api key on Arc0 Connect, under your brand. It goes straight into the vault.
  2. 02You set the rulesReads run, and writes like “execute Moderation Action” can wait for the user to approve.
  3. 03Any agent can actYour agent calls Moderation API through the Arc0 SDK or MCP, and so can Claude, ChatGPT and Cursor. Every call lands on the audit log.
POLICY.TS
await arc0.policies.set('moderation_api', {
  read: 'allow',
  write: 'ask',        // execute_action
  destructive: 'deny',  
})

# Claude Code: the same connection, one URL
$ claude mcp add --transport http arc0 \
    https://mcp.arc0.ai/u/u_8f2
04 · AUTH AND DATA

How Moderation API connects.

Users add their Moderation API api key on Arc0 Connect. It is encrypted in the vault, never shown to the model, and each user can rotate or revoke it at any time.

The same Moderation API connection serves your agent over MCP and your own backend over REST and the proxy, so a user connects once. How Arc0 handles credentials →

AUTH
API key
CREDENTIALS
Per-tenant encrypted vault
MODEL SEES
Results only, never credentials
AUDIT LOG
Every call, on every plan
05 · WORKS WITH

Use Moderation API from any agent.

Claude
ChatGPT
Cursor
Codex
VS Code
OpenAI Agents SDK
Claude Agent SDK
Vercel AI SDK
Mastra
LangGraph
07 · FAQ

Moderation API and Arc0, answered.

Q01

Can I use Moderation API with Claude, ChatGPT or Cursor?

Yes. Connect Moderation API to Arc0 once, then add your Arc0 MCP URL to Claude, ChatGPT, Cursor, Claude Code or any other remote-MCP client. Each assistant only gets the Moderation API actions you allow.

Q02

How do users connect Moderation API?

Users add their Moderation API api key on Arc0 Connect. It is encrypted in the vault, never shown to the model, and each user can rotate or revoke it at any time.

Q03

Which Moderation API actions can my agent take?

11 in total: 6 read, 5 write and 0 destructive, such as “execute Moderation Action”. Your policies decide which of them each agent may call.

Q04

Can I make my agent read-only in Moderation API?

Yes. Allow read actions and deny writes in the Moderation API policy. Your agent can still look things up, and any write it attempts is blocked and logged.

Q05

Can my own backend call Moderation API too?

Yes. The same Moderation API connection is available over REST and through the Arc0 proxy, so your product and your agent share one connection per user.

Get started

Plug Moderation API into your agent.

Your users connect Moderation API once, under your brand. Your agent gets 11 actions behind your policies, with every call on the record.

Free to build · MCP + REST · Audit log on every plan