Moderation API for AI agents
Moderation API detects harmful content and routes it through human review queues for platforms that need content moderation at scale. Connect it once through Arc0, and your agent, or Claude, ChatGPT and Cursor, can use it through one MCP endpoint, limited to what each user approved.
What agents do in Moderation API.
Evaluate content before publishing
Evaluate a piece of user content against moderation rules without storing it, a read-only check.
Check the review queue
List items waiting in the review queue to prioritize what a moderator looks at next.
Approve before taking a moderation action
Require approval before executing a moderation action, since it can remove or restrict user content.
11 Moderation API actions, graded by risk.
Every Moderation API action is tagged read, write or destructive, so one policy covers the whole app and new actions inherit the right default.
read
6Look things up. Allowed by default.
- moderation_api.get_accountGet Account and Quota
- moderation_api.get_wordlistGet Wordlist
- moderation_api.list_actionsList Moderation Actions
- moderation_api.list_wordlistsList Wordlists
- moderation_api.get_review_queueGet Review Queue
- moderation_api.list_review_queue_itemsList Review Queue Items
write
5Create and change things. Allow, or ask the user first.
- moderation_api.execute_actionExecute Moderation Action
- moderation_api.evaluate_contentEvaluate Content Without Storage
- moderation_api.modify_wordlist_wordsModify Wordlist Words
- moderation_api.set_review_queue_item_statusSet Review Queue Item Status
- moderation_api.submit_content_for_moderationSubmit Content for Moderation
destructive
0Delete, cancel or archive. Ask first, or deny outright.
- No destructive actions.
Moderation API in three steps.
- 01Your users connect Moderation APIThey add their Moderation API api key on Arc0 Connect, under your brand. It goes straight into the vault.
- 02You set the rulesReads run, and writes like “execute Moderation Action” can wait for the user to approve.
- 03Any agent can actYour agent calls Moderation API through the Arc0 SDK or MCP, and so can Claude, ChatGPT and Cursor. Every call lands on the audit log.
await arc0.policies.set('moderation_api', { read: 'allow', write: 'ask', // execute_action destructive: 'deny', }) # Claude Code: the same connection, one URL $ claude mcp add --transport http arc0 \ https://mcp.arc0.ai/u/u_8f2
How Moderation API connects.
Users add their Moderation API api key on Arc0 Connect. It is encrypted in the vault, never shown to the model, and each user can rotate or revoke it at any time.
The same Moderation API connection serves your agent over MCP and your own backend over REST and the proxy, so a user connects once. How Arc0 handles credentials →
- AUTH
- API key
- CREDENTIALS
- Per-tenant encrypted vault
- MODEL SEES
- Results only, never credentials
- AUDIT LOG
- Every call, on every plan
Use Moderation API from any agent.
Moderation API and Arc0, answered.
Can I use Moderation API with Claude, ChatGPT or Cursor?
Yes. Connect Moderation API to Arc0 once, then add your Arc0 MCP URL to Claude, ChatGPT, Cursor, Claude Code or any other remote-MCP client. Each assistant only gets the Moderation API actions you allow.
How do users connect Moderation API?
Users add their Moderation API api key on Arc0 Connect. It is encrypted in the vault, never shown to the model, and each user can rotate or revoke it at any time.
Which Moderation API actions can my agent take?
11 in total: 6 read, 5 write and 0 destructive, such as “execute Moderation Action”. Your policies decide which of them each agent may call.
Can I make my agent read-only in Moderation API?
Yes. Allow read actions and deny writes in the Moderation API policy. Your agent can still look things up, and any write it attempts is blocked and logged.
Can my own backend call Moderation API too?
Yes. The same Moderation API connection is available over REST and through the Arc0 proxy, so your product and your agent share one connection per user.
Plug Moderation API into your agent.
Your users connect Moderation API once, under your brand. Your agent gets 11 actions behind your policies, with every call on the record.
Free to build · MCP + REST · Audit log on every plan