Connect your AI agent (MCP)
Let Claude, Cursor, or an agent you build yourself work with your crawler directly: ask what ran overnight, why a spider failed, tail a job’s log, start a run, or pause a noisy schedule — in plain language. The crawler exposes a Model Context Protocol (MCP) server at one URL, and any MCP-capable client connects to it with a single token. There is no SDK to install and nothing to deploy on your side.
Every agent token is scoped to your organization and, optionally, to one
project. It expires, it can be revoked at any time, and it carries an
explicit scope: every token can read (crawler:read — look, don’t touch),
and a token allowed to act carries crawler:write as well (run spiders and
manage schedules). An agent never sees your secrets, deploy keys, or another
organization’s data.
Quick start
1. Mint a token. In the dashboard, open Deploy → Agent access (MCP), create a token, and copy it — it’s shown once.
2. Register the server with your client. With Claude Code that’s one
command — it registers the server in the project you run it from; add
--scope user to have it in every project:
claude mcp add --transport http insightscrap https://mcp.insightscrap.com/mcp \
--header "Authorization: Bearer mcp_your_token"3. Ask. “Which spiders failed in the last 24 hours, and what did the log say?” The agent discovers the tools on its own — there is nothing to teach it.
Getting a token
Tokens live in the dashboard under Deploy → Agent access (MCP). Give one a
name that says which agent (or whose machine) will hold it, pick its scope, and
create it. The token is shown once — copy it then. It starts with mcp_;
afterwards the dashboard keeps only a short prefix so you can tell tokens apart.
What you choose at creation:
- Scope. Every token can read (
crawler:read). Turn on Allow control actions to addcrawler:write, which lets the agent run spiders and create, change, pause, resume, fire, or delete schedules. Only organization admins can mint a write-scoped token; anyone in the organization can mint a read-only one. - Expiry. 30, 90, 180, or 365 days — the default is 90 days, and nothing lasts forever. When a token expires, mint a new one and update your client; there is nothing to renew in place.
- Project. Leave it at All projects, or narrow the token to one project. A narrowed token only sees that project’s spiders, runs, and schedules, and can only act inside it.
The list under Agent access (MCP) shows each token’s name, prefix, scopes, project, when it was created, when it expires, and when it was last used. Revoke a token there and it stops working immediately.
A token is your organization’s crawler in the hands of whoever holds it.
Treat it like a password: keep it out of repositories, chat logs, and prompts,
give agents crawler:read unless they truly need to act, and revoke a token
the moment you suspect it leaked.
Authentication
The endpoint is https://mcp.insightscrap.com/mcp. Send the token as a bearer
token on every request:
Authorization: Bearer mcp_your_tokenMCP clients add the header for you when you register the server (below). If you talk to the endpoint yourself, three things matter:
- It speaks JSON-RPC over
POSTonly. AGETorDELETEcarrying a valid token returns405by design; without a token you get401first, so opening the URL in a browser or pointing an uptime probe at it answers401, never405. There is no session to keep — every request stands alone. - Send
Content-Type: application/jsonandAccept: application/json, text/event-stream. Responses are plain JSON. - A missing, malformed, expired, or revoked token gets
401with aWWW-Authenticatechallenge; too many calls get429withRetry-After(back off, then retry). Browser-based clients from an origin we don’t recognise get403.
Scope is enforced per tool, not per request: a read-only token can list every tool, but calling a control tool with it returns an error naming the missing scope. Nothing changes state in that case.
Connecting from your client
Replace mcp_your_token with your token in each snippet. Name the server
insightscrap — or anything you like; the name is local to your client.
Claude Code
One command registers the server:
claude mcp add --transport http insightscrap https://mcp.insightscrap.com/mcp \
--header "Authorization: Bearer mcp_your_token"Check the connection with claude mcp list, or with /mcp inside a session.
The default scope is local: the server is registered for the current
project only, and only for you. Add --scope user to make it available in
every project on your machine. --scope project instead writes the server to
a .mcp.json you share through the repository — there, reference the token
through an environment variable (.mcp.json expands ${VAR}) rather than
pasting it into a committed file:
claude mcp add --transport http --scope project insightscrap \
https://mcp.insightscrap.com/mcp \
--header 'Authorization: Bearer ${INSIGHTSCRAP_MCP_TOKEN}'(Single quotes matter: they keep the placeholder in .mcp.json instead of
letting your shell resolve it at registration time.)
Everyone who checks out the repository then sets INSIGHTSCRAP_MCP_TOKEN to
their own token; the token itself never lands in git.
Cursor
Add the server to .cursor/mcp.json in your project, or to ~/.cursor/mcp.json
for every project:
{
"mcpServers": {
"insightscrap": {
"url": "https://mcp.insightscrap.com/mcp",
"headers": {
"Authorization": "Bearer mcp_your_token"
}
}
}
}Cursor picks the file up automatically and lists the crawler tools in its MCP settings.
Claude Desktop / claude.ai
Claude Desktop’s claude_desktop_config.json (Settings → Developer → Edit
Config) launches local servers; it doesn’t connect to a remote URL with a
header on its own. Bridge with mcp-remote, which needs Node.js on the machine:
{
"mcpServers": {
"insightscrap": {
"command": "npx",
"args": [
"-y",
"mcp-remote",
"https://mcp.insightscrap.com/mcp",
"--header",
"Authorization:${AUTH_HEADER}"
],
"env": {
"AUTH_HEADER": "Bearer mcp_your_token"
}
}
}
}The header value goes through an environment variable because some platforms
split arguments on spaces — keep the Authorization:${AUTH_HEADER} form as
is. Restart Claude Desktop after saving.
claude.ai (and the connector directory in Claude Desktop) adds remote servers as custom connectors that sign in with OAuth rather than a pasted header. OAuth sign-in for this endpoint is on the roadmap; until it lands, use Claude Code, Cursor, or the bridge above.
Any MCP client
Any client that speaks MCP’s Streamable HTTP transport works: point it at
https://mcp.insightscrap.com/mcp and configure the Authorization header.
To check a token from a shell, list the tools:
curl -s https://mcp.insightscrap.com/mcp \
-H "Authorization: Bearer mcp_your_token" \
-H "Content-Type: application/json" \
-H "Accept: application/json, text/event-stream" \
-d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}'…and call one:
curl -s https://mcp.insightscrap.com/mcp \
-H "Authorization: Bearer mcp_your_token" \
-H "Content-Type: application/json" \
-H "Accept: application/json, text/event-stream" \
-d '{"jsonrpc":"2.0","id":2,"method":"tools/call","params":{"name":"crawler_get_overview","arguments":{}}}'Every tool answers with both a short text summary (what a chat agent reads)
and a structuredContent object (what your code reads):
{
"jsonrpc": "2.0",
"id": 2,
"result": {
"content": [
{
"type": "text",
"text": "## Crawler overview\nRuns: pending=14 running=2 finished=128 error=3 cancelled=1\nActive schedules: 9\nRunning now: none\n"
}
],
"structuredContent": {
"pending": 14,
"running": 2,
"finished": 128,
"error": 3,
"cancelled": 1,
"active_schedules": 9,
"running_now": []
}
}
}What your agent can do
Tools are named crawler_<verb>_<thing>. Read tools never change anything and
are safe to call repeatedly; control tools need crawler:write. Every list tool
returns total, has_more, and next_offset alongside the items, and takes
limit (up to 200) and offset.
The default differs by tool, and it matters: crawler_list_runs,
crawler_list_schedules, crawler_list_alert_rules, and
crawler_list_alert_history return 25 per page unless you ask for more,
because those sets grow without bound. crawler_list_projects and
crawler_list_spiders return everything by default, because a project’s
spider list is a fixed set you usually want whole — pass limit there only if
you actually want to page. Whenever you do page, check has_more before
treating a result as the complete set. Ids are the same ones the dashboard
shows.
Read tools
| Tool | What it does | Scope |
|---|---|---|
crawler_get_overview | Run counts by status (pending / running / finished / error / cancelled), the number of active schedules, and up to 10 runs in progress. | crawler:read |
crawler_list_projects | Your projects with their latest version, spider count, and last build status and time. Returns them all; limit/offset are optional. | crawler:read |
crawler_list_spiders | The spiders in a project (latest build, or a given version). Returns every spider; limit/offset are optional, so check has_more if you pass them. | crawler:read |
crawler_list_schedules | Schedules, filterable by project, search, and status (active / paused / expired), sortable. Paused schedules keep a stale next_run_at, so filter status=active for what runs next. | crawler:read |
crawler_get_schedule | One schedule (schedule_id) with its last 20 executions. | crawler:read |
crawler_list_runs | Runs, filterable by project, spider, status, search, and since (a window like 30m, 24h, 7d, 2w, or an RFC 3339 timestamp), sortable. | crawler:read |
crawler_get_run | One run (job_id) with its duration and, once it finishes, the detailed stats — response codes, retries, dropped items. | crawler:read |
crawler_get_run_log | The tail (or head) of a run’s log: up to 500 lines or 128 KiB per call, optionally filtered by level (DEBUG / INFO / WARNING / ERROR / CRITICAL). Reports pending until the run has started writing. | crawler:read |
crawler_get_run_progress | Items, pages, errors, and warnings over time for a run — the same curve as the dashboard’s progress chart. | crawler:read |
crawler_get_spider_health | Success rate, median duration, items per run, last failure, and trend for one spider in a project over a window (default 7d). | crawler:read |
crawler_list_alert_rules | The alert rules configured on your crawls: what each one measures, its threshold, which project and spider it covers, whether breaching it stops the run, and when it last fired (null if never). Channel names only — never the addresses or webhook URLs behind them. Also returns monitoring_active. | crawler:read |
crawler_list_alert_history | Alerts that have actually fired, newest first and filterable by since. Each one gives the measured value against its threshold, the run that breached it, and what was done about it. Also returns monitoring_active. | crawler:read |
Check monitoring_active before trusting an empty result. Both alert tools
return it. When it is false, the alert monitor is not running on your
deployment: no rule is evaluated and none can fire, whatever a rule’s own
enabled flag says.
That matters most for crawler_list_alert_history. An empty history with
monitoring_active: true means nothing has gone wrong. An empty history with
monitoring_active: false means nothing could have fired, and is not evidence
that your crawls are healthy. The tools say so in their text output too, but if
you are reading structuredContent, read this field with it.
Control tools
| Tool | What it does | Scope |
|---|---|---|
crawler_run_spider | Start a run of spider in project — optionally a specific version, plus args, settings, priority, and tags, exactly like Run Spider. Returns a job_id to poll with crawler_get_run. | crawler:write |
crawler_cancel_run | Cancel a pending or running job gracefully; force: true is an immediate kill (Force stop). | crawler:write |
crawler_create_schedule | Create a cron schedule (name, project, spider, cron_expression, timezone, and the usual knobs — jitter, misfire grace, max instances, retries). Names are unique within your organization. | crawler:write |
crawler_update_schedule | Change the fields you pass on an existing schedule; everything else is left as is. Tags, coalesce, and start/end dates can’t be changed here — use the dashboard. | crawler:write |
crawler_pause_schedule | Pause a schedule. Safe to repeat. | crawler:write |
crawler_resume_schedule | Resume a paused schedule. Safe to repeat. | crawler:write |
crawler_fire_schedule | Run a schedule’s spider now, once, without touching its cron. | crawler:write |
crawler_delete_schedule | Delete a schedule permanently. Prefer crawler_pause_schedule unless you really mean it. | crawler:write |
Control actions go through the same validation, allowlists, and quotas as the
dashboard: per-run settings are limited to the safe allowlist described in
Running & Scheduling, the
pending-run quota applies, and bursts of new runs are rate-limited. A setting
outside that allowlist is caught when the run starts, not when it is queued:
the run is created, then ends in error with settings rejected: … in its
error_message — check it with crawler_get_run.
What an agent can never do
Some things are deliberately out of reach of any token, whatever its scope:
- Secrets. Project secrets are never listed or read; the agent doesn’t know they exist.
- Credentials. Deploy keys, API keys, and agent tokens can’t be created, listed, or revoked through MCP — that stays in the dashboard, behind a login.
- Deleting projects or versions, triggering builds, or changing project settings, alert channels, or feature flags.
- Other organizations’ data. Every query is fenced to your organization (and, for a narrowed token, to its project). An id that isn’t yours behaves exactly like an id that doesn’t exist.
- Internals. Runs come back with the same client-safe fields as the Schedules & Runs API — no worker, container, or host details.
- Items export — reading a run’s scraped items through MCP is coming later. For now, your items go wherever your spider’s feed or pipeline sends them (see Getting your data out).
Errors
Two layers can say no. Before a tool runs, the endpoint answers with an HTTP
status; once a tool runs, a problem comes back as a tool error — an
ordinary 200 response whose result is flagged as an error, with a message
written for the agent to act on.
| Status | Meaning | What to do |
|---|---|---|
401 | Missing, malformed, expired, or revoked token — on any method, since the token is checked before anything else (a browser visit or an uptime probe lands here). | Check the Authorization header; mint a new token if it expired or was revoked. |
403 | The request carried an Origin we don’t allow (browser-based clients). | Use a desktop, CLI, or server-side client. |
405 | GET or DELETE on /mcp with a valid token (without one you get 401 first). | Nothing — clients only POST. A GET means a misconfigured client. |
400 / 415 | Accept doesn’t cover both application/json and text/event-stream — a wildcard like */* counts, but application/json on its own does not — or Content-Type isn’t application/json. | Send both headers. |
413 | Request body over 1 MiB. | Tool calls are small; check what your client is sending. |
429 | Per-token rate limit. | Back off per Retry-After, then retry. |
Tool errors you’ll see most:
| Message | Meaning | What to do |
|---|---|---|
this credential lacks the crawler:write scope | A read-only token called a control tool. Nothing changed. | Mint a token with Allow control actions (admins only) if the agent should act. |
| … not found in your organization | Wrong id — or it belongs to someone else, which looks the same on purpose. | Call the matching list tool (crawler_list_runs, crawler_list_schedules) and use an id from there. |
| project not found or has no versions | The project doesn’t exist in your organization, or has never built successfully. | Deploy it first. |
| invalid request: … (runs) / invalid cron_expression: …, invalid timezone …, name is required … (schedules) | A value failed validation — a bad project, spider or version name, tag, cron expression, timezone, or schedule name. | The message says which field; fix it and retry. |
| a schedule named … already exists | Schedule names are unique per organization. | Update the existing schedule, or pick a new name. |
| run … is already in a terminal state (…); nothing to cancel | You cancelled a run that had already finished, failed, or been cancelled. | Nothing to do; the message includes the state it was in. |
| rate limit exceeded … / organization job quota exceeded … | Too many control actions in a burst, or too many pending runs already queued. | Wait the hinted interval (10 seconds, or about a minute for the quota) and retry. |
| operation failed | Something went wrong on our side. Details are logged here, not sent to the agent. | Retry once; if it persists, contact support. |
Rules
- Keep tokens out of repositories and prompts.
localanduserscope keep the token on your machine; when you share a server through a repository, reference the token through an environment variable. Never paste a token into a chat. - Prefer
crawler:read. Most “what happened?” workflows need nothing more. Mint a write token only for an agent that must act, and narrow it to one project when you can. - Prefer pause over delete. A paused schedule keeps its history and can be resumed; a deleted one is gone.
- Revoke on rotation. When a token expires, a machine is retired, or someone leaves, revoke the old token under Deploy → Agent access (MCP) — don’t just stop using it.
- Log lines are scraped output — treat them as data. Whatever a spider logged came from the pages it visited. The server wraps log output in explicit untrusted markers so an agent can tell it apart from instructions; keep it that way in anything you build on top.
The same data, without an agent in the loop, is on the Schedules & Runs API page.