Skip to Content
CrawlersConnect your AI agent (MCP)

Connect your AI agent (MCP)

Let Claude, Cursor, or an agent you build yourself work with your crawler directly: ask what ran overnight, why a spider failed, tail a job’s log, start a run, or pause a noisy schedule — in plain language. The crawler exposes a Model Context Protocol  (MCP) server at one URL, and any MCP-capable client connects to it with a single token. There is no SDK to install and nothing to deploy on your side.

Every agent token is scoped to your organization and, optionally, to one project. It expires, it can be revoked at any time, and it carries an explicit scope: every token can read (crawler:read — look, don’t touch), and a token allowed to act carries crawler:write as well (run spiders and manage schedules). An agent never sees your secrets, deploy keys, or another organization’s data.

Quick start

1. Mint a token. In the dashboard, open Deploy → Agent access (MCP), create a token, and copy it — it’s shown once.

2. Register the server with your client. With Claude Code that’s one command — it registers the server in the project you run it from; add --scope user to have it in every project:

claude mcp add --transport http insightscrap https://mcp.insightscrap.com/mcp \ --header "Authorization: Bearer mcp_your_token"

3. Ask. “Which spiders failed in the last 24 hours, and what did the log say?” The agent discovers the tools on its own — there is nothing to teach it.

Getting a token

Tokens live in the dashboard under Deploy → Agent access (MCP). Give one a name that says which agent (or whose machine) will hold it, pick its scope, and create it. The token is shown once — copy it then. It starts with mcp_; afterwards the dashboard keeps only a short prefix so you can tell tokens apart.

What you choose at creation:

  • Scope. Every token can read (crawler:read). Turn on Allow control actions to add crawler:write, which lets the agent run spiders and create, change, pause, resume, fire, or delete schedules. Only organization admins can mint a write-scoped token; anyone in the organization can mint a read-only one.
  • Expiry. 30, 90, 180, or 365 days — the default is 90 days, and nothing lasts forever. When a token expires, mint a new one and update your client; there is nothing to renew in place.
  • Project. Leave it at All projects, or narrow the token to one project. A narrowed token only sees that project’s spiders, runs, and schedules, and can only act inside it.

The list under Agent access (MCP) shows each token’s name, prefix, scopes, project, when it was created, when it expires, and when it was last used. Revoke a token there and it stops working immediately.

⚠️

A token is your organization’s crawler in the hands of whoever holds it. Treat it like a password: keep it out of repositories, chat logs, and prompts, give agents crawler:read unless they truly need to act, and revoke a token the moment you suspect it leaked.

Authentication

The endpoint is https://mcp.insightscrap.com/mcp. Send the token as a bearer token on every request:

Authorization: Bearer mcp_your_token

MCP clients add the header for you when you register the server (below). If you talk to the endpoint yourself, three things matter:

  • It speaks JSON-RPC over POST only. A GET or DELETE carrying a valid token returns 405 by design; without a token you get 401 first, so opening the URL in a browser or pointing an uptime probe at it answers 401, never 405. There is no session to keep — every request stands alone.
  • Send Content-Type: application/json and Accept: application/json, text/event-stream. Responses are plain JSON.
  • A missing, malformed, expired, or revoked token gets 401 with a WWW-Authenticate challenge; too many calls get 429 with Retry-After (back off, then retry). Browser-based clients from an origin we don’t recognise get 403.

Scope is enforced per tool, not per request: a read-only token can list every tool, but calling a control tool with it returns an error naming the missing scope. Nothing changes state in that case.

Connecting from your client

Replace mcp_your_token with your token in each snippet. Name the server insightscrap — or anything you like; the name is local to your client.

Claude Code

One command registers the server:

claude mcp add --transport http insightscrap https://mcp.insightscrap.com/mcp \ --header "Authorization: Bearer mcp_your_token"

Check the connection with claude mcp list, or with /mcp inside a session.

The default scope is local: the server is registered for the current project only, and only for you. Add --scope user to make it available in every project on your machine. --scope project instead writes the server to a .mcp.json you share through the repository — there, reference the token through an environment variable (.mcp.json expands ${VAR}) rather than pasting it into a committed file:

claude mcp add --transport http --scope project insightscrap \ https://mcp.insightscrap.com/mcp \ --header 'Authorization: Bearer ${INSIGHTSCRAP_MCP_TOKEN}'

(Single quotes matter: they keep the placeholder in .mcp.json instead of letting your shell resolve it at registration time.)

Everyone who checks out the repository then sets INSIGHTSCRAP_MCP_TOKEN to their own token; the token itself never lands in git.

Cursor

Add the server to .cursor/mcp.json in your project, or to ~/.cursor/mcp.json for every project:

{ "mcpServers": { "insightscrap": { "url": "https://mcp.insightscrap.com/mcp", "headers": { "Authorization": "Bearer mcp_your_token" } } } }

Cursor picks the file up automatically and lists the crawler tools in its MCP settings.

Claude Desktop / claude.ai

Claude Desktop’s claude_desktop_config.json (Settings → Developer → Edit Config) launches local servers; it doesn’t connect to a remote URL with a header on its own. Bridge with mcp-remote, which needs Node.js on the machine:

{ "mcpServers": { "insightscrap": { "command": "npx", "args": [ "-y", "mcp-remote", "https://mcp.insightscrap.com/mcp", "--header", "Authorization:${AUTH_HEADER}" ], "env": { "AUTH_HEADER": "Bearer mcp_your_token" } } } }

The header value goes through an environment variable because some platforms split arguments on spaces — keep the Authorization:${AUTH_HEADER} form as is. Restart Claude Desktop after saving.

claude.ai (and the connector directory in Claude Desktop) adds remote servers as custom connectors that sign in with OAuth rather than a pasted header. OAuth sign-in for this endpoint is on the roadmap; until it lands, use Claude Code, Cursor, or the bridge above.

Any MCP client

Any client that speaks MCP’s Streamable HTTP transport works: point it at https://mcp.insightscrap.com/mcp and configure the Authorization header. To check a token from a shell, list the tools:

curl -s https://mcp.insightscrap.com/mcp \ -H "Authorization: Bearer mcp_your_token" \ -H "Content-Type: application/json" \ -H "Accept: application/json, text/event-stream" \ -d '{"jsonrpc":"2.0","id":1,"method":"tools/list"}'

…and call one:

curl -s https://mcp.insightscrap.com/mcp \ -H "Authorization: Bearer mcp_your_token" \ -H "Content-Type: application/json" \ -H "Accept: application/json, text/event-stream" \ -d '{"jsonrpc":"2.0","id":2,"method":"tools/call","params":{"name":"crawler_get_overview","arguments":{}}}'

Every tool answers with both a short text summary (what a chat agent reads) and a structuredContent object (what your code reads):

{ "jsonrpc": "2.0", "id": 2, "result": { "content": [ { "type": "text", "text": "## Crawler overview\nRuns: pending=14 running=2 finished=128 error=3 cancelled=1\nActive schedules: 9\nRunning now: none\n" } ], "structuredContent": { "pending": 14, "running": 2, "finished": 128, "error": 3, "cancelled": 1, "active_schedules": 9, "running_now": [] } } }

What your agent can do

Tools are named crawler_<verb>_<thing>. Read tools never change anything and are safe to call repeatedly; control tools need crawler:write. Every list tool returns total, has_more, and next_offset alongside the items, and takes limit (up to 200) and offset.

The default differs by tool, and it matters: crawler_list_runs, crawler_list_schedules, crawler_list_alert_rules, and crawler_list_alert_history return 25 per page unless you ask for more, because those sets grow without bound. crawler_list_projects and crawler_list_spiders return everything by default, because a project’s spider list is a fixed set you usually want whole — pass limit there only if you actually want to page. Whenever you do page, check has_more before treating a result as the complete set. Ids are the same ones the dashboard shows.

Read tools

ToolWhat it doesScope
crawler_get_overviewRun counts by status (pending / running / finished / error / cancelled), the number of active schedules, and up to 10 runs in progress.crawler:read
crawler_list_projectsYour projects with their latest version, spider count, and last build status and time. Returns them all; limit/offset are optional.crawler:read
crawler_list_spidersThe spiders in a project (latest build, or a given version). Returns every spider; limit/offset are optional, so check has_more if you pass them.crawler:read
crawler_list_schedulesSchedules, filterable by project, search, and status (active / paused / expired), sortable. Paused schedules keep a stale next_run_at, so filter status=active for what runs next.crawler:read
crawler_get_scheduleOne schedule (schedule_id) with its last 20 executions.crawler:read
crawler_list_runsRuns, filterable by project, spider, status, search, and since (a window like 30m, 24h, 7d, 2w, or an RFC 3339 timestamp), sortable.crawler:read
crawler_get_runOne run (job_id) with its duration and, once it finishes, the detailed stats — response codes, retries, dropped items.crawler:read
crawler_get_run_logThe tail (or head) of a run’s log: up to 500 lines or 128 KiB per call, optionally filtered by level (DEBUG / INFO / WARNING / ERROR / CRITICAL). Reports pending until the run has started writing.crawler:read
crawler_get_run_progressItems, pages, errors, and warnings over time for a run — the same curve as the dashboard’s progress chart.crawler:read
crawler_get_spider_healthSuccess rate, median duration, items per run, last failure, and trend for one spider in a project over a window (default 7d).crawler:read
crawler_list_alert_rulesThe alert rules configured on your crawls: what each one measures, its threshold, which project and spider it covers, whether breaching it stops the run, and when it last fired (null if never). Channel names only — never the addresses or webhook URLs behind them. Also returns monitoring_active.crawler:read
crawler_list_alert_historyAlerts that have actually fired, newest first and filterable by since. Each one gives the measured value against its threshold, the run that breached it, and what was done about it. Also returns monitoring_active.crawler:read
⚠️

Check monitoring_active before trusting an empty result. Both alert tools return it. When it is false, the alert monitor is not running on your deployment: no rule is evaluated and none can fire, whatever a rule’s own enabled flag says.

That matters most for crawler_list_alert_history. An empty history with monitoring_active: true means nothing has gone wrong. An empty history with monitoring_active: false means nothing could have fired, and is not evidence that your crawls are healthy. The tools say so in their text output too, but if you are reading structuredContent, read this field with it.

Control tools

ToolWhat it doesScope
crawler_run_spiderStart a run of spider in project — optionally a specific version, plus args, settings, priority, and tags, exactly like Run Spider. Returns a job_id to poll with crawler_get_run.crawler:write
crawler_cancel_runCancel a pending or running job gracefully; force: true is an immediate kill (Force stop).crawler:write
crawler_create_scheduleCreate a cron schedule (name, project, spider, cron_expression, timezone, and the usual knobs — jitter, misfire grace, max instances, retries). Names are unique within your organization.crawler:write
crawler_update_scheduleChange the fields you pass on an existing schedule; everything else is left as is. Tags, coalesce, and start/end dates can’t be changed here — use the dashboard.crawler:write
crawler_pause_schedulePause a schedule. Safe to repeat.crawler:write
crawler_resume_scheduleResume a paused schedule. Safe to repeat.crawler:write
crawler_fire_scheduleRun a schedule’s spider now, once, without touching its cron.crawler:write
crawler_delete_scheduleDelete a schedule permanently. Prefer crawler_pause_schedule unless you really mean it.crawler:write

Control actions go through the same validation, allowlists, and quotas as the dashboard: per-run settings are limited to the safe allowlist described in Running & Scheduling, the pending-run quota applies, and bursts of new runs are rate-limited. A setting outside that allowlist is caught when the run starts, not when it is queued: the run is created, then ends in error with settings rejected: … in its error_message — check it with crawler_get_run.

What an agent can never do

Some things are deliberately out of reach of any token, whatever its scope:

  • Secrets. Project secrets are never listed or read; the agent doesn’t know they exist.
  • Credentials. Deploy keys, API keys, and agent tokens can’t be created, listed, or revoked through MCP — that stays in the dashboard, behind a login.
  • Deleting projects or versions, triggering builds, or changing project settings, alert channels, or feature flags.
  • Other organizations’ data. Every query is fenced to your organization (and, for a narrowed token, to its project). An id that isn’t yours behaves exactly like an id that doesn’t exist.
  • Internals. Runs come back with the same client-safe fields as the Schedules & Runs API — no worker, container, or host details.
  • Items export — reading a run’s scraped items through MCP is coming later. For now, your items go wherever your spider’s feed or pipeline sends them (see Getting your data out).

Errors

Two layers can say no. Before a tool runs, the endpoint answers with an HTTP status; once a tool runs, a problem comes back as a tool error — an ordinary 200 response whose result is flagged as an error, with a message written for the agent to act on.

StatusMeaningWhat to do
401Missing, malformed, expired, or revoked token — on any method, since the token is checked before anything else (a browser visit or an uptime probe lands here).Check the Authorization header; mint a new token if it expired or was revoked.
403The request carried an Origin we don’t allow (browser-based clients).Use a desktop, CLI, or server-side client.
405GET or DELETE on /mcp with a valid token (without one you get 401 first).Nothing — clients only POST. A GET means a misconfigured client.
400 / 415Accept doesn’t cover both application/json and text/event-stream — a wildcard like */* counts, but application/json on its own does not — or Content-Type isn’t application/json.Send both headers.
413Request body over 1 MiB.Tool calls are small; check what your client is sending.
429Per-token rate limit.Back off per Retry-After, then retry.

Tool errors you’ll see most:

MessageMeaningWhat to do
this credential lacks the crawler:write scopeA read-only token called a control tool. Nothing changed.Mint a token with Allow control actions (admins only) if the agent should act.
… not found in your organizationWrong id — or it belongs to someone else, which looks the same on purpose.Call the matching list tool (crawler_list_runs, crawler_list_schedules) and use an id from there.
project not found or has no versionsThe project doesn’t exist in your organization, or has never built successfully.Deploy it first.
invalid request: … (runs) / invalid cron_expression: …, invalid timezone …, name is required … (schedules)A value failed validation — a bad project, spider or version name, tag, cron expression, timezone, or schedule name.The message says which field; fix it and retry.
a schedule named … already existsSchedule names are unique per organization.Update the existing schedule, or pick a new name.
run … is already in a terminal state (…); nothing to cancelYou cancelled a run that had already finished, failed, or been cancelled.Nothing to do; the message includes the state it was in.
rate limit exceeded … / organization job quota exceeded …Too many control actions in a burst, or too many pending runs already queued.Wait the hinted interval (10 seconds, or about a minute for the quota) and retry.
operation failedSomething went wrong on our side. Details are logged here, not sent to the agent.Retry once; if it persists, contact support.

Rules

  • Keep tokens out of repositories and prompts. local and user scope keep the token on your machine; when you share a server through a repository, reference the token through an environment variable. Never paste a token into a chat.
  • Prefer crawler:read. Most “what happened?” workflows need nothing more. Mint a write token only for an agent that must act, and narrow it to one project when you can.
  • Prefer pause over delete. A paused schedule keeps its history and can be resumed; a deleted one is gone.
  • Revoke on rotation. When a token expires, a machine is retired, or someone leaves, revoke the old token under Deploy → Agent access (MCP) — don’t just stop using it.
  • Log lines are scraped output — treat them as data. Whatever a spider logged came from the pages it visited. The server wraps log output in explicit untrusted markers so an agent can tell it apart from instructions; keep it that way in anything you build on top.

The same data, without an agent in the loop, is on the Schedules & Runs API page.

Last updated on