fastCRW — Web Data Toolkit for AI Agents
SkillWeb & browsinginstall sets up the CRW skill and MCP server for all detected AI agents (Claude Code, Cursor, Gemini CLI, Codex, OpenCode, Windsurf, Roo Code).
Available today. Use it from your connected AI after setup.
No other account needed.
Connect ahel once, and every AI you use reads what you have installed.
Then ask your AI: use the fastCRW — Web Data Toolkit for AI Agents skill
About this capability
Scrape, crawl, map, and search the web using fastCRW's native /v1 API. Use when the user needs web page content, site-wide extraction, URL discovery, or web search results. Single binary, 14 MB RAM; /v2 exists separately for Firecrawl migration.
What this skill tells your AI
The instructions your AI receives, as published by us/crw in docs/agent-onboarding/SKILL.md and read by ahel’s review.
When to use this skill
Use this skill when:
- The user asks you to read, scrape, or fetch a web page
- You need to extract content from a URL for context or research
- The user wants to crawl an entire website or discover its pages
- You need to search the web and get page content (cloud mode)
- The user mentions Firecrawl — use native
/v1for new CRW work; use/v2only when migrating existing Firecrawl v2 SDK code
Installation
npx crw-mcp@latest install # installs the skill + MCP server into detected agents
npx crw-mcp@latest init # skill only
install sets up the CRW skill and MCP server for all detected AI agents (Claude Code, Cursor, Gemini CLI, Codex, OpenCode, Windsurf, Roo Code).
Authentication
- Embedded mode (default): No key needed — the MCP server runs a self-contained scraper in ~14 MB RAM. No server required.
- Cloud mode (fastcrw.com): Set
CRW_API_KEY=crw_live_...andCRW_API_URL=https://api.fastcrw.com. Get a free key at https://fastcrw.com with 1000 one-time lifetime credits (never resets, not monthly).
MCP Tools
Output bounds: By default, content is truncated to ~15 000 chars (
crw_scrape,crw_check_crawl_status,crw_parse_file, and any page content inlined intocrw_searchresults viascrapeOptions) andcrw_mapreturns ≤ 100 URLs and ≤ 100 sitemaps. Truncated results carry atruncated: truemarker (crw_mapalso addstotalDiscoveredfor links andtotalSitemapsfor sitemaps). PassmaxLength: 0orlimit: 0to opt out of bounding.crw_searchdoes not advertisemaxLength, so an agent will not discover it, but a hand-written client may still pass it.
crw_scrape
Scrape a single URL and return clean content.
Parameters:
url(required) — The URL to scrapeformats— Output formats:markdown(default),html,links,imagesonlyMainContent— Strip navs/footers/sidebars. Default:trueincludeTags— Only include content matching these CSS selectors (e.g.["article", "main"])excludeTags— Exclude content matching these CSS selectors (e.g.["nav", "footer"])renderJs— Force JavaScript rendering. Default: auto-detect (null)waitFor— Milliseconds to wait after page load before capturingrenderer— Renderer override (e.g."chrome")maxLength— Truncate output to this many chars.0= unbounded. Default: ~15 000
crw_crawl
Start an async BFS crawl from a URL. Returns a job ID — poll with crw_check_crawl_status.
Parameters:
url(required) — Starting URLmaxDepth— Maximum link depth. Default:2maxPages— Maximum pages to crawljsonSchema— JSON schema for structured extraction per pagerenderJs— Force JavaScript renderingwaitFor— Milliseconds to wait after page load before capturingrenderer— Renderer override
Returns: { "id": "job-uuid" } — use this ID with crw_check_crawl_status.
crw_check_crawl_status
Poll an async crawl job for results.
Parameters:
id(required) — The crawl job ID fromcrw_crawlmaxLength— Truncate each page's content fields to this many chars.0= unbounded. Default: ~15 000
Returns: { "status": "scraping|completed|failed", "data": [...] }
Browser Automation: Full interactive browser control (JavaScript rendering, click, fill, etc.) requires the separate crw-browse MCP server binary (
command: crw-browse). It exposes its own tools (goto,tree, and others) and is not part of this MCP server. Do not callcrw_browsehere — it is not a tool in crw-mcp and will return a JSON-RPC -32602 "Unknown tool" error.
crw_search
Search the web for current information, news, facts, or docs. Use whenever the answer may depend on up-to-date or external information. Returns ranked results (url/title/description/snippet); optionally scrape each result inline. Always available in proxy/cloud mode; in embedded mode only when a search backend is configured.
Parameters:
query(required) — The search querylimit— Maximum number of results to return. Default:5lang— Language code for results (e.g."en","tr")tbs— Time filter:qdr:h|qdr:d|qdr:w|qdr:m|qdr:y(past hour/day/week/month/year)sources— If set, group results by source:web,news,imagescategories— Category bias; e.g."pdf","github","research","news","images"scrapeOptions— Options for scraping each result page (e.g.{"formats": ["markdown"]})
crw_map
Discover all URLs on a website via sitemap + link extraction, without scraping content.
Parameters:
url(required) — The URL to mapmaxDepth— Discovery depth. Default:2useSitemap— Check sitemap.xml. Default:truecrawlFallback— Supplement sitemap discovery with a short BFS crawl. Default:true(false= sitemap-only)limit— Maximum URLs to return.0= unbounded. Default:100
Returns: { "links": ["url1", "url2", ...] }
crw_extract
Extract structured JSON from one or more URLs via a prompt and/or JSON schema. Async job, poll with crw_check_extract_status. Needs an LLM.
Parameters:
urls(required) — URLs to extract fromprompt— Free-text extraction objective (required unlessschemais given)schema— JSON schema constraining the extracted outputllmApiKey— BYOK LLM API keyllmProvider— LLM provider (used withllmApiKey)llmModel— LLM model (used withllmApiKey)
Returns: { "success": true, "id": "job-uuid", "status": "processing", "urls": 1 } — use the ID with crw_check_extract_status.
crw_check_extract_status
Poll an async extract job for results.
Parameters:
id(required) — The extract job ID fromcrw_extract
Returns: status and, when complete, a per-URL results array.
crw_cancel_extract
Idempotently request cancellation of an extract job. A claimed URL may finish
while status is cancelling; terminal cancelled preserves that result and
marks every untouched ordered slot cancelled.
Parameters:
id(required) — The extract job ID fromcrw_extract
Returns the same canonical status envelope as crw_check_extract_status.
crw_parse_file
Parse a local file (PDF) into markdown or structured output without fetching from the web.
Parameters:
contentBase64(required) — Base64-encoded file contentsfilename— Original filename (optional, e.g."report.pdf")formats— Output formats:markdown(default),plainText,links,images,json,summary(json/summary need a server LLM)jsonSchema— JSON schema for LLM extraction (whenformatsincludesjson)parsers— Document parsers to apply. Default:["pdf"]maxLength— Truncate output to this many chars.0= unbounded. Default: ~15 000
Common Patterns
Scrape a page for context:
crw_scrape(url="https://example.com", formats=["markdown"])
Crawl docs for RAG: First discover URLs, then crawl:
crw_map(url="https://docs.example.com") → get URL list
crw_crawl(url="https://docs.example.com", maxPages=50) → extract all content
crw_check_crawl_status(id="...") → poll until completed
Search the web:
crw_search(query="your search query", limit=5)
Search from the CLI (one-shot LLM-ready output):
When the crw binary is available, prefer the native field projection
over piping through jq — it's one call instead of two:
crw search "renewable energy 2024" --json --fields title,url,snippet --limit 3
Available fields: title, url, description, snippet, position,
score, category. --json is shorthand for --format json.
Common Edge Cases
- JavaScript-heavy sites: Set
renderJs: trueif the page is blank or returns a loading skeleton - Rate limiting: Cloud mode has per-plan rate limits. Check response headers for
X-RateLimit-* - Large crawls: Use
crw_mapfirst to estimate site size before committing to a largecrw_crawl - Timeout: Crawl jobs expire after 1 hour. Poll
crw_check_crawl_statusregularly
Links
- Cloud API: https://fastcrw.com — 1000 one-time lifetime free credits (never resets, not monthly)
- Docs: https://docs.fastcrw.com
- GitHub: https://github.com/us/crw
- Native API:
/v1/scrape,/v1/crawl,/v1/map, and/v1/searchare the recommended routes for new CRW integrations - Firecrawl migration:
/v2/*is a compatibility layer, not the default API for new builds
Signals
- GitHub stars
- 970
- Forks
- 71
- Last commit
- Sep 2026
Others that do the same job
Advanced
- Catalog kind
- skill
- Gateway key
crw- Source
- github.com/us/crw