crawlforge-mcp

CrawlForge MCP is a production-ready MCP server that gives AI agents the power to scrape websites, extract structured data, run deep research, bypass anti-bot…

Community: Submitted by a user or imported; check the owner before granting accessOnlineNo sign-inGlobalFreeRead-only

What it can do

  • Fetch Url: Fetch content from a URL with optional headers and timeout
  • Extract Text: Extract clean text content from a webpage
  • Extract Links: Extract all links from a webpage with optional filtering

What data it sees

Do you need an account

No: the server works without sign-in

CrawlForge MCP is a production-ready MCP server that gives AI agents the power to scrape websites, extract structured data, run deep research, bypass anti-bot detection, and process documents. It packages 27 specialized web scraping tools into a single MCP server with credit-based pricing and a free tier.

Server tool list (20)

Raw names from tools/list. Only developers need these.

fetch_urlFetch content from a URL with optional headers and timeout
extract_textExtract clean text content from a webpage
extract_linksExtract all links from a webpage with optional filtering
extract_metadataExtract metadata from a webpage (title, description, keywords, etc.)
scrape_structuredExtract structured data from a webpage using CSS selectors
search_webSearch the web using Google Search API (proxied through CrawlForge)
crawl_deepCrawl websites deeply using breadth-first search
map_siteDiscover and map website structure
extract_contentExtract and analyze main content from web pages with enhanced readability detection
process_documentProcess documents from multiple sources and formats including PDFs and web pages
summarize_contentGenerate intelligent summaries of text content with configurable options
analyze_contentPerform comprehensive content analysis including language detection and topic extraction
extract_structuredExtract structured data from a webpage using LLM-powered analysis and a JSON Schema. Falls back to CSS selector extraction when no LLM provider is configured.
batch_scrapeProcess multiple URLs simultaneously with support for async job management and webhook notifications
scrape_with_actionsExecute browser action chains before scraping content, with form auto-fill and intermediate state capture
deep_researchConduct comprehensive multi-stage research with intelligent query expansion, source verification, and conflict detection
track_changesEnhanced content change tracking with baseline capture, comparison, scheduled monitoring, advanced comparison engine, alert system, and historical analysis
generate_llms_txtAnalyze websites and generate standard-compliant LLMs.txt and LLMs-full.txt files defining AI model interaction guidelines
stealth_modeAdvanced anti-detection browser management with stealth features, fingerprint randomization, and human behavior simulation
localizationMulti-language and geo-location management with country-specific settings, browser locale emulation, timezone spoofing, and geo-blocked content handling