crawlforge-mcp
CrawlForge MCP is a production-ready MCP server that gives AI agents the power to scrape websites, extract structured data, run deep research, bypass anti-bot…
Community: Submitted by a user or imported; check the owner before granting accessOnlineNo sign-inGlobalFreeRead-only
What it can do
- Fetch Url: Fetch content from a URL with optional headers and timeout
- Extract Text: Extract clean text content from a webpage
- Extract Links: Extract all links from a webpage with optional filtering
What data it sees
Do you need an account
No: the server works without sign-in
CrawlForge MCP is a production-ready MCP server that gives AI agents the power to scrape websites, extract structured data, run deep research, bypass anti-bot detection, and process documents. It packages 27 specialized web scraping tools into a single MCP server with credit-based pricing and a free tier.
Server tool list (20)
Raw names from tools/list. Only developers need these.
| fetch_url | Fetch content from a URL with optional headers and timeout |
| extract_text | Extract clean text content from a webpage |
| extract_links | Extract all links from a webpage with optional filtering |
| extract_metadata | Extract metadata from a webpage (title, description, keywords, etc.) |
| scrape_structured | Extract structured data from a webpage using CSS selectors |
| search_web | Search the web using Google Search API (proxied through CrawlForge) |
| crawl_deep | Crawl websites deeply using breadth-first search |
| map_site | Discover and map website structure |
| extract_content | Extract and analyze main content from web pages with enhanced readability detection |
| process_document | Process documents from multiple sources and formats including PDFs and web pages |
| summarize_content | Generate intelligent summaries of text content with configurable options |
| analyze_content | Perform comprehensive content analysis including language detection and topic extraction |
| extract_structured | Extract structured data from a webpage using LLM-powered analysis and a JSON Schema. Falls back to CSS selector extraction when no LLM provider is configured. |
| batch_scrape | Process multiple URLs simultaneously with support for async job management and webhook notifications |
| scrape_with_actions | Execute browser action chains before scraping content, with form auto-fill and intermediate state capture |
| deep_research | Conduct comprehensive multi-stage research with intelligent query expansion, source verification, and conflict detection |
| track_changes | Enhanced content change tracking with baseline capture, comparison, scheduled monitoring, advanced comparison engine, alert system, and historical analysis |
| generate_llms_txt | Analyze websites and generate standard-compliant LLMs.txt and LLMs-full.txt files defining AI model interaction guidelines |
| stealth_mode | Advanced anti-detection browser management with stealth features, fingerprint randomization, and human behavior simulation |
| localization | Multi-language and geo-location management with country-specific settings, browser locale emulation, timezone spoofing, and geo-blocked content handling |