Skip to content
Alternatives
Software like Web crawling and data extraction
What else does this job. Matched on what each project does, not on who links to whom.
Closest first
- Crawloracrawlora.netCrawlora is a developer platform offering web scraping APIs and anti-bot solutions for extracting data from search engines, maps, commerce, and finance sites. It supports over 500 endpoints and integrates with AI agents and MCP workflows, enabling reliable data access for developers and businesses.
- Website Crawlerwebsitecrawler.orgWebsite Crawler is a SaaS platform designed for real-time crawling, analysis, and data extraction from websites, with a focus on identifying technical issues that are difficult to detect manually. It provides an online interface for auditing and monitoring websites, making it suitable for users seeking to optimize site performance, search presence, and technical health. The platform offers visual analysis through real-time pie charts, covering metrics such as loading times, internal and external links, HTTP status codes, and content breakdowns. Users can compare current and historical analysis with the crawl history utility, which logs all crawl jobs and their timestamps. Website Crawler generates detailed technical reports that help users address issues like internal redirects, dead or unused CSS/JavaScript, duplicate content, broken links, canonical link problems, missing heading tags, duplicate titles and meta tags, and missing image alt tags. It also provides features for identifying thin content, generating XML sitemaps with customizable options, and auditing security headers and SSL certificate status, including email alerts for impending certificate expiry. Data extraction is a core feature, allowing users to configure custom tag settings using CSS selectors or XPath to scrape specific information from web pages. Extracted data can be downloaded in CSV, JSON, Markdown, PDF, or spreadsheet formats, and the platform can produce LLM-ready JSON structured data. An API is available for retrieving data in JSON format for integration with other software projects. Users can adjust crawl behavior, such as crawler type, URLs processed per minute, and delays, as well as track usage credits and manage crawl history. The platform supports crawling JavaScript-heavy single-page applications by executing JavaScript to capture dynamic content. Website Crawler includes options to schedule automated crawls and send email notifications upon completion. It also allows users to add branding to PDF reports, configure email alerts, and manage blacklist or whitelist directives. The service offers a free plan supporting up to 10 URLs without requiring a credit card, with paid plans scaling up to 500,000 pages. Users can upgrade or cancel at any time, and all features are managed through an online dashboard.
- webclawwebclaw.ioWebclaw is a web scraping API for LLMs and AI agents. It turns websites into markdown, JSON, or structured data from a URL, with a stated focus on clean output and fewer tokens than raw HTML. The page also describes it as a drop-in Firecrawl replacement and says it works without a headless browser. Its listed surfaces include a cloud API with REST endpoints for scrape, crawl, and search, plus an MCP Server and a CLI tool. Webclaw says the same engine drives the API, CLI, and MCP server, and its features page mentions that every endpoint has its own page. The product text also says it handles bot protection and returns static pages in around 118 ms. The service is described for use cases including RAG, agents, research, and monitoring. It is mentioned alongside integrations with LangChain, Cursor, and n8n, and it says it can be plugged into Claude, Cursor, and agents through the MCP Server. The site also advertises live scraping with no signup and three free runs a day. Pricing-related text on the page includes free credits for open-source builders, a free trial prompt, and a statement to save 20% on any yearly plan.
- Crawleocrawleo.devCrawleo is a web API for AI applications. It is described as an answer engine for AI applications that can search the web, crawl URLs, and reach pages that other scrapers cannot, while returning clean Markdown that is ready for model use. The service exposes five endpoints: Bing Search, Google Search, Google Maps, Crawler, and Headful Browser. Bing Search includes crawling and returns page content as Markdown rather than only snippets. Google Search returns full SERP data, including organic results, knowledge graphs, People Also Ask, news, images, and shopping. Google Maps returns business listings, places, and landmarks with addresses, coordinates, ratings, and phone numbers, and is described as being built for local search and lead generation. The Crawler can crawl any URL with optional JavaScript rendering and return HTML, Markdown, or plain text. Headful Browser runs real Chromium with SOAX residential proxies, can get past Cloudflare, Akamai, and DataDome, and can return Markdown or HTML plus an optional screenshot. Crawleo is intended to connect with AI tools and editors already in use. The page names Claude, Cursor, Windsurf, GitHub Copilot, OpenAI, DeepSeek, Hugging Face, Ollama, LangChain, OpenClaw, n8n, Make, and Zapier. It ships a native MCP server, provides an n8n node and an OpenClaw skill, and can also be used as a plain REST API from any HTTPS-capable client. Pricing is monthly and credit-based, with Free, Developer, Pro, and Scale plans. Every account gets 500 free credits each month with no card required. The page also states that Crawleo does not train AI on user data and does not sell it. The company name shown is Crawleo LLC.
- ScrapeUpscrapeup.comScrapeUp is a web scraping API that leverages AI to extract structured data from any website or PDF. It manages proxies, browsers, and CAPTCHAs, allowing developers to describe their data needs in plain English and receive JSON output. The platform is designed for use cases like SEO monitoring, price tracking, and lead generation.
- WebScrapingAPIwebscrapingapi.comWebScrapingAPI provides hosted APIs and managed services for extracting web data from pages, search engines, Amazon, e-commerce sites, job boards, real estate portals, and other sources. It also offers residential proxy infrastructure for businesses that need scalable web data collection.
- CrawlEyesgithub.comCrawlEyes is an open-source toolkit for scraping websites, extracting full text, searching content, and applying semantic reranking. It provides an MCP server and command-line installation path for developers building AI agents and research workflows.
- berrycrawlberrycrawl.comBerrycrawl is a web data API designed for AI agents. It converts live web content into structured, agent-ready data through focused endpoints that handle scraping, crawling, searching, screenshot capture, document parsing, and data extraction. The service addresses the challenges of turning raw web pages into usable information without the overhead of managing browsers, proxies, or custom parsers. Users can point to a website to retrieve its complete public brand profile, including logos, colors, fonts, and social links. Separate endpoints support scraping page content, capturing clean screenshots that remove cookie banners, overlays, and chat widgets while loading lazy content and stitching tall pages, and parsing documents. Site mapping and crawling features discover URLs from sitemaps and links, then execute bounded crawls with controls for depth, concurrency, deduplication, and webhooks. Web search capabilities use advanced operators to locate current sources and return relevant page content within the same workflow. Data extraction accepts a JSON schema or description to produce structured JSON output from single pages or batches of URLs, with support for asynchronous jobs. It is delivered as a unified API with endpoints such as POST /api/v1/brand. The platform emphasizes one focused service for live web content, discovery, extraction, and company intelligence. New users receive 100 free credits with no card required.
- Botscraperbotscraper.comBotscraper provides managed web scraping and data extraction services for businesses. It supports price monitoring, competitive intelligence, lead generation, SERP scraping, ecommerce data, and financial market data collection.
Ranked by how close each one sits to Web crawling and data extraction in the index, not by popularity. Back to Web crawling and data extraction →