Web Crawler Data Bridge
Advanced search and retrieval for web crawler data. Supports WARC, wget, Katana, SiteOne, and InterroBot crawlers.
This MCP server by Ben Caulfield bridges web crawler data with AI language models, supporting five major crawler formats: WARC files, wget archives, InterroBot databases, Katana HTTP text files, and SiteOne captures. Built with Python and featuring full-text search with boolean queries, field-specific filtering by HTTP status and content type, resource pagination, and advanced extras like thumbnail generation for images, markdown conversion, contextual snippets, and XPath extraction. The implementation uses in-memory SQLite databases for fast indexing and search, supports multiple simultaneous crawler configurations, and integrates with Claude Desktop through JSON configuration, making it valuable for AI-powered website analysis, content auditing, SEO optimization, and research workflows that need to query and analyze previously crawled web content.
Source
Repository: https://github.com/pragmar/mcp-server-webcrawl
Maintain Web Crawler Data Bridge?
Let people know it's listed here — add the badge (live metrics, light/dark aware) or a plain link to your README or docs.
[Web Crawler Data Bridge on getagentictools](https://getagentictools.com/mcp/pragmar-mcp-server-webcrawl?ref=badge)