Web Crawler Data Bridge

Advanced search and retrieval for web crawler data. Supports WARC, wget, Katana, SiteOne, and InterroBot crawlers.

pragmar 44
Claude CodeClaude DesktopGeneric
View source ↗

This MCP server by Ben Caulfield bridges web crawler data with AI language models, supporting five major crawler formats: WARC files, wget archives, InterroBot databases, Katana HTTP text files, and SiteOne captures. Built with Python and featuring full-text search with boolean queries, field-specific filtering by HTTP status and content type, resource pagination, and advanced extras like thumbnail generation for images, markdown conversion, contextual snippets, and XPath extraction. The implementation uses in-memory SQLite databases for fast indexing and search, supports multiple simultaneous crawler configurations, and integrates with Claude Desktop through JSON configuration, making it valuable for AI-powered website analysis, content auditing, SEO optimization, and research workflows that need to query and analyze previously crawled web content.

Source

Repository: https://github.com/pragmar/mcp-server-webcrawl

Maintain Web Crawler Data Bridge?

Let people know it's listed here — add the badge (live metrics, light/dark aware) or a plain link to your README or docs.

[Web Crawler Data Bridge on getagentictools](https://getagentictools.com/mcp/pragmar-mcp-server-webcrawl?ref=badge)