Content Core

Extracts content from diverse media sources including URLs, documents, videos, audio files, and images using intelligent auto-dete

lfnovo 166 ↓ 231k
Claude CodeClaude DesktopGeneric
View source ↗

Content Core MCP Server provides intelligent content extraction from diverse media sources including URLs, documents (PDF, Office files, EPUB), videos, audio files, and images through a unified interface. Built by Luis Novo, it leverages multiple extraction engines with smart auto-detection - using Docling for documents when available, falling back to PyMuPDF, and supporting Firecrawl, Jina, or BeautifulSoup for web content based on API availability. The server handles complex workflows like YouTube transcript extraction, audio/video transcription via OpenAI Whisper, and OCR for images, making it valuable for research automation, content analysis, and building AI agents that need to process mixed media content without manual preprocessing.

Source

Repository: https://github.com/lfnovo/content-core

Maintain Content Core?

Let people know it's listed here — add the badge (live metrics, light/dark aware) or a plain link to your README or docs.

Content Core on getagentictools
[![Content Core on getagentictools](https://getagentictools.com/badge/mcp/lfnovo-content-core.svg)](https://getagentictools.com/mcp/lfnovo-content-core?ref=badge)