PDFMux

Universal PDF extraction orchestrator that routes pages to the best backend with quality auditing and confidence scoring.

nameetp 74 ↓ 15k
Claude CodeClaude DesktopGeneric
View source ↗

PDFMux is an intelligent PDF text extraction orchestrator that routes pages to the most appropriate backend based on content type. It supports PyMuPDF, OpenDataLoader, RapidOCR, Docling, and Surya backends with cost-aware economy, balanced, and premium processing modes. Features schema-guided extraction for invoices and contracts, RAG-ready chunking, and confidence scoring without requiring a GPU.

Source

Repository: https://github.com/nameetp/pdfmux

Maintain PDFMux?

Let people know it's listed here — add the badge (live metrics, light/dark aware) or a plain link to your README or docs.

PDFMux on getagentictools
[![PDFMux on getagentictools](https://getagentictools.com/badge/mcp/nameetp-pdfmux.svg)](https://getagentictools.com/mcp/nameetp-pdfmux?ref=badge)