DINO-X
Empower LLMs with fine-grained visual understanding — detect, localize, and describe anything in images with natural language prom
This DINO-X MCP server by IDEA Research provides AI agents with real-world visual perception capabilities through the DINO-X computer vision API, enabling object detection, localization, human pose estimation, and image captioning. The implementation offers three core detection modes: text-prompted object detection for finding specific items, universal object detection for comprehensive scene analysis, and human pose keypoint detection with 17-point skeletal tracking, all with optional detailed descriptions and visualization capabilities that save annotated images with bounding boxes and labels. Built with TypeScript and the canvas library for image processing, it integrates with DeepDataSpace's hosted DINO-X API and includes robust error handling, task polling for asynchronous processing, and flexible output formatting, making it valuable for AI applications requiring visual understanding, accessibility tools, security monitoring, sports analysis, and automated content moderation workflows.
Source
Repository: https://github.com/idea-research/dino-x-mcp
Maintain DINO-X?
Let people know it's listed here — add the badge (live metrics, light/dark aware) or a plain link to your README or docs.
[DINO-X on getagentictools](https://getagentictools.com/mcp/idea-research-dino-x-mcp?ref=badge)