AI Vision

Integrates with Google's Gemini and Vertex AI models to analyze images, compare multiple images, and process video content with in

tan-yong-sheng 66 ↓ 4.8k
Claude CodeClaude DesktopGeneric
View source ↗

AI Vision MCP provides AI assistants with powerful image and video analysis capabilities through Google's Gemini and Vertex AI models, supporting both Google AI Studio API keys and Vertex AI service accounts for flexible deployment options. Built by tan-yong-sheng with TypeScript, it offers three core tools: analyze_image for single image analysis, compare_images for multi-image comparison (up to 4 images), and analyze_video for video content analysis, with intelligent file handling that supports URLs, local file paths, and base64 data. The implementation features smart upload strategies that automatically choose between direct API calls for smaller files and cloud storage (Google Cloud Storage for Vertex AI, Files API for Gemini) for larger files, with comprehensive error handling, retry logic, and configurable processing limits, making it valuable for content moderation, visual data analysis, educational applications, and any workflow requiring AI-powered understanding of visual media.

Source

Repository: https://github.com/tan-yong-sheng/ai-vision-mcp

Maintain AI Vision?

Let people know it's listed here — add the badge (live metrics, light/dark aware) or a plain link to your README or docs.

[AI Vision on getagentictools](https://getagentictools.com/mcp/tan-yong-sheng-ai-vision-mcp?ref=badge)