Gemini Media Analysis

Provides image, audio, and video analysis tools using Google's Gemini AI for content description, transcription, and understanding

mario-andreschak 12
Claude CodeClaude DesktopGeneric
View source ↗

MCP Video Recognition Server provides tools for image, audio, and video analysis using Google's Gemini AI. Built with TypeScript and the MCP SDK, it offers three main tools: image recognition for describing visual content, audio recognition for transcription and analysis, and video recognition for understanding video content. The server supports both stdio and SSE transport methods, includes file caching to improve performance, and handles the complexities of waiting for video processing to complete. It's particularly useful for applications requiring media content analysis, automated descriptions, or accessibility features without requiring direct integration with Google's APIs.

Source

Repository: https://github.com/mario-andreschak/mcp_video_recognition

Maintain Gemini Media Analysis?

Let people know it's listed here — add the badge (live metrics, light/dark aware) or a plain link to your README or docs.

[Gemini Media Analysis on getagentictools](https://getagentictools.com/mcp/mario-andreschak-mcp-video-recognition?ref=badge)