EvalView

Regression testing for AI agents with golden baselines, CI/CD integration, and multi-framework support.

hidai25 123
Claude CodeClaude DesktopGeneric
View source ↗

Provides regression testing capabilities for AI agent workflows including golden baseline comparisons, CI/CD pipeline integration, and support for multiple frameworks like LangGraph, CrewAI, OpenAI, and Claude. Tests can evaluate tool usage, execution sequences, and output quality with optional LLM-as-judge scoring.

Source

Repository: https://github.com/hidai25/eval-view

Maintain EvalView?

Let people know it's listed here — add the badge (live metrics, light/dark aware) or a plain link to your README or docs.

[EvalView on getagentictools](https://getagentictools.com/mcp/hidai25-eval-view?ref=badge)