EvalView
Regression testing for AI agents with golden baselines, CI/CD integration, and multi-framework support.
Claude CodeClaude DesktopGeneric
Provides regression testing capabilities for AI agent workflows including golden baseline comparisons, CI/CD pipeline integration, and support for multiple frameworks like LangGraph, CrewAI, OpenAI, and Claude. Tests can evaluate tool usage, execution sequences, and output quality with optional LLM-as-judge scoring.
Source
Repository: https://github.com/hidai25/eval-view
Maintain EvalView?
Let people know it's listed here — add the badge (live metrics, light/dark aware) or a plain link to your README or docs.
[EvalView on getagentictools](https://getagentictools.com/mcp/hidai25-eval-view?ref=badge)