Evaluation Agentic Loops

1,399 loops for evaluation — ranked by real GitHub stars, refreshed continuously, never faked.

# Loop Stars
1 Ship And Babysit tinyhumansai/openhuman Commit, push to origin (fork), open PR to tinyhumansai/openhuman:main, then poll every ~5min for CodeRabbit comments and CI failu… 34k ★ 2 Ham eyaltoledano/claude-task-master This command initiates the HAM (Hamster Automated Management) workflow for task execution. 28k ★ 3 Fix Ci facebook/relay You are a specialized skill for checking GitHub CI status and fixing failing tests in the Relay project. 19k ★ 4 Hunt Github Issues compiler-explorer/compiler-explorer I need you to help me systematically hunt and fix bugs reported in GitHub. 19k ★ 5 Laputa Done refactoringhq/tolaria Mark a Laputa task as done: add completion comment, move to In Review, then self-dispatch the next task. 18k ★ 6 Yeet polarsource/polar Lint, type-check, and create a PR. 10k ★ 7 Commit Update Pr samuelclay/NewsBlur Commit, push, create or update a pull request, and watch CI until green 7.5k ★ 8 Dev Cycle entireio/cli Orchestrates the development cycle - developer implements, reviewer critiques, repeat until done 4.6k ★ 9 Loop nyldn/claude-octopus [advanced] Execute tasks in loops with conditions, iterative improvements until goals are met 3.7k ★ 10 Add Model Support TransformerLensOrg/TransformerLens Guided workflow for adding a new architecture adapter to TransformerBridge. 3.6k ★ 11 L Debug Trace lowdefy/lowdefy Analytically debug issues by tracing data flow through code - no logs, no test-and-see 3k ★ 12 Ralph dataplat/dbatools Generate a Ralph Wiggum-style iterative automation for large tasks. The Ralph Wiggum technique runs an AI CLI in a stateless loop… 2.8k ★ 13 Pr Light-Heart-Labs/ODS Create a new branch, commit changes, push, create PR, and fix all failed checks (except claude review) 2.5k ★ 14 Pr Check Light-Heart-Labs/ODS Run all local CI checks before creating a PR 2.5k ★ 15 Ralph Tasks BrighterCommand/Brighter Create ralph-tasks.md for unattended TDD implementation 2.4k ★ 16 The artifact-to-skill loop A reusable workflow for turning one proven artifact into a transferable skill, playbook, or procedure and validating it on a seco… 2.2k ★ 17 The Axelrod subagent arena loop A controlled tournament where two reasoning AI agents repeatedly choose to cooperate or defect, then are compared with players th… 2.2k ★ 18 The cross-run playbook loop A versioned AI-agent learning workflow that tests one recorded lesson at a time, keeps evidence across runs, and removes guidance… 2.2k ★ 19 The devil's-advocate loop A critic-and-builder workflow that attacks a design, tracks every objection, and requires evidence before an objection can be clo… 2.2k ★ 20 The easy onboarding loop A first-time-user test that starts with no saved account or browser state, fixes one confirmed onboarding obstacle, and retries t… 2.2k ★ 21 The epistemic frontier loop A bounded reasoning loop that separates facts from assumptions, tests falsifiable hypotheses, updates confidence, and selects the… 2.2k ★ 22 The full product evaluation loop A comprehensive product-quality workflow that evaluates realistic scenarios across every major capability, fixes weak outcomes, a… 2.2k ★ 23 The literature-search verification loop A bounded literature-search workflow that deduplicates papers across live sources, verifies DOI metadata, measures topical releva… 2.2k ★ 24 The loop-auditor loop A read-only portfolio audit that recomputes measured-loop performance, evaluates operational loops on their own terms, and recomm… 2.2k ★ 25 The multi-LLM convergence loop Alternates two AI systems from different providers to review a plan, document, or code change until both approve the exact same v… 2.2k ★ 26 The next-action confidence check A bounded AI-agent exit gate that verifies the current task, evaluates the next action, and returns control to the user before mo… 2.2k ★ 27 The promise-to-proof loop A product review that compares claims in marketing, documentation, demos, and AI answers with current evidence, then fixes or nar… 2.2k ★ 28 The quality streak loop A realistic product-testing workflow that turns every failure into documented regression coverage and restarts the success streak… 2.2k ★ 29 The Revolve versioned-experiment loop A Revolve workflow that improves prompts, code, or configurations through checkpointed experiments whose scores remain comparable… 2.2k ★ 30 The self-improving champion loop A prompt-optimization workflow that tests challengers on a working set, promotes only fresh holdout wins, and keeps the current c… 2.2k ★ 31 The Strip Miner loop An evidence-driven workflow-mining loop that finds repeated successes in authorized coding-agent history, rejects contradicted ca… 2.2k ★ 32 New Source mampfes/hacs_waste_collection_schedule Guide a contributor through adding support for a new waste-collection provider, from "I want my council added" to "PR submitted t… 2.1k ★ 33 Finalize amd/gaia Finalizes implementation by rebasing onto main, running the GAIA claude.yml code review (Opus), fixing issues, linting, and loopi… 1.5k ★ 34 Build Skill qdhenry/Claude-Command-Suite Create comprehensive Claude Code Skills through elicitation-driven development 1.3k ★ 35 Explain Code qdhenry/Claude-Command-Suite Analyze and explain code functionality 1.3k ★ 36 Parallel Feature Build qdhenry/Claude-Command-Suite Orchestrated parallel implementation of complex features using multiple agents, with dependency-aware batching and synchronized p… 1.3k ★ 37 Upgrade Llvm clice-io/clice Upgrade LLVM to a new version. Accepts the target version as argument (e.g., 22.1.4). 1.3k ★ 38 Issue Batch sceneview/sceneview Launch-and-go continuous issue-processing cycle — a replace-on-completion pipeline of 6-8 lean-clone background agents, fire-and-… 1.2k ★ 39 Recover Failed Ingest SemiAnalysisAI/InferenceX Recover a failed main-branch sweep ingest through the normal artifact-reuse path without rerunning GPU benchmarks 1.2k ★ 40 Dart Manage Pr dartsim/dart manage an open DART pull request through CI, review, and cleanup 1.2k ★ 41 Dart Review Pr dartsim/dart review a PR or address review feedback 1.2k ★ 42 Gh Issue Troubleshoot dotCMS/core Fix a dotCMS GitHub issue end-to-end — fetches the issue, researches the codebase, proposes a concrete code fix with before/after… 948 ★ 43 Apply DanSnow/vue-recaptcha Implement tasks from an OpenSpec change 896 ★ 44 X Audit tksuoran/erhe Conduct a Software Architecture, Foundations & Security Audit of a C++ project (KISS, DRY, YAGNI, Overengineering, Foundations, S… 883 ★ 45 X Debug tksuoran/erhe Solve any problem with Protocol D (RIDHV) -- deterministic root-cause analysis, not guess-and-check. A universal method, tuned he… 883 ★ 46 X Develop tksuoran/erhe Implement a C++ TDDAB plan, block by block, as a senior C++ developer -- RED (failing Catch2 test) -> GREEN (minimum code) -> VER… 883 ★ 47 X Review Plan tksuoran/erhe Review a C++ TDDAB plan for methodology compliance, dependency ordering, self-sufficiency, and conformity to the C++ foundations… 883 ★ 48 X Tddab tksuoran/erhe Plan a C++ feature with the TDDAB methodology (Test-Driven Development As Blocks) -- produce a block-by-block plan where each blo… 883 ★ 49 E2e Test zmeyer44/Locker Write and run a full Playwright E2E test suite for the current feature branch 765 ★ 50 SKILL ReflexioAI/claude-smart Update an existing pull request with new changes. Use when the user wants to update a PR, push follow-up changes to a PR, refresh… 714 ★ 51 Pr markuplint/markuplint Create and push a pull request 605 ★ 52 Address Comments iota-uz/iota-sdk Fetch unresolved PR comments and address them with code changes 445 ★ 53 Tdd iota-uz/iota-sdk Test-driven development: clarify requirements, write tests, implement code. RUN THIS COMMAND IN PLAN MODE ONLY 445 ★ 54 Maestro Grill catlog22/maestro-flow Use when stress-testing a plan, idea, or requirement against codebase reality before brainstorming 428 ★ 55 Maestro Plan catlog22/maestro-flow Use when creating, revising, or verifying an execution plan for a phase or task 428 ★ 56 Maestro Ralph Cli Execute catlog22/maestro-flow Skill execution wrapper for delegate — execute skill, return structured result 428 ★ 57 Maestro Swarm Workflow catlog22/maestro-flow Parallel workflow accelerator — route intent to fixed Workflow scripts for multi-agent concurrent execution 428 ★ 58 Odyssey Planex catlog22/maestro-flow Requirement-driven iterative cycle — plan, execute, strict verify, fix loop until acceptance criteria met 428 ★ 59 Quality Debug catlog22/maestro-flow Use when bugs, test failures, or unexpected behavior need systematic root cause investigation 428 ★ 60 Quality Refactor catlog22/maestro-flow Use when accumulated tech debt needs systematic identification and safe reduction 428 ★ 61 Quality Review catlog22/maestro-flow Use after execution to evaluate code quality across correctness, security, performance, and architecture 428 ★ 62 Quality Test catlog22/maestro-flow Use when implementation needs user acceptance testing with interactive verification and gap closure 428 ★ 63 Improve Compile Rate paiml/depyler | Issue | Action | |-------|--------| | Tests fail | Rollback, try narrower fix | | Regression | Rollback, add guards | | Same er… 356 ★ 64 Mantishack deonmenezes/mantishack One-shot MAXIMAL autonomous pentest — the full scan+validate pipeline PLUS a parallel red-team agent war-game, adversarially veri… 351 ★ 65 Handle Pr Comments dsifry/metaswarm Handle review comments on pull requests with appropriate responses and resolutions. 335 ★ 66 Test And Fix 0xquinto/bcherny-claude Run tests and fix any failures 334 ★ 67 Brainstorm CL-ML/open-collider You are the brainstorm orchestrator for Open Collider. You manage an iterative idea generation loop. 331 ★ 68 Overnight Prs Laer-Smart/2anki.net Autonomous overnight loop — verify issues against the codebase, close stale ones, open review-ready PRs (bug fixes + trio-decided… 323 ★ 69 Create Phase syahiidkamil/Software-Engineer-AI-Agent-Atlas Create a development phase the lean way — resolve ambiguity through Q&A, then capture it as a self-contained low-fidelity wirefra… 305 ★ 70 Create Test Cases syahiidkamil/Software-Engineer-AI-Agent-Atlas Author human-readable manual test cases (markdown) into docs/living-test-cases/ — the living artifact that /qa:manual-test-run ex… 305 ★ 71 Visual syahiidkamil/Software-Engineer-AI-Agent-Atlas Plan any change the lean way — resolve ambiguity through Q&A, optionally explore the codebase, then capture the result as a singl… 305 ★ 72 Cmt n9e/fe When user uses /cmt slash command, commit current git changes using fast mode by default, or strict mode when requested. Fast mod… 297 ★ 73 Fast mrgoonie/human-mcp Analyze and fix the issue [FAST] 289 ★ 74 Test mrgoonie/human-mcp Run test suite and fix issues 289 ★ 75 Dashclaw Quality ucsandman/DashClaw Recurring find-and-fix quality pass over the whole DashClaw app — browser smoke (frontend-verify) + code gates → triage → paralle… 279 ★ 76 Prune Deadcode Layr-Labs/eigenda Systematically identify and remove dead code from a directory or module. 262 ★ 77 Approve BrighterCommand/Darker Approve a specification phase 250 ★ 78 Ralph Implement BrighterCommand/Darker Unattended TDD implementation from ralph-tasks (auto mode + self-driving loop) 250 ★ 79 Tidy First BrighterCommand/Darker Separate structural and behavioral changes following Beck's Tidy First 250 ★ 80 Release Oc shaun0927/openchrome OpenChrome release workflow — issue verification, triage, review, fix own PRs, merge, and optionally publish 222 ★ 81 Sync Deps DataDog/chaos-controller Upgrade Go module dependencies and align shared dependencies with a reference repository. Loop until both conditions are met: zer… 208 ★ 82 Execute pcharbon70/term_ui CRITICAL: You are now executing the implementation. Work through the detailed task breakdown systematically, consulting agents fo… 193 ★ 83 Ralphy Validate wenqingyu/ralphy-openspec You are validating an OpenSpec change. 186 ★ 84 Busycommit wbern/agent-instructions Create multiple atomic git commits, one logical change at a time 168 ★ 85 Create Prd coleam00/ralph-loop-quickstart Create a comprehensive Product Requirements Document (PRD) for a new project with interactive discovery questions 158 ★ 86 Design Ingestor NomaDamas/AutoRAG-Research Design ingestion strategy with mandatory human review. Use after dataset inspection phase. 143 ★ 87 Implement Pipeline NomaDamas/AutoRAG-Research Orchestrate full pipeline implementation workflow from paper to validated code. Use when implementing a new retrieval or generati… 143 ★ 88 Cancel Ralph tobiasosborne/alethfeld | description | allowed-tools | hide-from-slash-command-tool | |---|---|---| | Cancel active Ralph Wiggum loop | Bash(test -f .cl… 142 ★ 89 SENTRY Solve Issue Tdd Comfy-Org/comfy-claude-prompt-library Can you get full context and info on this Sentry issue $ARGUMENTS 135 ★ 90 SOLVEBUG Solve Issue Tdd Comfy-Org/comfy-claude-prompt-library Your task is to fetch and solve the GitHub issue: $ARGUMENTS using test-driven development (TDD). 135 ★ 91 Update Version AztecProtocol/aztec-starter Update the Aztec version across the entire repo, update contract and TS code, and run tests. Usage: /update-version <new-version-… 119 ★ 92 Ralph Init pproenca/agent-tui Initialize a RALPH loop (.ralph/) with SPEC, PROMPT, TODO, and runner script 103 ★ 93 Do jayminwest/kotadb Universal entry point - delegates to appropriate workflow 101 ★ 94 Aliases TheBeardedBearSAS/claude-craft CLI aliases for frequently used commands 97 ★ 95 Ralph Run TheBeardedBearSAS/claude-craft Executer Claude en boucle continue jusqu'a completion de la tache (Ralph Wiggum v2.0) 97 ★ 96 Cherry Pick Fix istio-ecosystem/sail-operator Cherry-pick a failed automated cherry-pick from a bot-created issue 94 ★ 97 Rebase UT-ADL/autoware_mini Rebase current branch onto main and fix merge conflicts 92 ★ 98 Dev Loop goern/forgejo-mcp Spawn an implementer + verifier (+ optional planner) team that iterates on a scoped change until a deterministic check passes. Le… 88 ★ 99 Test And Fix solanabr/solana-ai-kit Run tests and automatically fix common issues 88 ★ 100 Flutter Template Refiner appboypov/pew-pew-plaza-packs Expert in iterative application-to-template refinement for Flutter applications. Use when stripping functionalities from a comple… 82 ★

Showing the top 100 of 1,399 — browse and filter all agentic loops.

The best new evaluation loops, weekly

One email every week with the tools worth your time — plus The Loop Engineering Field Guide free when you join.

Frequently asked questions

What are the best evaluation loops for coding agents?

Ranked by GitHub stars, the current top evaluation loops are: Ship And Babysit (34k ★), Ham (28k ★), Fix Ci (19k ★), Hunt Github Issues (19k ★), Laputa Done (18k ★).

How do I install a evaluation loop?

Loops are workflow patterns, not packages: open the loop's page, copy its prompt or setup, and run it with your agent. Each page includes the full loop definition and source.

How many evaluation loops are listed here?

1,399 evaluation loops are currently tracked, with metrics (stars, downloads, last-update dates) sourced from GitHub and package registries and refreshed continuously.

Related categories