AI agent evaluation toolkit for Copilot Studio. Plan evals, generate test cases, interpret results, and triage failures — grounded in Microsoft's Eval Scenario Library and Triage & Improvement Playbook.
★not rated 127 2mo agoA
tokens not measured
originalMIT
Extend Pydantic AI agents with batteries-included capabilities from pydantic-ai-harness, currently Code Mode for collapsing many tool calls into one sandboxed Python execution.
★not rated 127 4d agoA
tokens not measured
originalMIT
WAVES (Workers, Aggregate, Verify, Extend): wave-based orchestration for Codex. Decompose into independent slices, spawn parallel Codex subagents (or spawnagentsoncsv) as a bounded wave, verify evidence-backed handoffs, then synthesize. The manifest is the stop function: stop only on completion, stagnation, or budge.
★not rated 127▲
+1 1mo agoA
tokens not measured
originalApache-2.0
Carves safety-adjacent requests into a safe scope, executed by a genuinely isolated fresh-context subagent, with an adversarial verify pass and a descope ledger.
★not rated 125 22d agoA
tokens not measured
originalMIT
Git-native note graph for AI agent collaboration. Injects mycelium notes into context when files are read, teaches agents the protocol via SKILL.md, and nudges agents to leave notes after edits.
★not rated 124 4mo agoA
tokens not measured
originalMIT
File a task with nohuman and check on it without leaving Claude Code: taskadd and taskstatus over your local nohuman server (127.0.0.1:8420). nohuman plans, codes, tests and has the work reviewed by a second model, then opens the pull request and stops.
★not rated 122 4d agoA
tokens not measured
originalMIT
Test whether your Claude Code skills actually trigger, and benchmark any coding agent — author, run, and analyze the eval suite that proves it, locally or as a CI gate.
★not rated 119 5d agoA
tokens not measured
originalApache-2.0
Plugin de ingeniería de software: 8 agentes de núcleo, Selina si hay frontend y Lucius bajo demanda. Memoria persistente por proyecto, quality gates y flujos desde la idea hasta la entrega.
★not rated 119 19d agoA
tokens not measured
originalMIT
Agent Skills for Sonilo's licensed music, sound-effects, dubbing, and audio-ducking API — with structured-brief prompting guidance (pre-flight, music briefs, SFX action maps) folded into each skill.
★not rated 113 9d agoA
tokens not measured
originalMIT
Universal AI agents collection for Claude Code - 28+ specialized agents with strict boundary enforcement, project-specific agent creation, and intelligent task routing.
★not rated 111 7mo agoA
tokens not measured
originalMIT
PRD operating system: capture rough ideas, draft reviewable PRDs, run standard and adversarial review (Codex or a Claude senior-staff-engineer subagent, recorded truthfully), triage findings, decompose approved PRDs into atomic issue specs, and execute those issues with scope enforcement, stop-gate receipts, and…
★not rated 108
changed 2d agoA
tokens not measured
originalMIT
At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: