self improving agents agents

12 tagged self improving agents, measured the same way as everything else here.

Browse within: llm-agents 11skill-optimization 11

pr-quality-reviewer

01

Tencent/SkillHone

Agent

PR merge gate for skill-repo PRs. Runs the offline static check, produces a rubric score, posts the verdict as a Forgejo PR comment, and returns APPROVE or REQUESTCHANGES to the dispatching reviewer. Use only from skillhone-evaluation's reviewer flow, never from developer self-check.

137 23d ago A 67 tokens

issue-reporter

02

Tencent/SkillHone

Agent

Analyzes probe evaluation results and creates a single focused Forgejo issue describing the highest-impact failure pattern to fix next.

137 23d ago A 27 tokens

deduper

03

Tencent/SkillHone

Agent

You are the Deduper. You receive validated Q/A candidates from many seeds and produce the final benchmark set by removing structural duplicates and near-collisions. This is where "variety" stops being a per-seed concern and becomes a corpus-level concern.

137 23d ago A 0 tokens

verifier

04

hiphapis/loopcraft

Agent

Independent grader for loopcraft loop-tasks. Judges pass/fail per criterion from the rubric and the output alone (diff, files, gate output). Never receives the maker's reasoning or conversation. Read-only — makes no changes.

3 1mo ago A 49 tokens