mechanistic interpretability agents

4 tagged mechanistic interpretability, measured the same way as everything else here.

claim

01

zjunlp/Mechanist

Agent

The claim agent of /auto. Runs the /auto-claim skill under two orthogonal axes — BEHAVIORSOURCE (given / given-validation / discovery) sets where the behavior comes from and whether it is validated; MECHANISM (given / discovery) sets who picks the mechanism method. discovery generates ranked, novelty-checked ideas…

51 6d ago A 129 tokens original MIT

experiment

02

zjunlp/Mechanist

Agent

The experiment agent of /auto. Wraps the /auto-experiment skill, which folds mechanism-family routing inline before implementing, code-reviewing, and deploying the experiment suite. Supports two-step invocation — first call returns candidate families for the orchestrator's mini-prompt, second call (with chosenfamily)…

51 6d ago A 72 tokens original MIT

iteration

03

zjunlp/Mechanist

Agent

The iteration agent of /auto. Runs the /auto-iteration-loop skill — an autonomous review loop that consumes /auto-verify's four-state output (PASS / FAIL / INCONCLUSIVE / ZEROELIGIBLEVARIANTS) plus the orthogonal deferred bucket and routes each claim to the right back-edge (① variant-only fix / ② baseline-script fix /…

51 6d ago A 143 tokens original MIT