research toolkit skills

18 tagged research toolkit, measured the same way as everything else here.

Browse within: ai-coding-agents 18

vp-baseline-compare

01

khamidov17/vpstack

Skill Claude CodeCodex

Run B1 (McAdams) + B2 (neural) baselines on the same eval set as the user's anonymization system, return a delta table (EER / WER / linkability). Use when the user asks "how does my system compare to baseline?" or "is this better than B1?" or wants a quick sanity check of their anonymizer against canonical references.…

2 4mo ago A 109 tokens

vp-plan-eng-review

02

khamidov17/vpstack

Skill Claude CodeCodex

Voice-privacy engineering review. Layers on top of the generic /plan-eng-review with 18 VP-specific quality gates: GPLv3 isolation, runtime model fetch, held-out test split safety, the 5 reproducibility checks, three attacker conditions, VP2026 submission format (CSV + Mixed gender + rank/zip), recipe interface…

2 4mo ago A 157 tokens

vp-talk

03

khamidov17/vpstack

Skill Claude CodeCodex

Two-mode planning skill for voice anonymization work. Two-mode planning skill. Like gstack /office-hours — asks forcing questions, then writes a locked plan that downstream skills use. Mode R (Research): VP2026 benchmark — 8 forcing questions on open question, threat model, contribution claim, baseline, eval scope…

2 4mo ago A 218 tokens