forger-labs-hq/researchforge

A lab protocol for coding agents. Freeze baselines, run hypotheses in isolated worktrees, reject failures and validate improvements from Claude Code, Cursor or a direct API key (Gemini/Anthropic/openAI).

7Stars on the repository
26Mods indexed here, across every type
3d agoLast push, which is what freshness is scored on
Apache-2.0Licence, which decides whether bodies are shown

forger-labs-hq/researchforge

Skill Claude CodeCodex

Set up the experiment contract and run the frozen baseline benchmark. Use when an improve-repository project needs its evaluation defined, the contract approved, or the baseline measured.

not rated 7 +1 3d ago A 38 tokens original Apache-2.0

forger-labs-hq/researchforge

Skill Claude CodeCodex

Check that ResearchForge's dependencies (git, Python, optionally Docker) are available and explain any failures. Use when setup fails, before starting a project, or when the user asks whether their machine is ready.

not rated 7 +1 3d ago A 47 tokens original Apache-2.0

forger-labs-hq/researchforge

Skill Claude CodeCodex

Generate testable, evidence-linked hypotheses from the research landscape and import them for validation. Use after the landscape exists, when the user wants hypotheses, experiment ideas, or "what should we try?".

not rated 7 +1 3d ago A 46 tokens original Apache-2.0

forger-labs-hq/researchforge

Skill Claude CodeCodex

Synthesize stored papers into a research landscape — grouped directions with evidence claims — and import it for validation. Use after papers are stored, when the user wants directions, themes, or a map of the literature.

not rated 7 +1 3d ago A 47 tokens original Apache-2.0

researchforge-paper

05

forger-labs-hq/researchforge

Skill Claude CodeCodex

Build the research package — BibTeX citations, related work, evidence matrix, paper outline, reproducibility bundle, and experiment data. Use when the user wants publication materials or a research write-up bundle.

not rated 7 +1 3d ago A 45 tokens original Apache-2.0

forger-labs-hq/researchforge

Skill Claude CodeCodex

Search arXiv for papers relevant to the project objective and review what was stored. Use when the user wants literature, related work, or asks what papers ResearchForge found.

not rated 7 +1 3d ago A 40 tokens original Apache-2.0

researchforge-plan

07

forger-labs-hq/researchforge

Skill Claude CodeCodex

Design experiment variants for a hypothesis — write patches, import the plan for validation, and get it approved. Use after a baseline exists, when the user wants to plan or implement experiments.

not rated 7 +1 3d ago A 41 tokens original Apache-2.0

forger-labs-hq/researchforge

Skill Claude CodeCodex

Summarize a run's results — ranking, Pareto trade-offs, constraint violations, and rejected experiments — grounded strictly in recorded measurements. Use when the user asks how the experiments went or which variant won.

not rated 7 +1 3d ago A 46 tokens original Apache-2.0

researchforge-run

09

forger-labs-hq/researchforge

Skill Claude CodeCodex

Execute an approved experiment plan through the screening → full benchmark funnel, or resume an interrupted run. Use when the user says run the experiments, or a run was interrupted.

not rated 7 +1 3d ago A 38 tokens original Apache-2.0

researchforge-ship

10

forger-labs-hq/researchforge

Skill Claude CodeCodex

Ship a validated experiment — clean local branch reconstructed from the baseline, engineering report, and optional draft PR. Use when the user wants the winning change as a branch, a report, or a PR.

not rated 7 +1 3d ago A 45 tokens original Apache-2.0

researchforge-start

11

forger-labs-hq/researchforge

Skill Claude CodeCodex

Start or resume a ResearchForge project — explore a research idea or improve a repository with benchmarked experiments. Use when the user wants to begin research, set up ResearchForge, or asks "where was I?" in an existing project.

not rated 7 +1 3d ago A 51 tokens original Apache-2.0

forger-labs-hq/researchforge

Skill Claude CodeCodex

Run repeated validation benchmarks on a run's finalists so a result can honestly be called validated. Use after a run has a promising winner, or when the user asks to confirm/validate a result.

not rated 7 +1 3d ago A 44 tokens original Apache-2.0

forger-labs-hq/researchforge

Skill Claude CodeCodex

Drive the autonomous research loop round after round — ask the engine which node to expand, plan there, run, read the result, repeat. Use when the user wants to keep improving a repository over many rounds, or wants autorun without an AI API key.

not rated 7 +1 3d ago A 58 tokens original Apache-2.0

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: