Skill Claude CodeCodex
Set up the experiment contract and run the frozen baseline benchmark. Use when an improve-repository project needs its evaluation defined, the contract approved, or the baseline measured.
Skill Claude CodeCodex
Set up the experiment contract and run the frozen baseline benchmark. Use when an improve-repository project needs its evaluation defined, the contract approved, or the baseline measured.
Skill Claude CodeCodex
Check that ResearchForge's dependencies (git, Python, optionally Docker) are available and explain any failures. Use when setup fails, before starting a project, or when the user asks whether their machine is ready.
Skill Claude CodeCodex
Generate testable, evidence-linked hypotheses from the research landscape and import them for validation. Use after the landscape exists, when the user wants hypotheses, experiment ideas, or "what should we try?".
Skill Claude CodeCodex
Synthesize stored papers into a research landscape — grouped directions with evidence claims — and import it for validation. Use after papers are stored, when the user wants directions, themes, or a map of the literature.
Skill Claude CodeCodex
Build the research package — BibTeX citations, related work, evidence matrix, paper outline, reproducibility bundle, and experiment data. Use when the user wants publication materials or a research write-up bundle.
Skill Claude CodeCodex
Search arXiv for papers relevant to the project objective and review what was stored. Use when the user wants literature, related work, or asks what papers ResearchForge found.
Skill Claude CodeCodex
Design experiment variants for a hypothesis — write patches, import the plan for validation, and get it approved. Use after a baseline exists, when the user wants to plan or implement experiments.
Skill Claude CodeCodex
Summarize a run's results — ranking, Pareto trade-offs, constraint violations, and rejected experiments — grounded strictly in recorded measurements. Use when the user asks how the experiments went or which variant won.
Skill Claude CodeCodex
Execute an approved experiment plan through the screening → full benchmark funnel, or resume an interrupted run. Use when the user says run the experiments, or a run was interrupted.
Skill Claude CodeCodex
Ship a validated experiment — clean local branch reconstructed from the baseline, engineering report, and optional draft PR. Use when the user wants the winning change as a branch, a report, or a PR.
Skill Claude CodeCodex
Start or resume a ResearchForge project — explore a research idea or improve a repository with benchmarked experiments. Use when the user wants to begin research, set up ResearchForge, or asks "where was I?" in an existing project.
Skill Claude CodeCodex
Run repeated validation benchmarks on a run's finalists so a result can honestly be called validated. Use after a run has a promising winner, or when the user asks to confirm/validate a result.
Cursor rule
Set up the experiment contract and run the frozen baseline benchmark. Use when an improve-repository project needs its evaluation defined, the contract approved, or the baseline measured.
Cursor rule
Check that ResearchForge's dependencies (git, Python, optionally Docker) are available and explain any failures. Use when setup fails, before starting a project, or when the user asks whether their machine is ready.
Cursor rule
Generate testable, evidence-linked hypotheses from the research landscape and import them for validation. Use after the landscape exists, when the user wants hypotheses, experiment ideas, or "what should we try?".
Cursor rule
Synthesize stored papers into a research landscape — grouped directions with evidence claims — and import it for validation. Use after papers are stored, when the user wants directions, themes, or a map of the literature.
Cursor rule
Build the research package — BibTeX citations, related work, evidence matrix, paper outline, reproducibility bundle, and experiment data. Use when the user wants publication materials or a research write-up bundle.
Cursor rule
Search arXiv for papers relevant to the project objective and review what was stored. Use when the user wants literature, related work, or asks what papers ResearchForge found.
Cursor rule
Design experiment variants for a hypothesis — write patches, import the plan for validation, and get it approved. Use after a baseline exists, when the user wants to plan or implement experiments.
Cursor rule
Summarize a run's results — ranking, Pareto trade-offs, constraint violations, and rejected experiments — grounded strictly in recorded measurements. Use when the user asks how the experiments went or which variant won.
Cursor rule
Execute an approved experiment plan through the screening → full benchmark funnel, or resume an interrupted run. Use when the user says run the experiments, or a run was interrupted.
Cursor rule
Ship a validated experiment — clean local branch reconstructed from the baseline, engineering report, and optional draft PR. Use when the user wants the winning change as a branch, a report, or a PR.
Cursor rule
Start or resume a ResearchForge project — explore a research idea or improve a repository with benchmarked experiments. Use when the user wants to begin research, set up ResearchForge, or asks "where was I?" in an existing project.
Cursor rule
Run repeated validation benchmarks on a run's finalists so a result can honestly be called validated. Use after a run has a promising winner, or when the user asks to confirm/validate a result.