auto-research

auto-research is a cursor rule for Cursor from Project-VIC-International/Agentic-AI-Development-Course. It costs 1,682 tokens per session, scanned A, original, Apache-2.0.

A rule for improving software through repeated experiments: suggest one testable change, measure it against a baseline, and keep it only if the result is better.

In plain words
What is it for?
Use it to tune performance and document reproducible before-and-after comparisons.
Why use it?
It replaces guesswork with recorded evidence about whether a change actually improves speed or another measured result.

Cursor rule for Cursor

Written for Cursor: a Cursor rule (.mdc). Also seen: mentions subagents.

Good fit Use it to tune performance and document reproducible before-and-after comparisons.

Compare 6 cursor rules from other repositories ↓
Install with agentmods
npx agentmods add rules/project-vic-international/agentic-ai-development-course/auto-research
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Clone the repo
git clone --depth 1 https://github.com/Project-VIC-International/Agentic-AI-Development-Course

Made for: Cursor.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for auto-research

README.md
[![agentmods](https://agentmods.dev/badge/rules/project-vic-international/agentic-ai-development-course/auto-research.svg)](https://agentmods.dev/rules/project-vic-international/agentic-ai-development-course/auto-research)
Your own site
<a href="https://agentmods.dev/rules/project-vic-international/agentic-ai-development-course/auto-research"><img src="https://agentmods.dev/badge/rules/project-vic-international/agentic-ai-development-course/auto-research.svg" alt="Measured on agentmods" height="20"></a>
Per session 1,682 This file is loaded in full into every session.
When invoked 1,682 The same file — it is already loaded in full.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.01682 $0.01682
Opus 5 $0.00841 $0.00841
Sonnet 5 $0.00336 $0.00336
Haiku 4.5 $0.00168 $0.00168

Measured 8d ago against content hash bd46496c870a, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-08, from the pricing page.

Security

Grade A, and why

auto-research scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 8d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

part-2-lab/exercises/03-spec-driven-dev/example-cursor-rules/auto-research.mdc · 127 lines

How it starts

The opening of the file, as written. The whole thing — 127 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Auto-Research: Experiment-Driven Improvement

After a feature passes its spec requirements, use Andrej Karpathy's auto-research methodology to autonomously discover improvements. This is not unstructured brainstorming — it is a disciplined experiment loop that produces measured, defensible results.

For a CAC mission tool, this matters for two reasons:

  1. Performance — investigators need answers fast. A triage pass that takes 8 hours instead of 2 means children wait longer for help.
  2. Defensibility — every claimed improvement is backed by a logged, reproducible experiment. When a defense expert asks "how do you know your hash matcher is faster than the previous version?", the answer is a numbered experiment record with before / after measurements and the SHA of the branch that produced them.

The Loop

while improvements_possible:
    1. READ        — current spec thresholds + benchmark baselines
    2. HYPOTHESIZE — one concrete, testable change
    3. IMPLEMENT   — in an isolated branch or worktree
    4. MEASURE     — run benchmarks, compare to baseline
    5. EVALUATE    — keep if better, discard if not
    6. RATCHET     — update spec floor to new baseline if merged

The agent runs this loop. The investigator reviews each ratchet decision before it is merged.

Rules

1. Always Measure, Never Guess

  • Every hypothesis MUST have a measurable prediction ("this change should improve throughput by ~15%" or "this should reduce false positives below 0.5%").
  • Run the actual benchmarks. Do not eyeball code and declare it faster.
  • Record before / after numbers with the same workload, same hardware profile, and same measurement methodology. A faster benchmark on a different machine proves nothing.

2. One Variable at a Time

  • Each experiment changes exactly one thing: an algorithm, a buffer size, a parallelism strategy, a data structure, a model checkpoint, a confidence threshold.
  • Do not bundle multiple changes into one experiment. If you change two things and the metric improves, you cannot attribute the improvement to either change.

Read the full file on GitHub · 127 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 8d ago First seen · 127 lines · 1,682 tokens per session scan A bd46496c870a

Subscribe to this mod's changes

auto-research is a cursor rule published in the GitHub repository Project-VIC-International/Agentic-AI-Development-Course (5 stars, last pushed 4mo ago), licensed Apache-2.0. It adds 1,682 tokens to every session, about $0.0084 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.