Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add rules/rajitsaha/100xprism/evalgit clone --depth 1 https://github.com/rajitsaha/100xprismWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/rules/rajitsaha/100xprism/eval)<a href="https://agentmods.dev/rules/rajitsaha/100xprism/eval"><img src="https://agentmods.dev/badge/rules/rajitsaha/100xprism/eval.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00033 | $0.01042 |
| Opus 5 | $0.00016 | $0.00521 |
| Sonnet 5 | $0.00007 | $0.00208 |
| Haiku 4.5 | $0.00003 | $0.00104 |
Grade A, and why
eval scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 104 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Eval — Run skill evals and score them
Turns the dormant modules/<slug>/evals/evals.json files into a real, graded scorecard:
does the skill trigger on its prompts, and does its output satisfy each assertion?
The deterministic engine is scripts/eval-harness.py (discovery, validation, work-list,
scorecard rendering — no model calls). Grading is your job: fan the cases out to
subagents and have Haiku 4.5 judge every assertion with structured output. Installed path
is ~/100xprism/scripts/eval-harness.py; in a checkout it's scripts/eval-harness.py.
Phase 0 — Pick the target
# one module, everything, or just what changed on this branch:
python3 ~/100xprism/scripts/eval-harness.py validate --module <slug>
python3 ~/100xprism/scripts/eval-harness.py validate --all
python3 ~/100xprism/scripts/eval-harness.py validate --changed origin/main
Fix any structural errors before grading — a malformed eval file can't be scored.
Phase 1 — Get the work-list
python3 ~/100xprism/scripts/eval-harness.py plan --module <slug> --json
This emits { "modules": [ { "module", "cases": [ { id, prompt, expected_output, assertions[], files[] } ] } ] }. Each (case, assertion) is one unit of work.
Phase 2 — Grade with parallel subagents (Haiku 4.5)
For each case, dispatch one subagent (use the subagents skill / Agent tool, or a
Workflow fan-out) that:
-
Runs the prompt against the skill. Load the target skill (its SKILL.md) as context, then answer the case
promptexactly as the assistant would — this is the candidate response. Note whether the skill would have auto-triggered on that prompt (trigger accuracy) separately from output quality. -
Grades each assertion. Spawn a Haiku 4.5 grader (
model: claude-haiku-4-5) that, given the prompt, the candidate response, theexpected_output, and one assertion, returns structured output:{ "module": "<slug>", "case_id": <id>, "assertion": "<text>", "passed": true|false, "reason": "<one line>" }
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- yesterday First seen · 104 lines · 33 tokens per session scan A 87ae747dc880
eval is a cursor rule published in the GitHub repository rajitsaha/100xprism (10 stars, last pushed 4d ago), licensed MIT. It adds 33 tokens to every session and 1,042 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-03.
Other cursor rules, from other repositories
angular-20
This rule provides comprehensive best practices and coding standards for Angular development, focusing on modern TypeScript, standalone components, signals, and performance optimizations.
dev-standard
Apache Superset development standards and guidelines for Cursor IDE.
cli-error-handling
CLI command error handling patterns.
family-instance-domain-actions
Family instance domain action implementation patterns.
prefer-assertions-over-defensive-checks
Prefer assertions over defensive checks when data is guaranteed to be valid.
prefer-direct-imports-over-module-mocks
Prefer extracting a testable core over vi.mock / vi.resetModules when unit tests need to reach production logic entangled with config, env, or singletons.