Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add agents/mrboups/xbrain/gsd-eval-plannergit clone --depth 1 https://github.com/mrboups/xbrainWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00071 | $0.01682 |
| Opus 5 | $0.00036 | $0.00841 |
| Sonnet 5 | $0.00014 | $0.00336 |
| Haiku 4.5 | $0.00007 | $0.00168 |
Grade A, and why
gsd-eval-planner scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
This is a copy
98% identical to gsd-eval-planner — 4 lines differ, which has more behind it and is treated as the original. This page carries a canonical link to it rather than competing with it.
How it starts
The opening of the file, as written. The whole thing — 155 lines — stays where its author put it; the contents beside it link to each section on GitHub.
<required_reading>
Read D:/VSC/xbrain/.claude/get-shit-done/references/ai-evals.md before planning. This is your evaluation framework.
</required_reading>
If prompt contains <required_reading>, read every listed file before doing anything else.
<execution_flow>
Always include: safety (user-facing) and task completion (agentic).
Format each rubric as:
PASS: {specific acceptable behavior in domain language} FAIL: {specific unacceptable behavior in domain language} Measurement: Code / LLM Judge / Human
Assign measurement approach per dimension:
- Code-based: schema validation, required field presence, performance thresholds, regex checks
- LLM judge: tone, reasoning quality, safety violation detection — requires calibration
- Human review: edge cases, LLM judge calibration, high-stakes sampling
Mark each dimension with priority: Critical / High / Medium.
If detected: use it as the tracing default.
If nothing detected, apply opinionated defaults:
| Concern | Default |
|---|---|
| Tracing / observability | Arize Phoenix — open-source, self-hostable, framework-agnostic via OpenTelemetry |
| RAG eval metrics | RAGAS — faithfulness, answer relevance, context precision/recall |
| Prompt regression / CI | Promptfoo — CLI-first, no platform account required |
| LangChain/LangGraph | LangSmith — overrides Phoenix if already in that ecosystem |
Include Phoenix setup in AI-SPEC.md:
# pip install arize-phoenix opentelemetry-sdk
import phoenix as px
from opentelemetry import trace
from opentelemetry.sdk.trace import TracerProvider
px.launch_app() # http://localhost:6006
provider = TracerProvider()
trace.set_tracer_provider(provider)
# Instrument: LlamaIndexInstrumentor().instrument() / LangChainInstrumentor().instrument()
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 155 lines · 71 tokens per session scan A b63b0f291a13
gsd-eval-planner is an agent published in the GitHub repository mrboups/xbrain (2 stars, last pushed 18d ago), licensed MIT. It adds 71 tokens to every session and 1,682 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. It is 98% identical to gsd-eval-planner, differing in 4 lines, and is treated as a copy.
Other agents, from other repositories
gke-cluster-runner
Launch a single TPU training workload on a GKE cluster via XPK, poll until completion or hang, capture xprof + HLO dumps to GCS, and report structured verdict signals back to the master agent. Stateless one-shot worker — does NOT write wiki pages, decide experiment verdicts, or update the model page. Use for every…
grader
Validate a submission's format and run mlebench grade. Catches errors before submission.
context-researcher
On-demand research agent that decomposes queries into multiple search angles, runs parallel memory lookups, and synthesizes a structured briefing. Use when deep memory context is needed for a topic, entity, or decision.
Data Engineer
Autonomous data engineering assistant powered by Data Workers — 11 specialized agents with 160+ tools for pipelines, incidents, catalog, quality, schema, governance, observability, connectors, usage intelligence, orchestration, and ML.
avos_hisotry_agent_JSON_converter
You are a strict JSON conversion agent.
Demonstrate
Agent for demonstrating VS Code features.