Borrowing it
Nothing to install: this file belongs to ssf0409/tracelens. Take a copy, put it at the same path in your own repository, and replace the rules that are about this project with yours.
curl -O https://raw.githubusercontent.com/ssf0409/tracelens/main/CLAUDE.mdgit clone --depth 1 https://github.com/ssf0409/tracelensWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/instructions/ssf0409/tracelens/claude-md)<a href="https://agentmods.dev/instructions/ssf0409/tracelens/claude-md"><img src="https://agentmods.dev/badge/instructions/ssf0409/tracelens/claude-md.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.02536 | $0.02536 |
| Opus 5 | $0.01268 | $0.01268 |
| Sonnet 5 | $0.00507 | $0.00507 |
| Haiku 4.5 | $0.00254 | $0.00254 |
Grade A, and why
tracelens CLAUDE.md scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 6d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 278 lines — stays where its author put it; the contents beside it link to each section on GitHub.
TraceLens - Development Guide
Project Overview
TraceLens is an open source evaluation and regression-testing framework for AI agents. It turns agent runs into inspectable traces, graded outcomes, baseline comparisons, calibration data, and CI-ready reliability signals.
Keep this repository domain-agnostic. Checked-in docs and examples should work for external users without private project names, local absolute paths, or unpublished downstream integrations.
Ownership Boundary
TraceLens Owns
- Core data models:
Task,EvalSet,Trial,Transcript,Outcome - Execution primitives:
AgentAdapter,SimpleAdapter,HTTPAPIAdapter,EvaluationRunner(concurrency, timeouts, progress, checkpoint/resume) - Grader abstractions:
CodeGrader,LLMGrader,CompositeGrader - Built-in validators and budget/event-chain graders
- Statistical analysis:
pass@k,pass^k, bootstrap confidence intervals - Baseline management and regression detection
- Report rendering for markdown, JSON, HTML, and CI summaries
- Human-eval calibration: sample trial worksheets and reconcile human vs grader scores
- Reproducibility fingerprints via
DecisionSpec
Downstream Projects Own
- Domain task data and eval-set curation
- Agent invocation details and adapter subclasses
- Domain-specific graders and thresholds
- Baseline files and promotion policy
- CI policy for blocking, warning, or manual review
TraceLens evaluates evidence; it should not become the source of domain truth for a downstream project.
Key Files
src/tracelens/
├── core/
│ ├── task.py # Task, TaskLoader, EvalSet - test case definitions
│ ├── trial.py # Trial, TrialBatch - execution tracking
│ ├── grader.py # Grader ABCs - CodeGrader, LLMGrader, CompositeGrader
│ ├── transcript.py # Transcript - execution record
│ ├── decision_spec.py # DecisionSpec - reproducibility fingerprinting
│ └── outcome.py # Outcome - grading result (incl. grader_error flag)
├── execution/
│ ├── runner.py # EvaluationRunner - parallel execution, checkpoint/resume
│ ├── agent_adapter.py # AgentAdapter ABC, SimpleAdapter
│ ├── http_adapter.py # HTTPAPIAdapter for JSON endpoints
│ └── registry.py # Plugin loading via dotted import paths
├── statistics/
│ ├── pass_at_k.py # pass@k - capability ceiling
│ ├── consistency.py # pass^k - reliability measurement
│ ├── inference.py # Bootstrap CI, significance testing
│ └── latency.py # Latency aggregation helpers
├── baselines/
│ ├── manager.py # BaselineManager - store/retrieve/promote baselines
│ └── comparison.py # RegressionDetector - detect regressions
├── calibration/
│ ├── analyzer.py # CalibrationAnalyzer - grader vs human agreement
│ └── sampler.py # sample_for_review - select trials for human review
├── contracts/
│ └── contract.py # BehaviorContract - declarative grader generation
├── graders/
│ └── event_chain.py # Event-chain verifier
├── llm/
│ ├── provider.py # LLMProvider ABC and InMemoryProvider
│ └── factory.py # Provider factory policy
├── metrics/
│ ├── budgets.py # Latency/token/tool-call/trace consistency graders
│ └── validators.py # JSON schema, regex, contains, constraint graders
├── reporting/
│ └── generator.py # ReportGenerator - markdown, JSON, HTML, CI summary
└── cli/
├── main.py # run / report / sample / calibrate / reconcile
├── sample.py # Human review worksheet generation
└── calibrate.py # Human-vs-grader reconciliation
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 6d ago First seen · 278 lines · 2,536 tokens per session scan A 204cf87f2a87
tracelens CLAUDE.md is an instructions file published in the GitHub repository ssf0409/tracelens (2 stars, last pushed 12d ago), licensed MIT. It adds 2,536 tokens to every session, about $0.0127 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other instructions, from other repositories
flow-like copilot-instructions.md
Copilot instructions for Rheosoph/flow-like, covering flow-like wasm node development — rust, project overview, build & test, architecture and pin types.
next.js AGENTS.md
AGENTS.md instructions for vercel/next.js, covering next.js development guide, codebase structure, monorepo overview, core package: packages/next and other important packages.
codex AGENTS.md
AGENTS.md instructions for openai/codex, covering rust/codex-rs, the codex-core crate, code review rules, crate api surface and model visible context.
vscode buildNext.instructions.md
Working notes and architecture documentation for the new esbuild-based build system in build/next. Use when making changes to the new build pipeline (transpile/bundle commands, NLS plugin, source-map handling, resource copying, or self-hosting watch tasks).
vscode oss-third-party-notices.instructions.md
Instructions for microsoft/vscode, covering vs code oss third-party-notices pipeline, architecture, pipeline flow in ci, applying the notice (cutover) and fallback chain (never fail the build).
langchain AGENTS.md
AGENTS.md instructions for langchain-ai/langchain, covering global development guidelines for the langchain monorepo, corridor security analysis, project architecture and context, monorepo structure and development tools & commands.