auto-codabench CLAUDE.md

auto-codabench CLAUDE.md is an instructions file for coding agents from ihsaan-ullah/auto-codabench. It costs 2,697 tokens per session, scanned A, original, MIT.

Repository-specific instructions for AutoCodeBench, a Python library that helps create and validate Codabench competition bundles. Codabench is a platform for running machine-learning competitions.

In plain words
What is it for?
Installing, testing, validating, and understanding the AutoCodeBench library, its web interface, and its benchmark files.
Why use it?
It gives Claude Code the project structure, commands, testing rules, and authentication-free command paths needed to work consistently in the repository.

Instructions file

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add instructions/ihsaan-ullah/auto-codabench/claude-md
Clone the repo
git clone --depth 1 https://github.com/ihsaan-ullah/auto-codabench

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for auto-codabench CLAUDE.md

README.md
[![agentmods](https://agentmods.dev/badge/instructions/ihsaan-ullah/auto-codabench/claude-md.svg)](https://agentmods.dev/instructions/ihsaan-ullah/auto-codabench/claude-md)
Your own site
<a href="https://agentmods.dev/instructions/ihsaan-ullah/auto-codabench/claude-md"><img src="https://agentmods.dev/badge/instructions/ihsaan-ullah/auto-codabench/claude-md.svg" alt="Measured on agentmods" height="20"></a>
Per session 2,697 This file is loaded in full into every session.
When invoked 2,697 The same file — it is already loaded in full.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.02697 $0.02697
Opus 5 $0.01349 $0.01349
Sonnet 5 $0.00539 $0.00539
Haiku 4.5 $0.00270 $0.00270

Measured 4d ago against content hash e90ebe57f78c, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

auto-codabench CLAUDE.md scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

CLAUDE.md · 74 lines

How it starts

The opening of the file, as written. The whole thing — 74 lines — stays where its author put it; the contents beside it link to each section on GitHub.

CLAUDE.md

This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.

What this repo is

autocodabench: a pip-installable library for agentic authoring + pre-launch validation of Codabench competition bundles, built on the Claude Agent SDK. Target venue: JMLR MLOSS (see docs/ and the design discussion on branch jmlr-oss-direction). The package lives in src/autocodabench/; a Chainlit web UI (web/, deployed as an HF Space via Dockerfile) consumes the library; benchmark/ holds the pure-SDK end-to-end benchmarks (create-bench and validate-bench, both live under benchmark/).

The root README.md doubles as the HF Spaces metadata file — its YAML header configures the Space; don't remove it.

Commands

pip install -e .                          # editable install (also: pip install -e '.[dev]')
python -m pytest tests/                   # unit suite — fast, fully keyless, must stay that way

# Keyless CLI paths (work with no Claude auth at all):
autocodabench demo --out /tmp/demo        # rebuild+validate the demo bundle from a recorded run
autocodabench validate <bundle-dir-or-zip> [--facts facts.yaml]
autocodabench checks list                 # registered checks by tier, with citations

# Auth-requiring paths (subscription login preferred; ANTHROPIC_API_KEY second):
autocodabench auth status [--no-probe]    # active path + masked creds; verifies the SDK can sign in (live turn) unless --no-probe
autocodabench auth use <auto|subscription|api_key>   # choose; subscription hides any key from the SDK
autocodabench validate <bundle> --judged      # adds LLM-judged advisory checks
autocodabench plan-build-validate "<idea>" [--data D] [--pdf P]   # agentic plan→build→validate pipeline (alias: create; idea and/or a PDF proposal)
autocodabench plan "<idea>" [--data D]   # Phase 1 only → specs/implementation_plan.md
autocodabench build <plan.md | --run-dir D>  # Phase 2 only → build a bundle from a plan

# Any agentic command above accepts --backend: claude[:model] (default), ollama:<m>, openai:<m>, or URL#<m>.

python -m autocodabench.core.bundle_io    # core smoke test (demo bundle in a tempdir)
python -m autocodabench.mcp.server        # MCP stdio server (hangs on stdin — correct)
python scripts/make_demo_fixture.py       # regenerate the shipped replay fixture

cd web && chainlit run app.py --host 127.0.0.1 --port 8500 -h   # web UI (needs .env)

# Deploy the web UI to the Hugging Face Space (GitHub master is the source of
# truth; the script injects the HF README header and force-pushes to hf main):
scripts/deploy_hf.sh [--dry-run] [--yes] [src-ref]   # default src-ref: origin/master

# Benchmarks (pure-SDK; any backbone via --backend; needs Docker + a populated instrument):
python benchmark/autocodabench_create_bench/run.py --competition style-trans-fair --backend claude

Read the full file on GitHub · 74 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 4d ago First seen · 74 lines · 2,697 tokens per session scan A e90ebe57f78c

Subscribe to this mod's changes

auto-codabench CLAUDE.md is an instructions file published in the GitHub repository ihsaan-ullah/auto-codabench (2 stars, last pushed 1mo ago), licensed MIT. It adds 2,697 tokens to every session, about $0.0135 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other instructions, from other repositories

vscode buildNext.instructions.md

Working notes and architecture documentation for the new esbuild-based build system in build/next. Use when making changes to the new build pipeline (transpile/bundle commands, NLS plugin, source-map handling, resource copying, or self-hosting watch tasks).

microsoft/vscode · 6,785 tokens

spec-kit AGENTS.md

AGENTS.md instructions for github/spec-kit, covering agents.md, about spec kit and specify, quickstart — add a new integration in 5 steps, integration architecture and integrationmanifest — file tracking.

github/spec-kit · 7,104 tokens

codex AGENTS.md

AGENTS.md instructions for openai/codex, covering rust/codex-rs, the codex-core crate, code review rules, crate api surface and model visible context.

openai/codex · 5,182 tokens

langchain AGENTS.md

AGENTS.md instructions for langchain-ai/langchain, covering global development guidelines for the langchain monorepo, corridor security analysis, project architecture and context, monorepo structure and development tools & commands.

langchain-ai/langchain · 4,345 tokens

vscode oss-third-party-notices.instructions.md

Instructions for microsoft/vscode, covering vs code oss third-party-notices pipeline, architecture, pipeline flow in ci, applying the notice (cutover) and fallback chain (never fail the build).

microsoft/vscode · 5,001 tokens

next.js AGENTS.md

Instructions for vercel/next.js, covering next.js development guide, codebase structure, monorepo overview, core package: packages/next and other important packages.

vercel/next.js · 7,296 tokens