agentic_coding_lab CLAUDE.md

Instructions for a research framework that studies how test-driven development, or TDD, and prompt styles affect code written by AI. TDD means writing tests as part of the development process.

In plain words
What is it for?
It helps run or reanalyze research questions, monitor Docker-based coding experiments, aggregate measurements, and create overview snapshots.
Why use it?
It organizes experiments, analysis, and reports so researchers can repeat studies and compare results consistently.

Instructions file

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add instructions/marcoemrich/agentic_coding_lab/claude-md
Clone the repo
git clone --depth 1 https://github.com/marcoemrich/agentic_coding_lab
Per session 3,491 This file is loaded in full into every session.
When invoked 3,491 The same file — it is already loaded in full.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.03491 $0.03491
Opus 5 $0.01746 $0.01746
Sonnet 5 $0.00698 $0.00698
Haiku 4.5 $0.00349 $0.00349

Measured 2d ago against content hash 7d79ee0c5ca3, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

agentic_coding_lab CLAUDE.md scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

CLAUDE.md · 143 lines

How it starts

The opening of the file, as written. The whole thing — 143 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Agentic Coding Lab

Research framework for studying how TDD workflow structures and prompt styles affect AI-generated code quality. Runs Claude Code CLI in Docker containers against coding katas, then analyzes results via RQ-driven aggregation.

Full documentation: README.md — methodology, RQ frontmatter schema, workflow descriptions, model configurations, script reference, metrics glossary.

Skills

/run-rq RQ-N

End-to-end RQ orchestration: validate README → generate fill plan → start Docker batch → monitor progress → aggregate → propose findings updates. Pure orchestration calling existing repo scripts.

  • Triggers: "run-rq", "fill RQ-N", "RQ-N voranbringen", "Forschungsfrage N starten"
  • Details: .claude/skills/run-rq/SKILL.md

/reanalyze RQ-N

Re-run analysis pipeline on all runs matching an RQ, reaggregate metrics, and propose findings updates. No new runs — only refreshes existing data after pipeline changes.

  • Triggers: "reanalyze RQ-N", "reanalyse", "Runs neu analysieren"
  • Details: .claude/skills/reanalyze/SKILL.md

/build-overview

Generates a frozen experiment-overview snapshot across all RQs under research/reports/. Runs generate-snapshot-skeleton.py, then fills synthesis sections from findings.md files.

  • Details: .claude/skills/build-overview/SKILL.md

/exact-coding-baseline-export [date] [source-workflow]

Mint a new exact-coding-baseline-YYYY-MM-DD/ snapshot under research/workflow-dev/export/. Auto-detects the current correctness-oriented source workflow from research/workflow-dev/workflow-construction.md (the "Default für korrekheits-kritische Arbeit" recommendation), or takes an explicit source name. Copies source files, applies the HITL transformation (Step-8 checkpoints, autonomy-level switch, mode-neutral execution rule), and writes README + VERSION inside .claude/. Templates (HITL consumable, README, tdd-execution-mode) live in the skill directory.

  • Triggers: "exact-coding baseline export", "neue exact-coding baseline", "exact-coding-baseline-export"
  • Details: .claude/skills/exact-coding-baseline-export/SKILL.md

Read the full file on GitHub · 143 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 143 lines · 3,491 tokens per session scan A 7d79ee0c5ca3

Subscribe to this mod's changes

agentic_coding_lab CLAUDE.md is an instructions file published in the GitHub repository marcoemrich/agentic_coding_lab (11 stars, last pushed 16d ago), licensed MIT. It adds 3,491 tokens to every session, about $0.0175 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other instructions, from other repositories

vscode_abap_remote_fs CLAUDE.md

Instructions for marcellourbani/vscode_abap_remote_fs, covering abap fs development guidelines, workflow and constraints.

marcellourbani/vscode_abap_remote_fs · 101 tokens

sheal AGENTS.md

Instructions for liwala/sheal, covering agent instructions, agent operating policy, 1. tdd discipline — strict order, the only exception: spikes and 2. tests assert intent.

liwala/sheal · 2,989 tokens

digital-service-orchestra skills.instructions.md

Instructions for navapbc/digital-service-orchestra, a project described as: Workflow infrastructure plugin for Claude Code projects — TDD-driven sprint management, review gates, hook parameterization, and multi-stack config.

navapbc/digital-service-orchestra · 83 tokens

spec-driven-tdd AGENTS.md

Instructions for strelov1/spec-driven-tdd: This repo is a skill-pack. When implementing an OpenSpec change, invoke the spec-driven-tdd skill and follow its lifecycle: plan in OpenSpec, isolate in a worktree, implement each task via TDD → simplify → review, then finish + archive.

strelov1/spec-driven-tdd · 121 tokens

mcp-repo-onboarding AGENTS.md

Instructions for rogermt/mcp-repo-onboarding, covering agents.md — agent & copilot instructions, current status, output verification, project identity and ⚠️ critical: tdd required.

rogermt/mcp-repo-onboarding · 2,056 tokens

arxiv-agent-mcp AGENTS.md

AGENTS.md instructions for tbaraniuk/arxiv-agent-mcp, covering agents.md, roles, test-writer, implementer and per-task loop.

tbaraniuk/arxiv-agent-mcp · 451 tokens