Borrowing it
Nothing to install: this file belongs to kwakseongjae/oh-my-design. Take a copy, put it at the same path in your own repository, and replace the rules that are about this project with yours.
curl -O https://raw.githubusercontent.com/kwakseongjae/oh-my-design/main/.claude/skills/omd-lab-02-design-harness/SKILL.mdgit clone --depth 1 https://github.com/kwakseongjae/oh-my-designWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/kwakseongjae/oh-my-design/omd-lab-02-design-harness)<a href="https://agentmods.dev/skills/kwakseongjae/oh-my-design/omd-lab-02-design-harness"><img src="https://agentmods.dev/badge/skills/kwakseongjae/oh-my-design/omd-lab-02-design-harness/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/kwakseongjae/oh-my-design/omd-lab-02-design-harness"><img src="https://agentmods.dev/badge/skills/kwakseongjae/oh-my-design/omd-lab-02-design-harness.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector pass
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00013 | $0.01468 |
| Opus 5 | $0.00006 | $0.00734 |
| Sonnet 5 | $0.00003 | $0.00294 |
| Haiku 4.5 | $0.00001 | $0.00147 |
Grade A, and why
omd-lab-02-design-harness scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 12d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 132 lines — stays where its author put it; the contents beside it link to each section on GitHub.
OmD Lab #02 — Design Harness Refinement
omd:harness의 정교화 작업을 위한 실험실. 동일한 design task를 서로 다른 하네스 설정 (v1, v2, v3, ...)으로 돌려서 품질·비용·실패모드를 비교한다.
왜
Lab #01이 "DESIGN.md 유무"의 영향을 봤다면, Lab #02는 하네스 자체의 설정(prompt 변형, persona pool, eval rubric, asset 정책 등)이 산출물에 미치는 영향을 본다.
디렉토리 구조
skills/omd-lab-02-design-harness/
├── SKILL.md (this file)
├── playbooks/
│ ├── v1.md (현재 baseline — first full implementation)
│ ├── v2.md (다음 실험 가설)
│ └── ...
├── runs/
│ ├── v1-run-<ts>-<slug>/ ← omd:harness가 v1 설정으로 돌린 산출물 전체
│ ├── v2-run-<ts>-<slug>/
│ └── ...
├── compare/
│ ├── README.md (어떤 task로 어떤 v를 비교했는가)
│ └── <task-id>/
│ ├── index.html (v1 vs v2 vs v3 동시 비교 뷰)
│ └── metrics.json (집계 비교 지표)
└── postmortem-aggregate.md (전 v 누적 학습)
Lab Run 프로토콜
새 v 정의 (Lab 운영자 — 사용자 또는 Claude)
playbooks/v<N>.md작성. 가설을 한 줄로:## Hypothesis v2 raises persona ABANDON budget from 3s → 5s, expecting fewer false-abandon and more useful friction signal.- v이 v과 바꾸는 것 정확히 명시. 1개 변수만 바꾸는 게 원칙.
- 변경 점이 sub-agent 프롬프트에 있으면
playbooks/v<N>/agents-overrides/*.md로 patch 보관.
Lab Run 실행
# 운영자가 수동 실행
omd harness "<task>" --lab v2
# 또는 사용자가 자연어:
# "이 task를 lab v2 설정으로도 돌려서 v1과 비교해줘"
--lab v<N>이 들어오면:
runs/v<N>-run-<ts>-<slug>/디렉토리에 산출물 적재- v의 agent overrides가 있으면 그걸 임시로
.claude/agents/에 덮어씌운 채 실행 (run 종료 시 원복) - run.log에
lab_version: v<N>기록
비교 뷰 생성
운영자가 동일 task에 대해 v1, v2, ..., vN의 run을 끝내면:
omd lab compare --task "<task-slug>" --versions v1,v2,v3
이게 compare/<task-slug>/index.html을 만든다. 4-패널 (또는 N-패널) 비교:
- 각 v의 brief.md / 주요 wireframe / DESIGN.md.patch / persona ABANDON 요약 / 토큰 비용
- 공통 metrics.json: per-v {iterations, total_tokens, persona_abandon_rate, deterministic_pass_rate, jury_score, time_to_handoff}
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 12d ago First seen · 132 lines · 13 tokens per session scan A 1cd1f062a6b5
omd-lab-02-design-harness is a skill published in the GitHub repository kwakseongjae/oh-my-design (500 stars, last pushed 5d ago), licensed MIT. It adds 13 tokens to every session and 1,468 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
dembrandt
A TypeScript command-line tool that renders a web page with Playwright and examines its actual styles. It turns colours, typography, spacing, borders, shadows, animation curves, components, and breakpoints into structured design tokens.
extract-design
Extract the full design language from any website URL. Produces 8 output files including AI-optimized markdown, visual HTML preview, Tailwind config, React theme, shadcn/ui theme, Figma variables, W3C design tokens, and CSS variables. Also runs WCAG accessibility scoring. Use when user says 'extract design', 'get…
designlang-tokens
Use when styling UI for cal.com — references the extracted design system tokens instead of inventing colors, spacing, or typography.
brandmd
Extract a website's design system into a DESIGN.md file. Use when starting a new frontend project, rebuilding a site, or when the user wants AI-generated UI to match an existing brand.
montology
A repo's vocabulary as a database, enforced against the code by a tree-sitter scan — in every language it declares (Python, TypeScript/JS, Go, Rust, Swift, Java, Ruby, Elixir, C/C++, and more). Use BEFORE naming anything in code — a class, struct, function, type, module, table, column, endpoint, event, env var, CLI…
design-engineering
Premium design engineering skill for agentic workflows — produces high-end, distinctive UI designs using DESIGN.md as the portable contract across Pencil MCP (in-IDE canvas), Figma MCP (team handoff + design tokens), and Google Stitch (vibe exploration + AI generation). Enforces anti-generic principles, WCAG 2.2 AA…