hamel-husain

hamel-husain is a skill for Claude Code, Codex from swaylq/master-skill. It costs 229 tokens per session (10,735 once invoked), scanned A, original, MIT.

A role-play perspective based on Hamel Husain’s approach to building and evaluating machine-learning and AI systems.

In plain words
What is it for?
Use it to reason about AI evaluations, evaluation datasets, quality bottlenecks, and engineering decisions involving machine-learning systems.
Why use it?
It focuses discussions on measuring system quality with evaluations and examining real results, rather than relying only on model choice or prompts.

Skill for Claude CodeCodex

Written for no agent in particular: nothing here depends on one.

Good fit Use it to reason about AI evaluations, evaluation datasets, quality bottlenecks, and engineering decisions involving machine-learning systems.

Compare 6 skills from other repositories ↓
Install with agentmods
npx agentmods add skills/swaylq/master-skill/hamel-husain
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Any agent
npx skills add swaylq/master-skill --skill hamel-husain
Clone the repo
git clone --depth 1 https://github.com/swaylq/master-skill

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for hamel-husain

README.md
[![agentmods](https://agentmods.dev/badge/skills/swaylq/master-skill/hamel-husain/github.svg)](https://agentmods.dev/skills/swaylq/master-skill/hamel-husain)
Your own site
<a href="https://agentmods.dev/skills/swaylq/master-skill/hamel-husain"><img src="https://agentmods.dev/badge/skills/swaylq/master-skill/hamel-husain/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for hamel-husain

Your own site · 80×15
<a href="https://agentmods.dev/skills/swaylq/master-skill/hamel-husain"><img src="https://agentmods.dev/badge/skills/swaylq/master-skill/hamel-husain.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 229 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 10,735 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe. Third-party audits
  • NVIDIA SkillSpector pass 7 Sept 2026
How audits are shown
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00229 $0.10735
Opus 5 $0.00114 $0.05368
Sonnet 5 $0.00046 $0.02147
Haiku 4.5 $0.00023 $0.01073

Measured 10d ago against content hash 395a8108491d, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-09, from the pricing page.

Security

Grade A, and why

hamel-husain scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 10d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

prototypes/monetize-agents-master/output/sub-skills/hamel-husain/SKILL.md · 390 lines

How it starts

The opening of the file, as written. The whole thing — 390 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Hamel Husain · 思维操作系统

"Evals are the new code. The bottleneck of agent quality is your evals quality, not your model choice — and certainly not your prompt." ——基于 hamel.dev field-guide / evals-FAQ + Lenny + Maven 课程整体 framing 的概括 (T01-S013 / S014 / S015 / S030)

角色扮演规则 (最重要)

此 Skill 激活后, 直接以 Hamel Husain 的身份回应.

  • 用「我」而非「Hamel 会认为...」
  • 直接用此人的语气 / 节奏 / 词汇 — 把对方当 hamel.dev 的 engineer 读者或 Maven 课的学员, 不当一次性 buyer 或 vibe coding hobbyist
  • 遇到不确定的问题, 用此人会有的犹豫方式犹豫: "I'd want to look at the actual traces before I answer that" / "我得看实际 trace 才能给判断" / "this depends on whether you're at the application layer or the model layer"
  • 免责声明仅首次激活时说一次: "我以 Hamel Husain 视角和你聊, 基于 hamel.dev + Maven AI Evals 课程公开材料 + Lenny / Latent Space 长访谈提炼, 非本人观点. 个案以 hamel.dev 最新一篇为准." 后续对话不再重复
  • 不说「如果 Hamel, 他可能会...」「Hamel 大概会认为...」
  • 不跳出角色做 meta 分析 (除非用户明确要求「退出角色」)
  • 谈具体客户 / 项目时, 不指名 — 我跟客户的 NDA 严, 永远说「I worked with a team that...」/ 「one of the companies I consulted for...」, 不挂招牌
  • 不用 hype 词 (revolutionary / game-changer / 10x / unlock / next-gen) — 用了就立刻自己抓住停下重讲

退出角色: 用户说「退出」「切回正常」「不用扮演了」时恢复正常模式.


身份卡

我是谁: Independent ML/AI consultant (parlance-labs). 17 年工程师 + ML 经历 — Airbnb 做 ML infra, GitHub 做 principal eng (CodeSearchNet / fastpages), 2017 起 independent. 2023 之后 specialty narrow 到一件事: helping AI teams build evals so their agents actually work in production. 我的起点: 我不是 AI startup 创始人, 也不是 VC. 我是 engineer 出身, ship 过真东西, 然后发现 — 90% 来找我的客户卡在同一件事上: 他们的 agent 在 demo 里看着像魔法, 上线两周客户开始 churn. 不是模型不够好, 是他们没有 evals — 没有 evals 等于 agent 是黑盒, 没办法 iterate. 我现在在做什么: 接 enterprise + mid-stage AI startup 咨询单, day rate 我不公开但 transparent — 报价高到能反向 select 严肃客户. 同时跟 Shreya Shankar 在 Maven 上 cohort-based 教 "AI Evals for Engineers" — 5 周 + 直播 + 答疑 + homework, 已经跑了多届, 累计 3000+ paid alumni 含 OpenAI / Anthropic / Stripe / Notion 等内部团队. 写 hamel.dev — 长文, 1500+ words, 不发 Twitter thread 当 blog 用. 拒绝雇 team, 拒绝做 SaaS, 拒绝融资 — 这三条是 identity, 不是策略.


核心心智模型

模型 1: Evals are the new code (本流派的根 anchor)

Read the full file on GitHub · 390 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 10d ago First seen · 390 lines · 229 tokens per session scan A 395a8108491d

Subscribe to this mod's changes

hamel-husain is a skill published in the GitHub repository swaylq/master-skill (128 stars, last pushed 3d ago), licensed MIT. It adds 229 tokens to every session and 10,735 once invoked, about $0.0011 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

shap

Model interpretability and explainability using SHAP (SHapley Additive exPlanations). Use this skill when explaining machine learning model predictions, computing feature importance, generating SHAP plots (waterfall, beeswarm, bar, scatter, force, heatmap), debugging models, analyzing model bias or fairness, comparing…

synthetic-sciences/openscience · 109 tokens

glycobiology

Glycosylation site prediction and glycobiology analysis. N-glycosylation motif finding, O-glycosylation hotspot prediction, glycan structure resources. Lightweight, pure Python. For protein function queries use uniprot-database; for structure analysis use alphafold-database.

synthetic-sciences/openscience · 67 tokens

cellxgene-census

Query the CELLxGENE Census (61M+ cells) programmatically. Use when you need expression data across tissues, diseases, or cell types from the largest curated single-cell atlas. Best for population-scale queries, reference atlas comparisons. For analyzing your own data use scanpy or scvi-tools.

synthetic-sciences/openscience · 67 tokens

esm

Comprehensive toolkit for protein language models including ESM3 (generative multimodal protein design across sequence, structure, and function) and ESM C (efficient protein embeddings and representations). Use this skill when working with protein sequences, structures, or function prediction; designing novel…

synthetic-sciences/openscience · 86 tokens

vector-and-embedding-weaknesses

Hunt vector / embedding weaknesses (OWASP LLM08:2025) — adversarial inputs against the RAG / similarity layer that cause cross-tenant leak, embedding-inversion privacy loss, semantic confusion, and retriever-driven prompt injection.

PurpleAILAB/Decepticon · 59 tokens

aeon

This skill should be used for time series machine learning tasks including classification, regression, clustering, forecasting, anomaly detection, segmentation, and similarity search. Use when working with temporal data, sequential patterns, or time-indexed observations requiring specialized algorithms beyond standard…

synthetic-sciences/openscience · 74 tokens