qad

A procedure for ModelOpt Quantization-Aware Distillation, or QAD: training a model to recover accuracy lost when converting it to a smaller numerical format. It runs the process through Megatron Bridge on Slurm, a cluster job scheduler.

In plain words
What is it for?
Use it when explicitly running QAD, preparing its data, launching or resuming Slurm jobs, exporting checkpoints, or deciding how to recover a quantized model.
Why use it?
It provides a prescribed recovery process when post-training quantization creates a measured accuracy gap, while preventing unnecessary expensive runs.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/nvidia/model-optimizer/qad
Any agent
npx skills add NVIDIA/Model-Optimizer --skill qad
Clone the repo
git clone --depth 1 https://github.com/NVIDIA/Model-Optimizer

Made for: Claude Code, Codex.

Per session 69 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 1,332 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00069 $0.01332
Opus 5 $0.00034 $0.00666
Sonnet 5 $0.00014 $0.00266
Haiku 4.5 $0.00007 $0.00133

Measured 2d ago against content hash 53854abd981a, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

qad scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

plugins/modelopt/skills/qad/SKILL.md · 107 lines

How it starts

The opening of the file, as written. The whole thing — 107 lines — stays where its author put it; the contents beside it link to each section on GitHub.

ModelOpt Quantization-Aware Distillation

QAD is expensive. Run it only when the user explicitly authorizes QAD for the target model or run. A Day-0, PTQ, evaluation, comparison, or recipe-search request alone is not authorization to start QAD.

Follow the supported workflow

Before constructing commands, read:

  • examples/megatron_bridge/README.md, especially PTQ, data preparation, QAD, export, and Slurm usage
  • examples/megatron_bridge/{quantize.py,distill.py} via --help
  • the common skill's environment-setup.md, workspace-management.md, and slurm-setup.md; also its remote-execution.md for remote Slurm

Treat the example README and --help output as authoritative for mutable flags, commands, containers, and checkpoint formats. This skill supports Slurm only.

Execute in this order

  1. Confirm the gap. Reuse only validated, comparable BF16/PTQ results and the exact benchmark configuration from preceding evaluation or recipe search; run missing, invalid, or non-comparable baselines. Confirm the target benchmarks and their context-length needs. Stop if the PTQ gap to BF16 is already below 1%.
  2. Reproduce PTQ and verify compatibility. In the target runtime, require AutoBridge.can_handle() for the target model and PTQ through quantize.py to succeed while preserving the exact preceding PTQ config or recipe: format, layer selection, calibration data/count, sequence length, and seed. A changed quantization setting is a new PTQ candidate and must be evaluated before QAD. In the master-rank .quant_summary.txt, require finite positive amax for enabled static quantizers; accept dynamic/format-defined None only when the recipe intends it. Treat the summary as rank-local under model parallelism.
  3. Choose topology explicitly. Derive the smallest fitting node count and TP/PP/CP/EP from student and teacher architecture, the chosen sequence length, and available GPU memory. Prefer CP before TP for small long-context models; keep EP=1 for dense models and ETP=1 because the current distill.py workflow does not support expert tensor parallelism. For MoE require:

Read the full file on GitHub · 107 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 107 lines · 69 tokens per session scan A 53854abd981a

Subscribe to this mod's changes

qad is a skill published in the GitHub repository NVIDIA/Model-Optimizer (3,612 stars, last pushed 2d ago), licensed Apache-2.0. It adds 69 tokens to every session and 1,332 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

systematic-debugging

Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.

obra/superpowers · 21 tokens

next-cache-components-adoption

Turn on Cache Components in a Next.js app and resolve the blocking routes it surfaces. Use when the user wants to enable, adopt, or migrate to Cache Components, flip the cacheComponents flag, work through a flood of blocking-prerender / instant validation errors, run the cache-components-instant-false codemod, or…

vercel/next.js · 95 tokens

babysit-pr

Babysit a GitHub pull request after creation by continuously polling review comments, CI checks/workflow runs, and mergeability state until the PR is merged/closed or user help is required. Diagnose failures, retry likely flaky failures up to 3 times, auto-fix/push branch-related issues when appropriate, and keep…

openai/codex · 114 tokens

imagegen

Generate or edit raster images when the task benefits from AI-created bitmap visuals such as photos, illustrations, textures, sprites, mockups, or transparent-background cutouts. Use when Codex should create a brand-new image, transform an existing image, or derive visual variants from references, and the output…

openai/codex · 113 tokens

cpu-profile-analysis

Analyze V8/Chrome CPU profiles (.cpuprofile) and DevTools trace files (Trace-.json). Use when: profiling performance, investigating slow functions, comparing code paths, finding bottlenecks, analyzing timeToRequest, understanding call trees from sampling profiler data, analyzing layout/paint/rendering, investigating…

microsoft/vscode · 71 tokens

next-cache-components-optimizer

Drive a Next.js route to instant navigation by setting up an agentic loop, under Cache Components / PPR, on initial load (hard navigation) and client-side navigation (soft navigation). Encode the goal as a failing @next/playwright instant() e2e and work it to green, one verified route at a time; the shipped test then…

vercel/next.js · 170 tokens