llm-inference skills

147 tagged llm-inference, measured the same way as everything else here.

Browse within: llm-agents 61kv-cache 60diffusion 27disaggregated-serving 27kubernetes 27omni 27routing-engine 27llms 25gpu-computing 19gateway 17generative-ai 17prompt 17proxy 17inference 13

ai-dynamo/dynamo

Skill Claude CodeCodex

Consults the repository performance rules and applicable Dynamo and engine guides to select one evidence-backed optimization proposal, then writes the generator's knowledge-consult.md reasoning record. Use after perf-analyzer completes a valid AIPerf analysis and before create-optimization-hypothesis materializes a…

7.9k +27 today A 66 tokens

ai-dynamo/dynamo

Skill Claude CodeCodex

Deploys one assigned DynamoGraphDeployment and proves it with an OpenAI-compatible smoke test. Use when user-interviewer has captured the user-provided baseline DGD or hypothesis-challenger has approved a later DGD.

7.9k +27 today A 50 tokens

dynamo-docs

03

ai-dynamo/dynamo

Skill Claude CodeCodex

Adds, updates, moves, or removes content on the Dynamo Fern docs site — standard docs pages, catalog-driven recipe and feature-benchmark pages, examples, recipes, and translations — keeping everything in line with the documentation style guide. Use for any change under docs/, recipes/, or examples/ (new page, edit…

7.9k +27 changed today A 133 tokens

build-brightstaff

04

katanemo/plano

Skill Claude CodeCodex

Build the brightstaff native binary. Use when brightstaff code changes.

7.0k +12 13d ago A 19 tokens original Apache-2.0

build-wasm

05

katanemo/plano

Skill Claude CodeCodex

Build the WASM plugins for Envoy. Use when WASM plugin code changes.

7.0k +12 13d ago A 21 tokens original Apache-2.0

plano-agent-skills

06

katanemo/plano

Skill Claude CodeCodex

Best practices for building agents and agentic applications with Plano, including configuration, routing, orchestration, guardrails, observability, and deployment.

7.0k +12 13d ago A 35 tokens original Apache-2.0

add-cuda-kernel

07

flashinfer-ai/flashinfer

Skill Claude CodeCodex

Step-by-step tutorial for adding new CUDA kernels to FlashInfer.

6.3k +21 today A 18 tokens original Apache-2.0

neuron-core/neuron-ai

Skill Claude CodeCodex

Write tests for Neuron AI agents, RAG systems, workflows, and tools using the built-in testing utilities. Use this skill when the user mentions testing agents, writing unit tests, mocking AI providers, testing tool execution, verifying RAG retrieval, testing workflow behavior, or creating test cases for Neuron AI…

2.1k +10 today A 94 tokens original MIT

neuron-tool-creator

11

neuron-core/neuron-ai

Skill Claude CodeCodex

Create custom tools, toolkits, and MCP integrations for Neuron AI agents. Use this skill when the user mentions creating tools, building toolkits, extending Tool class, defining tool properties, implementing tool execution, MCP server integration, Model Context Protocol, connecting external tools, or tool guidelines.…

2.1k +10 today A 100 tokens original MIT

neuron-core/neuron-ai

Skill Claude CodeCodex

Build custom Neuron AI workflows with nodes, events, middleware, and human-in-the-loop patterns. Use this skill whenever the user mentions workflows, orchestration, event-driven systems, custom agents, complex multi-step processes, human-in-the-loop patterns, or wants to build a custom agentic system from scratch.…

2.1k +10 today A 89 tokens original MIT

add-unit-test

13

xLLM-AI/xllm

Skill Claude CodeCodex

Add or update xLLM unit tests in the repository. Use when Codex needs to create a new C++/CUDA/NPU/MLU unit test, place a test under tests/, wire it into CMake with cctest, update an existing test target, choose platform gates, or validate test naming and dependencies against current xLLM test conventions.

1.5k yesterday A 76 tokens original Apache-2.0

code-review

14

xLLM-AI/xllm

Skill Claude CodeCodex

Review code changes for quality, security, performance, and correctness following project-specific standards. Use when reviewing pull requests, examining git diffs, or when the user asks for a code review. This skill should be used proactively — when the user asks for a review without specifying commits, automatically…

1.5k yesterday A 71 tokens original Apache-2.0

xLLM-AI/xllm

Skill Claude CodeCodex

Use when the user wants to add, modify, debug, or review an xLLM TileLang Ascend kernel or specialization, including Python kernel definitions, generated Ascend-C source, runtime wrapper dispatch, TileLang CMake wiring, and NPU tests.

1.5k yesterday A 60 tokens original Apache-2.0

ai-compass

16

tingaicompass/AI-Compass

Skill Claude CodeCodex

Search and answer questions from the local AI-Compass AI knowledge base. Use when the user asks about AI models, tools, agents, RAG, multimodal systems, evaluation, learning paths, project resources, or recent AI developments covered by this repository; distinguish weekly updates from long-term topic knowledge…

930 +4 5d ago A 75 tokens

atlas-release

17

Avarok-Cybersecurity/atlas

Skill Claude CodeCodex

The Atlas build → verify → image → publish pipeline, plus the upstream-sync PR automation. Turns a merged commit into a serve-matrix-verified avarok/atlas-gb10 image that users can pull. Use when cutting an image, closing the main→:latest staleness gap, wiring the release gate, or auto-syncing the fork and opening a…

672 2d ago A 110 tokens AGPL-3.0

Avarok-Cybersecurity/atlas

Skill Claude CodeCodex

Enforce the measurement discipline for any performance claim (tok/s, TTFT, TPOT, wall, accuracy). Invoke BEFORE measuring, comparing, or quoting a perf number, and before writing one into a commit message, PR comment, BENCH.toml, or report. Born from the 2026-08-15 decode-rate flip-flop (29→34→13 asserted in sequence…

672 2d ago A 128 tokens AGPL-3.0

astrea

19

warpfront/hipfire

Skill Claude CodeCodex

Use for hipfire quant calibration, imatrix-driven experiments, KLD/PPL quality evaluation, k-map/format selection, MQ/HFQ/HFP/MFP tradeoff work, ParoQuant-style weight transform planning, and KV policy planning. Use when deciding whether a calibrated model candidate should be promoted, rejected, packaged, or sent…

565 3d ago A 80 tokens

warpfront/hipfire

Skill Claude CodeCodex

Use Kernel Atlas to collect phase-aware hipfire measurements and render ISA Fit View visualizations for AMD GPU kernels, quant formats, and architectures. Use when a user asks how MQ/HFQ/HFP/Q8 quants occupy hardware, asks for an ASCII ISA visualization, wants to compare gfx1010/gfx1030/gfx11/gfx12 kernel fit, or…

565 3d ago A 108 tokens

rebase-onto-modular

21

warpfront/hipfire

Skill Claude CodeCodex

Use when porting a hipfire feature/fix branch authored against pre-0.1.20 master onto post-modular master. Walks through the engine→hipfire-runtime + per-arch-crate split mechanically, then surfaces semantic conflicts that need human judgment.

565 3d ago A 61 tokens

RightNow-AI/AutoMegaKernel

Skill Claude CodeCodex

Use when optimizing or generating a CUDA megakernel for a HuggingFace Llama-family model with AutoMegaKernel (AMK), drives the correctness-gated propose -> eval -> keep/revert loop (or hands off to the unattended autoresearch driver).

137 +2 2mo ago A 57 tokens original MIT

onboard-model

23

cloudrift-ai/emmy

Skill Claude CodeCodex

Onboard or periodically reverify and benchmark a Hugging Face model on an exact target GPU platform. Use when asked to add a model recipe, refresh a maintained recipe on a supplied GPU server, benchmark serving, create reproducible experiments and a durable results report, fully qualify and tune the model's Emmy…

80 2d ago A 84 tokens original Apache-2.0

cloudrift-ai/emmy

Skill Claude CodeCodex

Use this skill when the user asks to re-run an article's benchmarks, reproduce blog post numbers, validate that an article URL still holds, check whether the latest code still performs like a published post, or otherwise compare re-measured Emmy results with published results. It fetches the article, finds its…

80 2d ago A 95 tokens original Apache-2.0