Skill Claude CodeCodex
Build superfast AI decision layers with semantic-router. ALWAYS use for Route, SemanticRouter, HybridRouter, dynamic routes with function calling, intent classification, Pinecone/Qdrant index.
Skill Claude CodeCodex
Build superfast AI decision layers with semantic-router. ALWAYS use for Route, SemanticRouter, HybridRouter, dynamic routes with function calling, intent classification, Pinecone/Qdrant index.
Skill Claude CodeCodex
Serve LLMs with SGLang for structured generation and high-throughput inference. Use when launching SGLang server, using RadixAttention, constrained JSON/regex output, or structured decoding.
Skill Claude CodeCodex
Offline speech processing with sherpa-onnx: ASR (streaming/non-streaming), TTS (Piper/Kokoro/Matcha/VITS), VAD, speaker diarization, speech enhancement. Use when running speech-to-text, text-to-speech, speaker identification, voice activity detection, or audio processing locally without internet.
Skill Claude CodeCodex
Optimize LLM inference with NVIDIA TensorRT-LLM. Use when running trtllm-build, converting HF checkpoints, serving with FP8/INT4, or maximizing GPU throughput.
Skill Claude CodeCodex
Deploy and serve embedding/reranker models with HuggingFace TEI. Use when launching TEI Docker, configuring embeddings API, choosing embedding models, tuning batch performance, or building RAG pipelines with TEI.
Skill Claude CodeCodex
Build local RAG pipelines with sentence-transformers, FAISS, ChromaDB, Qdrant. Use when generating embeddings, setting up semantic search, or optimizing retrieval with re-ranking and hybrid search.
Skill Claude CodeCodex
Deploy ML models on NVIDIA Triton Inference Server. Use when writing config.pbtxt, structuring modelrepository, building ensemble pipelines, configuring dynamic batching, writing tritonclient code, or debugging Triton model loading failures.
Skill Claude CodeCodex
Train, predict, export, and deploy YOLO models with Ultralytics. Use when running yolo train/predict/val/export, building custom object detection/segmentation/classification/pose/OBB pipelines, preparing data.yaml datasets, or deploying YOLO to ONNX/TensorRT/CoreML/edge devices.
Skill Claude CodeCodex
Fine-tune LLMs 2x faster with 70% less VRAM using Unsloth. Use when using FastLanguageModel, Unsloth SFT/DPO/GRPO, 2x faster fine-tuning, or exporting to GGUF/vLLM.
Skill Claude CodeCodex
Deploy and serve LLMs locally with vLLM or TGI. Use when launching vllm serve, running TGI Docker, configuring tensor parallelism, serving quantized models, using OpenAI-compatible API, tuning KV cache, or diagnosing OOM errors.
Instructions file CodexOpenCode
AGENTS.md instructions for jayll1303/AIEKit, covering aie-skills, install and skills (34).