jayll1303

35 mods across 1 repository, 18 stars between them.

semantic-router

25

jayll1303/AIEKit

Skill Claude CodeCodex

Build superfast AI decision layers with semantic-router. ALWAYS use for Route, SemanticRouter, HybridRouter, dynamic routes with function calling, intent classification, Pinecone/Qdrant index.

18 2mo ago A 40 tokens

sglang-serving

26

jayll1303/AIEKit

Skill Claude CodeCodex

Serve LLMs with SGLang for structured generation and high-throughput inference. Use when launching SGLang server, using RadixAttention, constrained JSON/regex output, or structured decoding.

18 2mo ago C 44 tokens

sherpa-onnx

27

jayll1303/AIEKit

Skill Claude CodeCodex

Offline speech processing with sherpa-onnx: ASR (streaming/non-streaming), TTS (Piper/Kokoro/Matcha/VITS), VAD, speaker diarization, speech enhancement. Use when running speech-to-text, text-to-speech, speaker identification, voice activity detection, or audio processing locally without internet.

18 2mo ago A 76 tokens

tensorrt-llm

28

jayll1303/AIEKit

Skill Claude CodeCodex

Optimize LLM inference with NVIDIA TensorRT-LLM. Use when running trtllm-build, converting HF checkpoints, serving with FP8/INT4, or maximizing GPU throughput.

18 2mo ago A 46 tokens

jayll1303/AIEKit

Skill Claude CodeCodex

Deploy and serve embedding/reranker models with HuggingFace TEI. Use when launching TEI Docker, configuring embeddings API, choosing embedding models, tuning batch performance, or building RAG pipelines with TEI.

18 2mo ago A 50 tokens

text-embeddings-rag

30

jayll1303/AIEKit

Skill Claude CodeCodex

Build local RAG pipelines with sentence-transformers, FAISS, ChromaDB, Qdrant. Use when generating embeddings, setting up semantic search, or optimizing retrieval with re-ranking and hybrid search.

18 2mo ago A 48 tokens

triton-deployment

31

jayll1303/AIEKit

Skill Claude CodeCodex

Deploy ML models on NVIDIA Triton Inference Server. Use when writing config.pbtxt, structuring modelrepository, building ensemble pipelines, configuring dynamic batching, writing tritonclient code, or debugging Triton model loading failures.

18 2mo ago A 50 tokens

ultralytics-yolo

32

jayll1303/AIEKit

Skill Claude CodeCodex

Train, predict, export, and deploy YOLO models with Ultralytics. Use when running yolo train/predict/val/export, building custom object detection/segmentation/classification/pose/OBB pipelines, preparing data.yaml datasets, or deploying YOLO to ONNX/TensorRT/CoreML/edge devices.

18 2mo ago A 71 tokens

unsloth-training

33

jayll1303/AIEKit

Skill Claude CodeCodex

Fine-tune LLMs 2x faster with 70% less VRAM using Unsloth. Use when using FastLanguageModel, Unsloth SFT/DPO/GRPO, 2x faster fine-tuning, or exporting to GGUF/vLLM.

18 2mo ago A 62 tokens

vllm-tgi-inference

34

jayll1303/AIEKit

Skill Claude CodeCodex

Deploy and serve LLMs locally with vLLM or TGI. Use when launching vllm serve, running TGI Docker, configuring tensor parallelism, serving quantized models, using OpenAI-compatible API, tuning KV cache, or diagnosing OOM errors.

18 2mo ago C 62 tokens

AIEKit AGENTS.md

35

jayll1303/AIEKit

Instructions file CodexOpenCode

AGENTS.md instructions for jayll1303/AIEKit, covering aie-skills, install and skills (34).

18 2mo ago A 311 tokens