Low Latency skills

14 tagged Low Latency, measured the same way as everything else here.

Browse within: Production 12FP8 7High Throughput 7INT4 7In-Flight Batching 7Inference Optimization 7Inference Serving 7Multi-GPU 7NVIDIA 7TensorRT-LLM 7serverless 6Auto-Scaling 5Managed Service 5hybrid-search 5

pinecone

01

ihatesea69/HieuNghi-AI-Skills

Skill Claude CodeCodex

Managed vector database for production AI applications. Fully managed, auto-scaling, with hybrid search (dense + sparse), metadata filtering, and namespaces. Low latency (<100ms p95). Use for production RAG, recommendation systems, or semantic search at scale. Best for serverless, managed infrastructure.

not rated 3 6mo ago A 63 tokens MIT

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: