llmops skills

138 tagged llmops, measured the same way as everything else here.

Browse within: openai 40large-language-models 35self-hosted 35agentic-workflows 18copilot 18mkdocs 18newsletter 18agent-orchestration 15coding-agents 14annotations 12feedback-loop 12hypothesis-testing 12prompt-engineering 12regression-testing 12

slack-gif-creator

01

BerriAI/litellm

Skill Claude CodeCodex

Knowledge and utilities for creating animated GIFs optimized for Slack. Provides constraints, validation tools, and animation concepts. Use when users request animated GIFs for Slack like "make me a GIF of X doing Y for Slack.".

58k yesterday A 51 tokens

infra-scaling

04

langfuse/langfuse

Skill Claude CodeCodex

Tune and review Langfuse autoscaling for web, web-iso, and web-ingestion. Use for Terraform scale settings, RPM targets, scaling bounds, task counts, cost/performance tradeoffs, or Datadog evidence in the infrastructure repo.

34k yesterday A 55 tokens

skill-creator

05

langfuse/langfuse

Skill Claude CodeCodex

Guide for creating effective skills. This skill should be used when users want to create a new skill (or update an existing skill) that extends Codex's capabilities with specialized knowledge, workflows, or tool integrations.

34k yesterday A 46 tokens

turborepo

06

langfuse/langfuse

Skill Claude CodeCodex

Configure and troubleshoot Turborepo monorepos. Use for turbo.json, task pipelines, dependsOn, caching, filters, affected packages, CI optimization, environment variables, package boundaries, or shared internal packages.

34k yesterday A 50 tokens

promptfoo/promptfoo

Skill Claude CodeCodex

Standards for creating redteam plugins and graders. Use when creating new plugins, writing graders, or modifying attack templates.

25k yesterday A 29 tokens original MIT

promptfoo-evals

08

promptfoo/promptfoo

Skill Claude CodeCodex

Write, refine, run, and QA promptfoo evaluation suites: promptfooconfig.yaml, prompts, providers, vars, tests, assertions, model-graded rubrics, transforms, datasets, exports, and CI gates. Use for non-redteam eval coverage, regression tests, or new eval matrices. Do not use for adversarial redteam plugin or strategy…

25k yesterday A 79 tokens original MIT

promptfoo/promptfoo

Skill Claude CodeCodex

Create or refine promptfoo redteam setup configs: purpose, targets, plugins, strategies, frameworks, multi-input target inputs, policy text, grader guidance, contexts, and static-code-derived target/threat mapping. Use when preparing a red team scan plan from live probes, code evidence, or provider configs, or when…

25k yesterday A 99 tokens original MIT

trulens-diagnosis

10

truera/trulens

Skill Claude CodeCodex

Diagnose low evaluation scores and generate actionable improvement recommendations.

3.5k 3d ago A 17 tokens original MIT

truera/trulens

Skill Claude CodeCodex

Instrument LLM apps with TruLens OTEL-based tracing - from setup to debugging and optimization.

3.5k 3d ago A 25 tokens original MIT

running-all-tests

13

intentee/paddler

Skill Claude CodeCodex

Runs every test suite in the paddler workspace on the fastest available device. Use when the user asks to run the tests, run all the tests, run the full test suite, or check that everything still passes.

1.7k 1mo ago A 47 tokens original Apache-2.0

running-coverage

14

intentee/paddler

Skill Claude CodeCodex

Runs every test suite in the paddler workspace on the fastest available device, and produces code coverage report. Use when the user asks to run the code coverage, or to check the coverage.

1.7k 1mo ago A 42 tokens original Apache-2.0

phoenix-cli

15

Arize-ai/openinference

Skill Claude CodeCodex

Debug LLM applications using the Phoenix CLI. Fetch traces, analyze errors, structure trace review with open coding and axial coding, inspect datasets, review experiments, query annotation configs, and use the GraphQL API. Use whenever the user is analyzing traces or spans, investigating LLM/agent failures, deciding…

1.2k 2d ago A 113 tokens original Apache-2.0

Arize-ai/openinference

Skill Claude CodeCodex

Review Python OpenInference instrumentation code for correctness and completeness. Use this skill when reviewing a Python instrumentor package — whether it's a new instrumentor, a PR that modifies one, or when the user asks to audit/review/check an existing instrumentor's code quality. Trigger on phrases like "review…

1.2k 2d ago A 100 tokens original Apache-2.0

genai-conformance

17

Arize-ai/openinference

Skill Claude CodeCodex

Run, interpret, and iterate on the OpenInference GenAI conformance MVP at python/openinference-instrumentation/scripts/conformance/. Use when the user mentions GenAI conformance, OTel GenAI semantic conventions, Weaver registry live-check, the dual-write conversion (genaiconversion.py, enablegenaisemconv), genai.…

1.2k 2d ago A 95 tokens original Apache-2.0

add-model

18

adaline/gateway

Skill Claude CodeCodex

Add a new model to an existing provider.

605 1mo ago A 11 tokens original MIT

review-provider

19

adaline/gateway

Skill Claude CodeCodex

Review a provider implementation for correctness and consistency with other providers.

605 1mo ago A 14 tokens original MIT

research

20

SynaLinks/synalinks

Skill Claude CodeCodex

From idea to production in just few lines: Graph-Based Programmable Neuro-Symbolic LM Framework - a production-first LM framework built with decade old Deep Learning best practices.

455 yesterday not scanned tokens not measured original Apache-2.0

docs-design-tokens

21

rhesis-ai/rhesis

Skill Claude CodeCodex

Colour, font and design-token rules for the docs site — when hex is banned, when it is correct, and which surfaces ignore the theme. Use when styling docs components or editing CSS under docs/src.

391 3d ago A 46 tokens

worktree

22

rhesis-ai/rhesis

Skill Claude CodeCodex

Manage git worktrees with their own dev ports, symlinked .env files, playground, and simulations. Use when the user asks to create, set up, list, enter, or remove a worktree.

391 3d ago A 46 tokens

rhesis

23

rhesis-ai/rhesis

Skill Claude CodeCodex

Design, run, and analyze AI test suites on Rhesis — explore endpoints, build test foundations from a spec, create requirements and metrics, execute tests, and analyze results. Use when testing an AI endpoint, pasting a Product Requirements Document (PRD) or product spec, or working with Rhesis via MCP.

391 3d ago A 67 tokens

onboard-model

24

cloudrift-ai/emmy

Skill Claude CodeCodex

Onboard or periodically reverify and benchmark a Hugging Face model on an exact target GPU platform. Use when asked to add a model recipe, refresh a maintained recipe on a supplied GPU server, benchmark serving, create reproducible experiments and a durable results report, fully qualify and tune the model's Emmy…

80 yesterday A 84 tokens original Apache-2.0