datasets skills

95 tagged datasets, measured the same way as everything else here.

Browse within: ai-monitoring 39ai-observability 39aiengineering 39Evaluation 34gpt 33HuggingFace 10MLOps 10hub 10models 10hf 9academic 5polish 5

phoenix-cli

01

Arize-ai/phoenix

Skill Claude CodeCodex

Debug LLM applications using the Phoenix CLI. Fetch traces, analyze errors, structure trace review with open coding and axial coding, inspect datasets, review experiments, query annotation configs, and use the GraphQL API. Use whenever the user is analyzing traces or spans, investigating LLM/agent failures, deciding…

11k +28 today A 113 tokens

phoenix-github

02

Arize-ai/phoenix

Skill Claude CodeCodex

Manage GitHub issues, labels, project boards, sprint operations, and roadmap health for the Arize-ai/phoenix repository. Use when filing roadmap issues, triaging bugs, applying labels, running sprint close-out and rollover, auditing board hygiene, checking ticket-load balance across the team, keeping roadmap epics up…

11k +28 today A 90 tokens

pxi-eval-dataset

03

Arize-ai/phoenix

Skill Claude CodeCodex

Generate synthetic evaluation datasets for the PXI eval harness (evals/pxi/). Use whenever the user asks to create, author, draft, expand, or audit an eval dataset for a PXI tool, skill, or behavior — including phrases like "write evals for ", "test PXI behavior", "synthetic dataset for PXI", "cover this tool with…

11k +28 today A 139 tokens

agent-improve

04

langwatch/langwatch

Skill Claude CodeCodex

Turns production evidence into tested improvements for your AI agent. Forms hypotheses from real traces and analytics, explains the reasoning behind each one, then executes with the user: scenario tests that reproduce production failures, prompt and code changes as reviewable PRs, new evaluators and monitors that…

3.5k 2d ago A 85 tokens original Apache-2.0

connect-agent

05

langwatch/langwatch

Skill Claude CodeCodex

Connect the codebase's AI agent to LangWatch agent simulations over HTTP, so test suites run against it from the platform. Finds or adds the agent's chat endpoint, wires authentication for scenario traffic, makes the server adopt the W3C traceparent header so the judge reads the agent's own traces, registers the agent…

3.5k 2d ago A 97 tokens original Apache-2.0

prompt-optimization

06

langwatch/langwatch

Skill Claude CodeCodex

Improve a prompt on the evaluations workbench through a measured loop. Score the baseline first, then duplicate the target column, form a hypothesis from failing rows, edit the copy's prompt draft, run, compare pass rate and cost, and repeat until the numbers hold. Use when the user asks to optimize or improve a…

3.5k 2d ago A 105 tokens original Apache-2.0

firstdata

07

MLT-OSS/FirstData

Skill Claude CodeCodex

Find official portals, APIs, and download paths for authoritative primary data sources (governments, international organizations, research institutions, etc.). Use when users need to know "where to find this data from an official source", "which source is more authoritative", or "how to cite primary data". Covers…

182 15d ago A 78 tokens original MIT

benchmark-radar

08

ktwu01/benchmark-radar

Skill Claude CodeCodex

Find, inspect, and check AI benchmark records with the Benchmark Radar CLI. Use when a request needs benchmark discovery, details, recent Radar evidence, or local data health; do not assume why the user needs the results.

123 2d ago A 48 tokens original MIT

datenoio/internacia-db

Skill Claude CodeCodexCursor

Edit Internacia country and intblock YAML safely — validation, enrichment, provenance, and OpenSpec gates. Use when modifying data/countries/, data/intblocks/, data/blocktypes/, running validate scripts, enrich.py, or proposing schema changes in internacia-db.

11 12d ago A 60 tokens original MIT

internacia-query

10

datenoio/internacia-db

Skill Claude CodeCodexCursor

Query Internacia reference datasets — countries, borders, org membership, entity linking. Use when looking up country codes, UN members, NATO/EU/ASEAN rosters, border neighbors, Wikidata links, or joining on stable identifiers. Prefer DuckDB or Parquet over source YAML.

11 12d ago A 62 tokens original MIT

first-ask

11

asterixix/polish-academic-mcp

Skill Claude CodeCodex

Interactive, input-tool powered, task refinement workflow: interrogates scope, deliverables, constraints before carrying out the task; Requires the Joyride extension.

4 1mo ago A 34 tokens original MIT

mcp-cli

12

asterixix/polish-academic-mcp

Skill Claude CodeCodex

Interface for MCP (Model Context Protocol) servers via CLI. Use when you need to interact with external tools, APIs, or data sources through MCP servers, list available MCP servers/tools, or call MCP tools from command line.

4 1mo ago A 49 tokens original MIT

mcp-security-audit

13

asterixix/polish-academic-mcp

Skill Claude CodeCodex

Audit MCP (Model Context Protocol) server configurations for security issues. Use this skill when: Reviewing .mcp.json files for security risks Checking MCP server args for hardcoded secrets or shell injection patterns Validating that MCP servers use pinned versions (not @latest) Detecting unpinned dependencies in MCP…

4 1mo ago A 153 tokens original MIT

webhound-research

14

WebhoundAI/webhound-mcp

Skill Claude CodeCodex

Run budget-controlled Webhound reports or datasets when a question deserves a real, inspectable investigation. Use for market maps, due diligence, source verification, evidence-backed comparisons, cited reports, or structured web datasets where missing information could change a decision.

1 14d ago A 54 tokens original MIT

clean-example

15

frankxai/starlight-evals

Skill Claude CodeCodex

A minimal, deliberately clean SKILL.md fixture used to CI-test scripts/lint-skills.mjs against a known-good file. It has valid frontmatter, no dollar-digit sequences, no secret-shaped strings, and no personal paths.

0 4d ago A 50 tokens