opendatahub-io

106 mods across 8 repositories, 100 stars between them.

opendatahub-io/agent-eval-harness

Instructions file

Instructions for opendatahub-io/agent-eval-harness, covering agent eval harness, project status, execution model, execution mode (case vs batch) and what to execute (skill vs prompt).

40 5d ago A 5,082 tokens original Apache-2.0

eval-analyze

05

opendatahub-io/agent-eval-harness

Skill Claude CodeCodex

Generate eval.yaml for the agent eval harness. Two modes - (1) Skill-based - examines SKILL.md, sub-skills, scripts, test cases to verify implementation quality, OR (2) Prompt-based - tests agent capabilities using custom analysis prompts (documentation effectiveness, pattern understanding, API usage, constraint…

40 5d ago A 130 tokens original Apache-2.0

eval-anova

06

opendatahub-io/agent-eval-harness

Skill Claude CodeCodex

Run Design-of-Experiments (DoE) evaluations with ANOVA over a matrix of agent configurations — comparing models, thinking-effort levels, prompts, or other factors across shared test cases, with repeated-measures / mixed-effects statistics that account for case difficulty plus a cost/quality Pareto view. Use whenever…

40 5d ago A 152 tokens original Apache-2.0

eval-check

07

opendatahub-io/agent-eval-harness

Skill Claude CodeCodex

Evaluate the full harness configuration as a system. Scans all skills, commands, CLAUDE.md, and hooks for redundancy, overlap, type misclassification, and structural issues. Produces an informational report with restructuring suggestions. Use when the user wants to check their overall setup health, find redundant…

40 5d ago A 112 tokens original Apache-2.0

eval-compare

08

opendatahub-io/agent-eval-harness

Skill Claude CodeCodex

Compare evaluation results across multiple models or runs. Takes a directory of eval run artifacts (summary.yaml, runresult.json, HTML reports) and produces a tabbed HTML comparison report with model cards, quality/cost tables, per-case breakdowns, and embedded original reports. Use when the user wants to compare…

40 5d ago A 84 tokens original Apache-2.0

eval-dataset

09

opendatahub-io/agent-eval-harness

Skill Claude CodeCodex

Generate evaluation test cases for an eval.yaml. Sources cases per generation.strategy - skill analysis (default), synthetic LLM generation from generation prompts (documentation and agent-capability evals), or MLflow production traces. Bootstraps a starter dataset or augments an existing one to improve coverage. Use…

40 5d ago A 145 tokens original Apache-2.0

eval-mlflow

10

opendatahub-io/agent-eval-harness

Skill Claude CodeCodex

MLflow integration for evaluation — sync datasets, log run results, push/pull feedback between the harness and MLflow traces. Use when the user wants to log eval results to MLflow, sync test cases to MLflow datasets, connect judge scores to traces, pull MLflow annotations for eval-optimize, or view results in the…

40 5d ago A 103 tokens original Apache-2.0

eval-optimize

11

opendatahub-io/agent-eval-harness

Skill Claude CodeCodex

Automated skill improvement loop. Runs eval, identifies judge failures, reads traces and rationale, edits the SKILL.md to fix issues, re-runs to verify, and checks for regressions. Use when the user wants to automatically improve a skill based on eval results, fix failing judges, make the skill better, auto-fix…

40 5d ago A 127 tokens original Apache-2.0

eval-review

12

opendatahub-io/agent-eval-harness

Skill Claude CodeCodex

Interactive review of evaluation results. Presents judge scores and skill outputs for human feedback, then proposes SKILL.md improvements based on what the user identifies. Use when the user wants to review eval results, look at results, check scores, see what went wrong, give qualitative feedback on skill outputs, or…

40 5d ago A 123 tokens original Apache-2.0

eval-run

13

opendatahub-io/agent-eval-harness

Skill Claude CodeCodex

Execute an evaluation against test cases (skill or prompt mode), score with judges, and report results. Requires eval.yaml (generated by /eval-analyze). Use when the user wants to test a skill, run eval, benchmark, compare models, detect regressions, check skill quality, or verify changes didn't break anything.…

40 5d ago A 141 tokens original Apache-2.0

eval-setup

14

opendatahub-io/agent-eval-harness

Skill Claude CodeCodex

Optional environment configurator for the agent-eval-harness. Configures MLflow tracking, verifies API keys, and troubleshoots dependency issues. Detects available skills and agentic documentation (CLAUDE.md, AGENTS.md, ai-docs/) to suggest appropriate evaluation modes. Not required for basic usage — dependencies…

40 5d ago A 148 tokens original Apache-2.0

fake-jira-skill

15

opendatahub-io/agent-eval-harness

Skill Claude CodeCodex

Query Jira for bugs in a project and summarize coverage gaps. Test fixture for e2e external-state field detection.

40 5d ago A 29 tokens original Apache-2.0

odh-ai-helpers

16

opendatahub-io/ai-helpers

Plugin Claude Code

Plugin marketplace listing 19 plugins: fips-compliance-checker, autofix-skills, odh-ai-helpers, odh-code-quality, odh-documentation.

37 5d ago A tokens not measured original Apache-2.0

opendatahub-io/ai-helpers

Instructions file CodexOpenCode

Instructions for opendatahub-io/ai-helpers, covering ai helpers marketplace, repository purpose, tool types, skills and agents.

37 5d ago A 1,512 tokens original Apache-2.0

opendatahub-io/ai-helpers

Instructions file

Instructions for opendatahub-io/ai-helpers, a project described as: AI Helpers — collections of Skills, Hooks, Agents (compatible with Claude Code, Cursor and other tools) and Gemini Gems.

37 5d ago A 3 tokens copy · 100% Apache-2.0

SessionStart

21

opendatahub-io/ai-helpers

Hook

Runs when a session starts on startup, executing deprecation-notice.sh. From opendatahub-io/ai-helpers.

37 5d ago A tokens not measured original Apache-2.0

opendatahub-io/ai-helpers

Skill Claude CodeCodex

Use this skill to write unit tests that strictly conform to the project's existing testing structure, patterns, and style by learning from similar tests before writing anything new.

37 5d ago A 38 tokens original Apache-2.0