run-eval

run-eval is a skill for Claude Code, Codex from EliBarak12/Elliot. It costs 36 tokens per session (970 once invoked), scanned A, original, MIT.

An evaluation workflow tests the quality and correctness of tools in an Elliot connector. Elliot is a platform for turning APIs and databases into tools that AI agents can use.

In plain words
What is it for?
Use it to run an existing YAML or JSON evaluation suite against a connector. Tests can check for errors, minimum rows, required fields, rejected parameter values, and token limits.
Why use it?
It replaces ad-hoc checking with repeatable test cases for successful results, required fields, invalid inputs, and estimated token use. This helps reveal tools that return wrong data or handle errors poorly.

Skill for Claude CodeCodex

Part of the elliot plugin — 10 skills, 2 hooks, 1 MCP server shipped together

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/elibarak12/elliot/run-eval
Any agent
npx skills add EliBarak12/Elliot --skill run-eval
Clone the repo
git clone --depth 1 https://github.com/EliBarak12/Elliot

Made for: Claude Code, Codex.

Or install elliot, the plugin that ships this one along with the rest of its 10 skills, 2 hooks, 1 MCP server.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for run-eval

README.md
[![agentmods](https://agentmods.dev/badge/skills/elibarak12/elliot/run-eval.svg)](https://agentmods.dev/skills/elibarak12/elliot/run-eval)
Your own site
<a href="https://agentmods.dev/skills/elibarak12/elliot/run-eval"><img src="https://agentmods.dev/badge/skills/elibarak12/elliot/run-eval.svg" alt="Measured on agentmods" height="20"></a>
Per session 36 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 970 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00036 $0.00970
Opus 5 $0.00018 $0.00485
Sonnet 5 $0.00007 $0.00194
Haiku 4.5 $0.00004 $0.00097

Measured 4d ago against content hash 0127d0b6d238, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

run-eval scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/run-eval/SKILL.md · 92 lines

How it starts

The opening of the file, as written. The whole thing — 92 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Run Eval Workflow

Available eval suites

!ls connectors/*.eval.yaml .elliot/eval/*.json 2>/dev/null || echo "(no eval suites found — see example at connectors/my-saas.eval.yaml)"

Steps

1. Find or create the eval suite

elliot_run_eval accepts either:

  • path — a path to an *.eval.yaml (preferred) or *.json suite file, OR
  • suite_id — the bare id of a JSON suite under .elliot/eval/<id>.json.

If a *.eval.yaml already lives next to the connector, call elliot_run_eval(path="connectors/<slug>.eval.yaml").

If none exists, offer to create one. The canonical shape is YAML:

name: My Connector Evals
connector: my-connector
version: "1.0.0"
cases:
  - id: list-items-no-filter
    description: Returns at least one row with the expected fields
    tool_id: list_items
    arguments: {}
    expect:
      no_error: true
      min_rows: 1
      fields_present: [id, name]
      max_token_estimate: 500
  - id: list-items-rejects-bad-status
    description: A bad enum value is rejected, not silently ignored
    tool_id: list_items
    arguments: { status: not-a-real-status }
    expect:
      error_code: INVALID_PARAM_VALUE

Cover the error paths, not just the happy path. A tool that returns good rows for good input but silently accepts bad input is not agent-ready — the agent gets an empty or wrong result with no signal. Add at least one case per tool that asserts a bad argument is rejected: set expect.error_code to the code you expect (INVALID_PARAM_VALUE for a bad enum/bound, MISSING_PARAM for an omitted required param, UNKNOWN_PARAM for a stray key). The case passes only if the tool raises that code, and fails if the call succeeds — so you prove the contract rejects what it should. (In a legacy JSON suite the equivalent is "expect_error": "INVALID_PARAM_VALUE" on the case.)

Evaluate your skills, not just your tools. A deterministic skill is served as one callable tool, so a case's tool_id can name a skill id — the runner executes the whole step chain end-to-end (each step bound to the last) and the same expect block applies to the skill's final output. Add a case for every multi-step workflow you ship (fields_present on the fields the last step should return, max_token_estimate on the one-call cost) so a connector's headline workflows are validated before publish, not just their individual steps.

Read the full file on GitHub · 92 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 4d ago First seen · 92 lines · 36 tokens per session scan A 0127d0b6d238

Subscribe to this mod's changes

run-eval is a skill published in the GitHub repository EliBarak12/Elliot (11 stars, last pushed 4d ago), licensed MIT. It adds 36 tokens to every session and 970 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

agent-code-analyzer

Agent skill for code-analyzer - invoke with $agent-code-analyzer.

ruvnet/ruflo · 19 tokens

agui-dotnet-streaming-chat

Get started with the AG-UI .NET SDK: bootstrap and run your first streaming-chat app (client + server) with the AG-UI .NET NuGet packages (AGUI.Client, AGUI.Server, AGUI.Formatting, AGUI.Abstractions). USE FOR: which packages to install and how to wire them; constructing an AGUIChatClient against an endpoint and…

ag-ui-protocol/ag-ui · 223 tokens

agui-dotnet-sample-step

Add a GettingStarted sample Step (a Server/Client pair) to the AG-UI .NET SDK that demonstrates one protocol feature the way we want users to write it. USE FOR: adding a new samples/GettingStarted/StepNN Server+Client pair, wiring it into AGUI.slnx and the integration-test project, giving it a deterministic…

ag-ui-protocol/ag-ui · 168 tokens

agui-dotnet-protobuf

Use the protobuf wire transport (instead of the default Server-Sent Events) for an AG-UI connection with the AG-UI .NET SDK — a compact binary event stream negotiated via the Accept header. USE FOR: making an AGUIChatClient prefer protobuf by wiring an AGUIEventStreamHandler with ProtobufEventStreamFormatter (then…

ag-ui-protocol/ag-ui · 162 tokens

revdiff-plan

Review the last Codex assistant message (plan, analysis, or proposal) with inline annotations in a TUI overlay. Extracts the most recent response from Codex rollout files and opens it in revdiff for review and annotation. Activates on "revdiff-plan", "review plan with revdiff", "annotate plan", "review last response"…

umputun/revdiff · 84 tokens

nw-ddd-eventsourcing

Event Sourcing and CQRS as DDD implementation patterns — when to use, aggregate event streams, projections, snapshots, sagas, upcasting, conflict resolution.

nWave-ai/nWave · 38 tokens