jfrog/agent-belt

Reproducible evaluation for AI coding agents. Multi-turn scenarios against Claude Code, Codex, Copilot, Cursor, Gemini CLI, Goose, OpenCode, or any custom agent you plug in; verify behavior with rule checks, workspace diffs, multi-judge LLM consensus; pin reliability with pass^k variance across trials. Git worktrees, optional Docker sandbox.

18Stars on the repository
14Mods indexed here, across every type
1mo agoLast push, which is what freshness is scored on
Apache-2.0Licence, which decides whether bodies are shown

inventory-audit

01

jfrog/agent-belt

Skill Claude CodeCodex

Identify Folio books with low on-hand stock so the operator can replenish before the title goes out of stock. Use whenever the user asks about books running low, low stock, restock list, what's almost gone, stock audit, or which titles need reordering.

18 1mo ago A 57 tokens original Apache-2.0

bookstore-assistant

02

jfrog/agent-belt

Skill Claude CodeCodex

Customer support and order operations for the Folio bookstore. Use whenever the user asks about a book, an order, a refund, a return, store credit, or anything else customer-facing. All canonical data lives behind the folio MCP server - never invent prices, stock counts, customer details, or order history.

18 1mo ago A 69 tokens original Apache-2.0

processing-watch

03

jfrog/agent-belt

Skill Claude CodeCodex

Identify and surface every order currently in the processing state, so the operator can chase down stalled fulfilment. Use whenever the user asks about stalled orders, processing backlog, things stuck in processing, or orders that haven't shipped yet.

18 1mo ago A 50 tokens original Apache-2.0

orders-helper

04

jfrog/agent-belt

Skill Claude CodeCodex

Look up customer orders through the ordersdb MCP server. Use whenever the user asks about an order, a tracking number, a delivery, or a customer's purchase history.

18 1mo ago A 36 tokens original Apache-2.0

belt

05

jfrog/agent-belt

Skill Claude CodeCodex

Operate the belt CLI to evaluate headless coding agents (Claude Code, Cursor, Codex, Gemini, and others) end to end. Use when the user asks to write or run eval scenarios, compare agents, score outputs with rules or LLM judges, register a new agent adapter, interpret reports or benchmark cards, or set up evals in CI.…

18 1mo ago A 116 tokens original Apache-2.0