workspace-eval

workspace-eval is a skill for Claude Code, Codex from alecs5am/ralphy. It costs 303 tokens per session (1,490 once invoked), scanned A, original, Apache-2.0.

A wrapper that scores one project against the custom evaluation rules configured for its workspace. It produces a per-criterion scorecard and a report, then routes failed criteria to repair.

In plain words
What is it for?
Use it after rendering a project when you need to check it against a workspace rubric. It does not apply fixes or replace the generic evaluator.
Why use it?
It shows exactly which workspace-specific requirements pass or fail, instead of giving only a single overall result.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/alecs5am/ralphy/workspace-eval
Any agent
npx skills add alecs5am/ralphy --skill workspace-eval
Clone the repo
git clone --depth 1 https://github.com/alecs5am/ralphy

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for workspace-eval

README.md
[![agentmods](https://agentmods.dev/badge/skills/alecs5am/ralphy/workspace-eval.svg)](https://agentmods.dev/skills/alecs5am/ralphy/workspace-eval)
Your own site
<a href="https://agentmods.dev/skills/alecs5am/ralphy/workspace-eval"><img src="https://agentmods.dev/badge/skills/alecs5am/ralphy/workspace-eval.svg" alt="Measured on agentmods" height="20"></a>
Per session 303 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 1,490 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00303 $0.01490
Opus 5 $0.00151 $0.00745
Sonnet 5 $0.00061 $0.00298
Haiku 4.5 $0.00030 $0.00149

Measured 5d ago against content hash 6f6947a62e4c, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

workspace-eval scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 5d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.agents/skills/workspace-eval/SKILL.md · 76 lines

How it starts

The opening of the file, as written. The whole thing — 76 lines — stays where its author put it; the contents beside it link to each section on GitHub.

workspace-eval — run the universe rubric

A thin wrapper over the #469 runner. Your input is a project id + its workspace's custom rubric; your output is the per-criterion scorecard plus, on a non-clean verdict, a handoff to repair. The engine is ralphy workspace eval — this skill is only the agent-facing surface.

ALSO FIRE

  • After a render, when the user wants a standalone score against the universe's bar without running the full four-stage /universe-studio orchestrator.
  • When the user drops a <project> and asks whether it clears the workspace's own criteria (distinct from the generic /evaluator gates).

DO NOT FIRE

  • For the generic structural / audio / caption / vision gates on an arbitrary mp4 — that is /evaluator. This skill scores against the per-workspace CUSTOM rubric.
  • For applying fixes — that is /fixer / the repair loop (#409). This skill only produces the scorecard and routes failures onward.
  • For a workspace WITHOUT an evaluators.json rubric — there are no custom criteria to score; the runner reports "no rubric configured". Route to /evaluator for the built-in gates instead.
  • As the per-stage gate inside the four-stage flow — that is /universe-studio, which drives this same runner per stage.

HARD INVARIANTS

Inherited from AGENTS.md:

  1. ralphy is the only entry-point. Run the eval through ralphy workspace eval — never call the model or probe the mp4 directly.
  2. Append-only. The runner archives any existing workspace-eval.json / workspace-eval-report.md to .vN before writing the fresh one; never --force-overwrite without the user asking.
  3. Read MODELS.md before overriding the vision model (--model).

Run it

ralphy workspace eval <project>

Flags:

  • --no-vision — deterministic criteria only, NO model call (a free pass; use it to check the code-checked criteria without spending).
  • --model <id> — override the deep-vision model (default google/gemini-3.1-pro-preview; check MODELS.md first).
  • --workspace <slug> — score against a different workspace's rubric (default: the project's registered workspace).
  • --video <path> — override the scored video (default <project>/render/final.mp4).

Read the full file on GitHub · 76 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 5d ago First seen · 76 lines · 303 tokens per session scan A 6f6947a62e4c

Subscribe to this mod's changes

workspace-eval is a skill published in the GitHub repository alecs5am/ralphy (129 stars, last pushed 10d ago), licensed Apache-2.0. It adds 303 tokens to every session and 1,490 once invoked, about $0.0015 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

imaging-data-commons

Query and download public cancer imaging data from NCI Imaging Data Commons using idc-index. Use for accessing large-scale radiology (CT, MR, PET) and pathology datasets for AI training or research. No authentication required. Query by metadata, visualize in browser, check licenses.

synthetic-sciences/openscience · 62 tokens

clinical-reports

Write comprehensive clinical reports including case reports (CARE guidelines), diagnostic reports (radiology/pathology/lab), clinical trial reports (ICH-E3, SAE, CSR), and patient documentation (SOAP, H&P, discharge summaries). Full support with templates, regulatory compliance (HIPAA, FDA, ICH-GCP), and validation…

synthetic-sciences/openscience · 71 tokens

clinical-decision-support

Generate professional clinical decision support (CDS) documents for pharmaceutical and clinical research settings, including patient cohort analyses (biomarker-stratified with outcomes) and treatment recommendation reports (evidence-based guidelines with decision algorithms). Supports GRADE evidence grading…

synthetic-sciences/openscience · 97 tokens

gget

Fast CLI/Python queries to 20+ bioinformatics databases. Use for quick lookups: gene info, BLAST searches, AlphaFold structures, enrichment analysis. Best for interactive exploration, simple queries. For batch processing or advanced BLAST use biopython; for multi-database Python workflows use bioservices.

synthetic-sciences/openscience · 66 tokens

flow-cytometry-analysis

Complete flow cytometry analysis pipeline. FCS file handling, compensation, manual/automated gating, immunophenotyping, CFSE proliferation analysis, cell cycle analysis (Dean-Jett-Fox), and apoptosis assays. Extends flowio with analytical workflows. For raw FCS parsing only use flowio.

synthetic-sciences/openscience · 67 tokens

histolab

Lightweight WSI tile extraction and preprocessing. Use for basic slide processing tissue detection, tile extraction, stain normalization for H&E images. Best for simple pipelines, dataset preparation, quick tile-based analysis. For advanced spatial proteomics, multiplexed imaging, or deep learning pipelines use pathml.

synthetic-sciences/openscience · 62 tokens