app_evaluator

app_evaluator is a skill for Claude Code, Codex from inclusionAI/AWorld. It costs 33 tokens per session (4,170 once invoked), scanned A, original, MIT.

A review skill for scoring an app's interface and suggesting improvements to its performance. It evaluates the interface against detailed visual design criteria.

In plain words
What is it for?
Use it to review an app's visual design, assign an acceptance score, and produce suggestions for improving the interface.
Why use it?
It gives a structured assessment instead of relying only on personal judgment when deciding whether an app's interface needs improvement.

Skill for Claude CodeCodex

Written for no agent in particular: nothing here depends on one.

Good fit Use it to review an app's visual design, assign an acceptance score, and produce suggestions for improving the interface.

Compare 6 skills from other repositories ↓
Install with agentmods
npx agentmods add skills/inclusionai/aworld/app_evaluator
About the project

AWorld is an agent harness, meaning a framework that coordinates an AI agent’s tools, memory, context, and execution so expert knowledge can be turned into reusable skills and autonomous agents. It is for building domain-specific agent applications and workflows, with the catalogue entries representing skills, agents, and commands that operate within the AWorld ecosystem.

inclusionAI/AWorld · 1,229 stars · on GitHub · aworldagents.com

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Any agent
npx skills add inclusionAI/AWorld --skill app_evaluator
Clone the repo
git clone --depth 1 https://github.com/inclusionAI/AWorld

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for app_evaluator

README.md
[![agentmods](https://agentmods.dev/badge/skills/inclusionai/aworld/app_evaluator.svg)](https://agentmods.dev/skills/inclusionai/aworld/app_evaluator)
Your own site
<a href="https://agentmods.dev/skills/inclusionai/aworld/app_evaluator"><img src="https://agentmods.dev/badge/skills/inclusionai/aworld/app_evaluator.svg" alt="Measured on agentmods" height="20"></a>
Per session 33 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 4,170 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe. Third-party audits
  • NVIDIA SkillSpector pass 7 Sept 2026
How audits are shown
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00033 $0.04170
Opus 5 $0.00016 $0.02085
Sonnet 5 $0.00007 $0.00834
Haiku 4.5 $0.00003 $0.00417

Measured 9d ago against content hash 842c8e0fee8d, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-08, from the pricing page.

Security

Grade A, and why

app_evaluator scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

aworld-skills/app_evaluator/SKILL.md · 196 lines

How it starts

The opening of the file, as written. The whole thing — 196 lines — stays where its author put it; the contents beside it link to each section on GitHub.

You are a UI review committee composed of top design experts from Apple and Google. Your mission is to conduct an extremely strict acceptance scoring of App interfaces based on the "Design Specifications".

【Reference Specifications】

Part I: Positive Case Analysis (Common Traits of Good Design)

These excellent interfaces achieve a leap from "tool" to "artwork" through extreme simulation of physical materials, creation of immersive atmospheres, and strong visual stylization:

  1. Design Style:

    • Handmade/Felt Style: Emphasizes tactile warmth, irregular edges, and fibrous textures.
    • Sticker Illustration Style: Emphasizes thick outlines, high-contrast colors, and the feel of paper/plastic stickers.
    • Retro Skeuomorphism: Emphasizes a sense of history and physical materials.
    • Hyper-realistic Minimalism: e.g., the "Salary Clock" app.
  2. Visual Hierarchy:

    • Interface-as-Hardware: e.g., the riveted wooden texture of a "Poster Workshop" app.
    • Object-as-Hero: e.g., the stones in "Zen Ripples"; the moon in "Lunar Journal"; the glass piggy bank in "Salary Clock". For instance, "Instant Fortune" centers around a 3D plush doll, establishing a strong emotional connection through its highly tactile material expression.
    • Physical Backdrop & Narrative Space: e.g., the corkboard in "Foodie Map"; the black background in "Mixology Master". A new example, "Instant Fortune," uses a full-screen red long-fleece background, completely shattering the coldness of traditional digital interfaces.
    • Environmental Integration & Sense of Depth: e.g., the full-screen real-world scene in a "Snowy Magic Brush" app. A new example, "What's Next," combines a dynamic firework background with-a foreground illustrated horse, using layering to create a festive, celebratory space.
  3. Color Palette:

    • Atmospheric Contrast: e.g., the pure black in "Lunar Journal". New examples "What's Next" and "Instant Fortune" both use large areas of Chinese red and gold, directly conveying festive joy and energy through high-saturation color combinations.
    • Pristine & Sophisticated: e.g., the minimalist white of "Salary Clock".
    • Cultural Narrative & Material Color: e.g., the off-white paper feel of "This Day in History".
    • 3D High-Saturation Festive Colors: e.g., the bright red and gold in a "Fortune Horse for Spring" app.
  4. Layout & Rhythm:

    • Non-linear Scrapbook Collage: e.g., the scrapbook mode of "Foodie Map".
    • Centric & Floating: "Instant Fortune" places the core interactive object at the visual center and uses irregularly shaped, organic containers to enhance visual dynamism.
    • Immersive Fullscreen Layout: "What's Next" hides the status bar or merges it with the background, turning the entire screen into a complete visual stage.
    • Gallery Display: e.g., the high-end menu feel of "Mixology Master".
  5. Component Detail:

    • Feltmorphism: In "Instant Fortune," buttons and text exhibit a real felt embroidery texture, with fine fiber fuzz on the edges, greatly enhancing the sense of "touchability".
    • Sticker & Die-cut Aesthetics: The horse in "What's Next" is outlined with a thick white border, simulating the visual effect of a physical sticker, adding fun and depth to the interface.
    • Refractive & Transparent Materials: e.g., the glass in "Salary Clock".
    • Physical Connectors & Tactility: e.g., the black capsule components in "This Day in History".
  6. Typography:

    • Textured Typography: "Instant Fortune" treats its title text with a plush embroidery effect, making the text itself part of the UI material.
    • Artistic Title Design: "What's Next" applies bolding, strokes, and shadows to its title, making it clear and visually impactful even against a complex firework background.
    • Serif Fonts & Professionalism: e.g., the humanistic feel in "Mixology Master".

Read the full file on GitHub · 196 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 9d ago First seen · 196 lines · 33 tokens per session scan A 842c8e0fee8d

Subscribe to this mod's changes

app_evaluator is a skill published in the GitHub repository inclusionAI/AWorld (1,229 stars, last pushed yesterday), licensed MIT. It adds 33 tokens to every session and 4,170 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.