validation-playbook

A validation guide for a Linux desktop application and its update system. It lists checks for shell scripts, JavaScript, Rust code, package security, builds, and common desktop actions.

In plain words
What is it for?
Use it to run baseline tests, inspect build outputs and file hashes, check signed update data and package digests, and smoke-test login, project opening, terminals, file picking, notifications, and quitting.
Why use it?
It provides a repeatable way to detect broken builds, tampered or incomplete packages, and failures in important user workflows.

Agent

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add agents/ilysenko/codex-desktop-linux/validation-playbook
Clone the repo
git clone --depth 1 https://github.com/ilysenko/codex-desktop-linux
Per session 0 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 574 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00000 $0.00574
Opus 5 $0.00000 $0.00287
Sonnet 5 $0.00000 $0.00115
Haiku 4.5 $0.00000 $0.00057

Measured 3d ago against content hash aa436a2036aa, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

validation-playbook scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

docs/agents/validation-playbook.md · 69 lines

How it starts

The opening of the file, as written. The whole thing — 69 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Validation playbook

Fast source checks

bash -n install.sh scripts/lib/*.sh launcher/start.sh.template
bash tests/scripts_smoke.sh
node --test scripts/lib/upstream-linux-package.test.js
node --test scripts/patch-linux-window-ui.test.js scripts/lib/linux-features.test.js linux-features/*/test.js
cargo test -p codex-update-manager
cargo clippy -p codex-update-manager --all-targets -- -D warnings

Source-security matrix

Cover valid, corrupt, unsigned, and wrong-key InRelease; Packages digest mismatch; package digest mismatch; wrong name/version/architecture; unsupported architecture; and incomplete official payload. These are unit-tested with local fixtures and fail closed.

Baseline build

Build with features.example.json, inspect .codex-linux/build-info.json, and compare the official and staged resources/app.asar hashes. Confirm no runtime replacement, external CLI, or local content server appears in the staged tree. Confirm the desktop name is ChatGPT Community, its icon has the community mark, and the package/bin/path identity is still codex-desktop.

Smoke-test login, project open, terminal, file picker, URI launch, tray, notifications, clean quit, and a second launch on GNOME Wayland, KDE Wayland, and X11. Verify official/custom coexistence and shared-profile single-instance behavior.

When upgrading an installation made by the former Linux port, also exercise the recognized Browser/Chrome cache migration. Confirm the official clients use /tmp/codex-browser-use, the Chrome app-server parent is not group-writable, and unrelated/user-authored plugin caches are unchanged.

Features

Build every retained feature independently and run its adjacent test. Enabled drift must reject a candidate; disabled drift must not be probed. Retired IDs are ignored and arbitrary unknown IDs fail.

Package matrix

Inspect deb, RPM, pacman, AppImage, and Nix outputs on both architectures. Check official ELF/runtime payload, dependencies, desktop identity, AppArmor path, updater payload, and absence of official package-manager configuration. AppImage must never inject --no-sandbox.

Read the full file on GitHub · 69 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 3d ago First seen · 69 lines · 0 tokens per session scan A aa436a2036aa

Subscribe to this mod's changes

validation-playbook is an agent published in the GitHub repository ilysenko/codex-desktop-linux (3,744 stars, last pushed yesterday), licensed MIT. It costs nothing until one of its globs matches a file; then it loads 574 tokens. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.