Tutti is a real-time collaboration space where people can work alongside multiple AI agents, including by sharing rooms, editing together, and borrowing agents. It is intended for users who want human-agent or multi-agent collaboration across devices, with open-source and VM versions available. The catalogue entries provide skills and instructions for working with Tutti.
Borrowing it
Nothing to install: this file belongs to tutti-os/tutti. Take a copy, put it at the same path in your own repository, and replace the rules that are about this project with yours.
curl -O https://raw.githubusercontent.com/tutti-os/tutti/main/.codex/skills/tutti-test-audit/SKILL.mdgit clone --depth 1 https://github.com/tutti-os/tuttiWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/tutti-os/tutti/tutti-test-audit)<a href="https://agentmods.dev/skills/tutti-os/tutti/tutti-test-audit"><img src="https://agentmods.dev/badge/skills/tutti-os/tutti/tutti-test-audit/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/tutti-os/tutti/tutti-test-audit"><img src="https://agentmods.dev/badge/skills/tutti-os/tutti/tutti-test-audit.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector pass
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00074 | $0.01618 |
| Opus 5 | $0.00037 | $0.00809 |
| Sonnet 5 | $0.00015 | $0.00324 |
| Haiku 4.5 | $0.00007 | $0.00162 |
Grade A, and why
tutti-test-audit scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 12d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 166 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Tutti Test Audit
Make every changed test earn its maintenance cost. A green test, a coverage increase, or a larger test count is not a quality verdict.
Read the root and closest scoped AGENTS.md, then read
docs/conventions/unit-testing.md before a non-trivial test change or review.
Follow docs/conventions/testing.md#validation-selection for commands and
validation scope.
Establish the evidence map
Before editing, inspect:
- the production owner, public entry point, relevant caller and callee;
- sibling implementations that share the same invariant;
- existing tests, fixtures, conformance suites, and repository checks;
- the changed-aware and platform lane that will select the evidence;
- relevant history for a regression or a suspicious existing test.
For Agent Host lifecycle semantics, also read
packages/agent/host/README.md and start with the conformance contract. Do not
reimplement lifecycle tests in an adapter.
Pass the authoring gate
Answer all six questions before adding or materially rewriting a test:
- Protected contract: What observable behavior, invariant, compatibility promise, or prior failure matters?
- Credible failure: What plausible faulty implementation must make the test fail?
- Coverage gap: Why would existing evidence not catch that failure? Can an existing table, scenario, or fixture express it without duplication?
- Owner and observer: Which module owns the decision, and what is the lowest boundary that can observe the risk without mocking it away?
- Production seam: Does the test require an export, flag, wrapper, global, or injection hook that no production caller needs? If so, move to the real boundary or redesign the owner.
- Execution: Which blocking lane selects the test, and which native OS must execute it?
A missing answer means the test is not ready. “No new test” is a valid outcome when stronger existing evidence already protects the contract or the proposed assertion only restates a type, literal, implementation detail, or presentation choice.
What ships with it
1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 12d ago First seen · 166 lines · 74 tokens per session scan A 8217ecf085b0
tutti-test-audit is a skill published in the GitHub repository tutti-os/tutti (3,717 stars, last pushed 7d ago), licensed Apache-2.0. It adds 74 tokens to every session and 1,618 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
mcp-app-verification
Comprehensive verification checklists for MCP Apps. Tests with basic-host reference, validates handler-before-connect, text fallback, resource URI linking, single-file bundling, host styling, CSP, and legacy pattern detection.
behavior-contract
Bug condition/postcondition formalization as testable Behavior Contracts. Defines invariants that must be preserved across fixes.
quality-hooks
Language-specific auto-lint/format/typecheck pipeline. Supports Python (ruff+pyright), TypeScript (prettier+eslint+tsc), Go (gofmt+golangci-lint). Auto-fix and convergence loops.
eval-harness
Evaluation harness for testing agent and skill quality through structured benchmarks, regression tests, and quality scoring.
hook-management
Session-scoped hook lifecycle management with enable/disable/status controls, execution profiling, and color-coded performance alerts.
spec-execution
6-phase iterative specification execution workflow covering implementation, testing, review, improvement, commit, and progress tracking with quality-gated convergence.