Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add asattelmaier/gnome-ui-mcp --skill gnome-ui-automationgit clone --depth 1 https://github.com/asattelmaier/gnome-ui-mcpWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/asattelmaier/gnome-ui-mcp/gnome-ui-automation)<a href="https://agentmods.dev/skills/asattelmaier/gnome-ui-mcp/gnome-ui-automation"><img src="https://agentmods.dev/badge/skills/asattelmaier/gnome-ui-mcp/gnome-ui-automation/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/asattelmaier/gnome-ui-mcp/gnome-ui-automation"><img src="https://agentmods.dev/badge/skills/asattelmaier/gnome-ui-mcp/gnome-ui-automation.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00045 | $0.00738 |
| Opus 5 | $0.00023 | $0.00369 |
| Sonnet 5 | $0.00009 | $0.00148 |
| Haiku 4.5 | $0.00005 | $0.00074 |
Grade A, and why
gnome-ui-automation scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 12d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 73 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Core Concepts
Desktop lifecycle: AT-SPI connects lazily on the first tool call. The server requires a running GNOME Wayland session with accessibility enabled (gsettings set org.gnome.desktop.interface toolkit-accessibility true).
Element identification: Use find_elements to search for UI elements by text, role, or application. Each element has a unique id (e.g. 0/1/2) that represents its path in the AT-SPI tree. Element IDs can become stale after UI changes.
Input backends: The server tries Mutter Remote Desktop first for keyboard and mouse input, falling back to AT-SPI if unavailable. This is transparent to the caller.
Workflow Patterns
Before interacting with an element
- Discover:
list_applicationsto see running apps - Find:
find_elementswith a text query and optional role filter - Verify: Check the element's
id,role, andboundsin the result - Act: Use
click_element,activate_element, ortype_text - Confirm:
wait_for_elementorscreenshotto verify the action took effect
Clicking elements
- Prefer
click_elementoverclick_atwhen you have an element ID - Use
activate_elementwhen click doesn't work (tries action, keyboard, then mouse fallback) - Use
find_and_activatefor a single find-then-activate step - Check
input_injectedandeffect_verifiedin the response to confirm success
Typing text
type_texttypes at the current keyboard focusset_element_textreplaces the full text content of an editable elementtype_intocombines OCR-based label finding with typing (for forms)
Menu navigation
navigate_menuwalks a menu path like["File", "Save As..."]- Each step finds and activates the menu item, waiting for submenus to appear
Window management
list_windowsto find windows,close_windowto close the focused onemove_window,resize_window,snap_windowfor layout controltoggle_window_statefor fullscreen/maximize/minimize
Efficient discovery
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 12d ago First seen · 73 lines · 45 tokens per session scan A bb789abd52a4
gnome-ui-automation is a skill published in the GitHub repository asattelmaier/gnome-ui-mcp (1 stars, last pushed 27d ago), licensed MIT. It adds 45 tokens to every session and 738 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
computer-use
A set of rules for controlling Windows desktop applications through screen observation and simulated mouse and keyboard actions. It also describes how to find and launch installed applications and verify the results.
web-navigator
A browser-automation routing guide for web tasks. It directs navigation, page inspection, and actions such as clicking or filling forms to the appropriate browser tool.
computer-use
Linux/X11 desktop control — inspect, click, type into visible windows via AT-SPI + xdotool. Returns @eN refs from the focused window's accessibility tree.
agent-desktop
Reliable computer use via native OS accessibility trees. Use when an AI agent needs to see and operate desktop applications (click buttons, fill forms, navigate menus, read UI state, toggle checkboxes, scroll, drag, type text, take screenshots, manage windows, use clipboard, manage notifications). Covers 59 command…
agent-desktop-ffi
C-ABI bindings over agent-desktop's PlatformAdapter. Consumers (Python ctypes, Swift, Node ffi-napi, Go cgo, C++, Ruby fiddle) link libagentdesktopffi.{dylib,so,dll} and call ad functions directly instead of spawning the CLI binary per call. The canonical observe-act workflow is: adinit → adadaptercreate[withsession]…
tauri-pilot
Inspect, interact with, and test a running Tauri v2 app via CLI. Communicates over Unix socket using JSON-RPC 2.0. Use when testing UI, automating interactions, or debugging a Tauri app.