nubecita: Skill for Claude Code

.claude/skills/run-startup-bench/SKILL.md

run-startup-bench is a skill for Claude Code from kikin81/nubecita. It costs 71 tokens per session (2,143 once invoked), scanned A, original, MIT.

A project-specific procedure for running Nubecita's Android startup macrobenchmark and comparing its results. A macrobenchmark measures app behavior such as cold-start time on a real device.

In plain words
What is it for?
Use it to run the startup benchmark, compare results, measure profile impact, or regenerate the startup baseline profile.
Why use it?
It reduces common benchmark failures by checking connected devices and applying the required device setup first. It helps determine whether startup changes or a baseline profile improve cold launch performance.

Skill for Claude Code

Written for Claude Code: installed under .claude/.

This is kikin81/nubecita's own configuration. It tells Claude Code how to work on nubecita itself, so it is not a mod to install elsewhere. Copy it as a starting point and replace the rules that are about this project. Everything nubecita configures →

Reuse

Borrowing it

Nothing to install: this file belongs to kikin81/nubecita. Take a copy, put it at the same path in your own repository, and replace the rules that are about this project with yours.

Copy the file
curl -O https://raw.githubusercontent.com/kikin81/nubecita/main/.claude/skills/run-startup-bench/SKILL.md
Clone the repo
git clone --depth 1 https://github.com/kikin81/nubecita

Made for: Claude Code.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for run-startup-bench

README.md
[![agentmods](https://agentmods.dev/badge/skills/kikin81/nubecita/run-startup-bench/github.svg)](https://agentmods.dev/skills/kikin81/nubecita/run-startup-bench)
Your own site
<a href="https://agentmods.dev/skills/kikin81/nubecita/run-startup-bench"><img src="https://agentmods.dev/badge/skills/kikin81/nubecita/run-startup-bench/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for run-startup-bench

Your own site · 80×15
<a href="https://agentmods.dev/skills/kikin81/nubecita/run-startup-bench"><img src="https://agentmods.dev/badge/skills/kikin81/nubecita/run-startup-bench.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 71 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 2,143 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00071 $0.02143
Opus 5 $0.00036 $0.01071
Sonnet 5 $0.00014 $0.00429
Haiku 4.5 $0.00007 $0.00214

Measured 11d ago against content hash 780ca8dd9723, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-11, from the pricing page.

Security

Grade A, and why

run-startup-bench scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 11d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.claude/skills/run-startup-bench/SKILL.md · 164 lines

How it starts

The opening of the file, as written. The whole thing — 164 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Project-local skill for running :benchmark's StartupBenchmark and reading the comparison cleanly. Companion to benchmark/README.md (the long-form workflow doc).

Announce at start: "I'm using the run-startup-bench skill to run the macrobench and compare cells."

When to use

Trigger on any of:

  • "Run the startup bench" / "run macrobench" / "kick off the startup tests"
  • "Measure the profile's impact" / "is the baseline profile working"
  • "Compare bench results" / "is COLD startup faster than before"
  • "Regenerate the baseline profile" (the gen task — different command but same device hygiene)

If the user wants FeedScrollBenchmark instead, do NOT use this skill — that bench needs the asset-backed fake (nubecita-crmi.5 epic) which isn't shipped yet.

Pre-flight (always)

The two most expensive failure modes of :benchmark runs eat ~5–15 min each. Run all four before any bench command:

  1. Device connected, single instance.

    adb devices -l
    

    If a Pixel Tablet or other secondary device shows up over TCP (Wi-Fi ADB auto-pair), adb disconnect <ip:port> it — multi-device confuses gradle's picker and the next test fails with "No connected devices!".

  2. Stay-awake on USB (prevents mid-run sleep that drops ADB):

    adb -s <serial> shell settings put global stay_on_while_plugged_in 7
    
  3. Pin the serial in the gradle invocation:

    -Pandroid.testInstrumentationRunnerArguments.androidx.benchmark.targetDeviceSerial=<serial>
    
  4. Only the production generator needs sign-in. The gen task targets :app's bench flavor by default (fake repos, fake SignedIn, mock feed) — no sign-in, never routes to Login, but it produces a profile that's missing ~13k real network/crypto/serialization rules, so don't ship it. For the representative, shippable profile pass -PbaselineProfileEnvironment=production: that drives the real signed-in cold start, so first install productionNonMinifiedRelease and sign in (it fails fast if Splash routes to Login). The bench (not the generator) wipes app data each run, so its sign-in never needs to persist.

Read the full file on GitHub · 164 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 11d ago First seen · 164 lines · 71 tokens per session scan A 780ca8dd9723

Subscribe to this mod's changes

run-startup-bench is a skill published in the GitHub repository kikin81/nubecita (5 stars, last pushed today), licensed MIT. It adds 71 tokens to every session and 2,143 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

android_ui_verification

Automated end-to-end UI testing and verification on an Android Emulator using ADB.

sickn33/agentic-awesome-skills · 22 tokens

agent-device

Automates Apple-platform apps (iOS, tvOS, macOS), Android devices, and Amazon Vega OS TV apps in Vega Virtual Devices. Use when navigating apps, taking snapshots/screenshots where supported, driving TV remotes, tapping, typing, scrolling, extracting UI info, collecting evidence, or planning agent-device CLI commands.

callstack/agent-device · 69 tokens

dogfood

Systematically explore and test a mobile app on iOS/Android with agent-device to find bugs, UX issues, and other problems. Use when asked to dogfood, QA, exploratory test, find issues, bug hunt, or test this app on mobile.

callstack/agent-device · 55 tokens

ios-simulator

Verify and debug native, React Native, Expo, or Flutter apps on an iOS Simulator with agent-device. Use when an agent needs to launch an app, inspect its live UI, tap, type, scroll, validate a code change, collect failure evidence, or reproduce a workflow on an iPhone or iPad Simulator.

callstack/agent-device · 69 tokens

solopi-ai

A command-line framework for testing Android apps and devices with SoloPi, including on-device or cloud AI decision models. It manages devices, test cases, recorded interactions, replays, performance history, and evidence.

alipay/SoloPi · 127 tokens

eas-simulator

EAS service (paid). Run and control a user's app on a remote iOS/Android simulator hosted on EAS cloud. Read before running any eas simulator: commands - it has the current syntax for this experimental API. Use whenever the user needs a simulator they can't run locally - 'run my app on a cloud simulator', 'use eas…

expo/skills · 249 tokens