Deploys and validates AgentCore Runtime blueprints — handles the multi-stage process of Terraform foundation, container build, AgentCore Runtime provisioning, Cognito auth wiring, WebSocket proxy deployment, and integration testing. Use for blueprints in domains/agent-runtime/.
Analyzes benchmark JSON results and updates the benchmark report. Use when new benchmark runs complete or when comparing results across serving configurations.
Reviews a blueprint for completeness and coherence — checks that READMEs reference real files, configs match docker images, scripts exist, and cross-references between artifacts are valid.
Adversarial reviewer that red-teams a new spec or validation gate for knowledge that should have carried over from prior blueprint lessons but didn't. Use before a RALPH loop starts (Stage 0) or when reviewing a freshly written spec. Asks "which hard-won lesson from a previous deployment did this spec forget?" — NOT a…
Runs after a successful deployment or benchmark session to extract cross-cutting lessons and elevate them to steering files. Use when a RALPH loop completes or a capacity block session ends.
Deploys and validates infrastructure for a blueprint — handles the multi-stage process of Terraform apply, storage setup, capacity reservations, model staging, and pre-flight validation.
Drafts a new spec file from a brief description, following the project template and conventions from existing specs.
★not rated 4 1mo agoA25 tokens
At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: