High-level workflow skills must close the loop by proving they can complete realistic tasks, not only by sounding plausible. Use capture-the-flag style evals for workflows such as logging in, uploading attachments and starting a chat, or granting a group access to a Workplace Agent.
Use this workflow when an operator explicitly asks for live Tessl evidence for a skill, or asks to use Tessl's scenario-generation skill to improve the local eval suite.
Use this contract whenever Skills SDK work needs runtime, judge, or Tessl proof. It encodes the promotion pipeline for skill changes and separates lanes that must not be substituted for one another.
Route ambiguous Harness Engineering requests to one lifecycle stage when users ask where to start, resume, plan, implement, review, debug, schedule a heartbeat, or resolve domain terminology.
Scan Codex session history for skill failures, usage patterns, and coverage gaps. Use when the user wants daily skill-health monitoring or evidence-backed recommendations about installing, improving, merging, or pruning skills.
Capture a completed Codex workflow as a reusable SKILL.md package by analyzing session context plus optional session-collector evidence, interviewing the user with structured prompts, and writing a validated skill artifact. Use when the user asks to skillify or operationalize a repeatable process.
Use when the user asks about Shachar Azriel's AI Native DevCon talk on executable specs, verification layers for agentic coding, planner/verifier separation, requirement mapping, product-gap detection, and safe review architecture.
Use when the user asks about Macey Baker and Baruch Sadogursky's workshop-style AI Native DevCon session on turning repeated agent work into skills, rules, scripts, hooks, and evals.
Use when the user asks about Christopher Batey's talk 'Building Product Teams in the Age of AI: What We Had to Relearn Every Quarter' (Latent Space, 2026) — including questions about running AI-assisted product engineering teams, his three pillars (path to production at AI speed, training/evaluating AI-enabled…
Answers questions about, retrieves safe excerpts from, explains concepts from, and summarizes key arguments in Birgitta Böckeler's talk "State of Play: AI Coding Assistants" (AI Native Dev conference, 2026). Use when the user asks about the last 12 months in AI coding assistants, model-task fit, LLM statelessness…
Use when the user asks about Justin Cormack's AI Native DevCon talk on tests, observability, AI-generated behavior, evidence, instrumentation, and keeping AI systems honest when test signals are incomplete.
Use when the user asks about Patrick Debois's talk "Coding Agents Don't Scale Themselves. Neither Do Your Teams. The Rise of Agent Enablement." — including questions about agent enablement teams, the three pillars (Enablement, Platform, Governance), the Context Development Lifecycle applied to org charts, AI product…
Answers questions about Brian Douglas's talk on training AI on your own code. Use when a user asks about Brian Douglas's pipeline for capturing agent sessions, extracting skills from traces, fine-tuning small local models, tapes/steros tooling, SFT vs DPO decisions, or wants to apply his agent telemetry and training…
Answers questions about, applies frameworks from, and generates artifacts based on Tammuz Dubnov's talk "When Our PM Started Writing Code: What Merge Rate Taught Us About AI Adoption." Grounds every response in the bundled transcript and outline files. Use when a user asks about AI-native org design, merge rate…
Summarizes, explains, and answers questions about Dave Farley's talk 'Vibe Coding — Is this really the best we can do?', including key arguments, frameworks, and recommendations. Provides verbatim-cited explanations of Farley's three properties of programming languages, audits AI-coding setups against his…
Use when the user asks about Maximiliano Firtman's ("Maxi") AI Native DevCon talk on Web MCP and the agentic web — including questions about how Web MCP differs from MCP, exposing frontend tools to AI agents, the imperative vs declarative Web MCP APIs, tool design guidance (name/description/input schema/execute), the…
Assists with questions about Hannah Foxwell's talk 'The Reinvention of the Dev Team'. Use when a user asks about Foxwell's arguments on agentic software development, engineering team composition, AI-driven velocity, dev-to-PM ratios, the three anchors (build something worth building, speed requires safety, people…