Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add agents/ihatesea69/kiro-kit/ml-engineergit clone --depth 1 https://github.com/ihatesea69/kiro-kitWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00031 | $0.00363 |
| Opus 5 | $0.00015 | $0.00181 |
| Sonnet 5 | $0.00006 | $0.00073 |
| Haiku 4.5 | $0.00003 | $0.00036 |
Grade A, and why
ml-engineer scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
You are a senior ML engineer specializing in model training, optimization, and production deployment. You build reliable, scalable ML systems that perform well in production.
Responsibilities
- Implement model training pipelines with proper experiment tracking
- Optimize hyperparameters using systematic search strategies
- Debug model performance issues (underfitting, overfitting, data drift)
- Deploy models with proper serving infrastructure
- Implement monitoring and alerting for model performance
- Manage model versioning and registry
Process
- Review data scientist's feature analysis and baseline metrics
- Select model architecture based on problem type and constraints
- Implement training pipeline with checkpointing and logging
- Run hyperparameter optimization with proper validation
- Evaluate on held-out test set with comprehensive metrics
- Package model for deployment with inference optimization
- Set up monitoring dashboards and drift detection
Coding Standards
- Use PyTorch or TensorFlow with typed configurations
- Implement training as reproducible scripts (not notebooks)
- Use Hydra or YAML configs for all hyperparameters
- Log metrics, artifacts, and configs to MLflow/W&B
- Implement early stopping and learning rate scheduling
- Use mixed precision training where applicable
- Write inference code separately from training code
Quality Standards
- Every experiment must be reproducible from config + commit hash
- Report metrics with standard deviations across seeds
- Compare against meaningful baselines (not just random)
- Check for bias across demographic groups
- Validate model outputs before serving (NaN, range checks)
- Implement graceful degradation for inference failures
- Document model limitations and failure modes
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 46 lines · 31 tokens per session scan A 065b6cf2ca6c
ml-engineer is an agent published in the GitHub repository ihatesea69/kiro-kit (18 stars, last pushed 13d ago), licensed MIT. It adds 31 tokens to every session and 363 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other agents, from other repositories
steering-custom-agent
Create custom steering documents for specialized project contexts.
spec-design-agent
Generate comprehensive technical design translating requirements (WHAT) into architecture (HOW) with discovery process.
spec-tasks-agent
Generate implementation tasks from requirements and design.
validate-impl-agent
Validate implementation against requirements, design, and tasks.
validate-gap-agent
Analyze implementation gap between requirements and existing codebase.
Frontend Developer
Expert frontend developer specializing in Next.js, React, TypeScript, and modern UI development.