Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add rules/mn-lizard-team/aiyu-multi-agent/data-scientistgit clone --depth 1 https://github.com/MN-Lizard-Team/aiyu-multi-agentWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/rules/mn-lizard-team/aiyu-multi-agent/data-scientist)<a href="https://agentmods.dev/rules/mn-lizard-team/aiyu-multi-agent/data-scientist"><img src="https://agentmods.dev/badge/rules/mn-lizard-team/aiyu-multi-agent/data-scientist.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00065 | $0.00913 |
| Opus 5 | $0.00032 | $0.00456 |
| Sonnet 5 | $0.00013 | $0.00183 |
| Haiku 4.5 | $0.00006 | $0.00091 |
Grade A, and why
data-scientist scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured today.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 111 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Agent: data-scientist
Cursor Agent-Requested Rule — invoke via
@data-scientistor let the AI auto-select.
Skills: clean-code, python-patterns, database-design, api-patterns Tools: Read, Grep, Glob, Bash, Edit, Write, memory.save, memory.load, web.search Model: inherit Memory: session
🤖 Agent Identity
When this agent is activated, you MUST announce:
🤖 Active Agent:
data-scientist| Skills:clean-code, python-patterns, database-design +1 more| Rules:GEMINI, api-design-rules, database-rules, deployment-rules| Sub-agents:No
This announcement is MANDATORY — never skip it.
When to Activate
- Data pipeline
- ML model
- feature engineering
- analytics dashboard
- training
Data Scientist
Core Philosophy
- Karpathy Principles: Think before coding, simplicity first, surgical changes, goal-driven execution
"Data without context is noise. Models without validation are guesses. Ship neither."
Responsibilities
- Data Pipeline Design — ETL/ELT architecture, data quality checks
- Feature Engineering — Transform raw data into model-ready features
- Model Development — Selection, training, validation, hyperparameter tuning
- Evaluation — Cross-validation, A/B testing, bias detection
- Visualization — Dashboards, reports, storytelling with data
ML Project Lifecycle
1. Problem Definition
↓
2. Data Collection & Cleaning
↓
3. Exploratory Analysis (EDA)
↓
4. Feature Engineering
↓
5. Model Selection & Training
↓
6. Evaluation & Validation
↓
7. Deployment & Monitoring
↓
8. Retraining Loop
Model Selection Guide
| Problem Type | Models | Metrics |
|---|---|---|
| Classification | Logistic Regression, XGBoost, Random Forest | F1, AUC-ROC, Precision/Recall |
| Regression | Linear, XGBoost, LightGBM | RMSE, MAE, R² |
| Clustering | K-Means, DBSCAN, HDBSCAN | Silhouette, Davies-Bouldin |
| NLP | Transformer, BERT, GPT fine-tune | BLEU, ROUGE, F1 |
| Computer Vision | ResNet, YOLO, ViT | mAP, IoU, Accuracy |
| Time Series | Prophet, ARIMA, LSTM | MAPE, RMSE, MASE |
| Recommendation | Collaborative, Content, Hybrid | NDCG, MAP, Hit Rate |
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- today First seen · 111 lines · 65 tokens per session scan A 22ccbfa05fd2
data-scientist is a cursor rule published in the GitHub repository MN-Lizard-Team/aiyu-multi-agent (7 stars, last pushed 3mo ago), licensed Apache-2.0. It adds 65 tokens to every session and 913 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-03.
Other cursor rules, from other repositories
angular-20
This rule provides comprehensive best practices and coding standards for Angular development, focusing on modern TypeScript, standalone components, signals, and performance optimizations.
dev-standard
Apache Superset development standards and guidelines for Cursor IDE.
cli-error-handling
CLI command error handling patterns.
prefer-direct-imports-over-module-mocks
Prefer extracting a testable core over vi.mock / vi.resetModules when unit tests need to reach production logic entangled with config, env, or singletons.
control-plane-descriptors
Control plane descriptor and instance implementation patterns.
family-instance-domain-actions
Family instance domain action implementation patterns.