awesome-cursor-rules-mdc is a generator that creates Cursor MDC rule files from structured library information, using semantic search and language models to gather and organize guidance. Developers use it to produce reusable rules for libraries in Cursor, and the catalogue includes 200 of those rules.
Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
git clone --depth 1 https://github.com/sanjeed5/awesome-cursor-rules-mdcWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/rules/sanjeed5/awesome-cursor-rules-mdc/lightgbm)<a href="https://agentmods.dev/rules/sanjeed5/awesome-cursor-rules-mdc/lightgbm"><img src="https://agentmods.dev/badge/rules/sanjeed5/awesome-cursor-rules-mdc/lightgbm.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.02840 | $0.02840 |
| Opus 5 | $0.01420 | $0.01420 |
| Sonnet 5 | $0.00568 | $0.00568 |
| Haiku 4.5 | $0.00284 | $0.00284 |
Grade A, and why
lightgbm scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 314 lines — stays where its author put it; the contents beside it link to each section on GitHub.
lightgbm Best Practices
LightGBM is our go-to for high-performance tabular modeling. These rules ensure our LightGBM implementations are fast, reliable, and maintainable.
1. Code Structure & Reproducibility
Always encapsulate model creation and ensure runs are repeatable.
1.1. Use Scikit-learn API & Model Builder Functions
Always use LGBMClassifier or LGBMRegressor for sklearn compatibility. Wrap model instantiation in a function for clean, version-controlled hyperparameter management.
❌ BAD: Inline model creation with hardcoded parameters
# my_script.py
import lightgbm as lgb
model = lgb.LGBMClassifier(n_estimators=100, learning_rate=0.1, max_depth=7)
model.fit(X_train, y_train)
✅ GOOD: Function-based model creation with type hints and external config
# models/lgbm_model.py
import lightgbm as lgb
from typing import Dict, Any
import yaml # Or use a dataclass for params
def build_lgbm_classifier(params: Dict[str, Any]) -> lgb.LGBMClassifier:
"""Builds a LightGBM Classifier with specified hyperparameters."""
return lgb.LGBMClassifier(**params)
# config/lgbm_params.yaml
# classifier_v1:
# objective: binary
# n_estimators: 500
# learning_rate: 0.05
# num_leaves: 31
# max_depth: 6
# random_state: 42
# n_jobs: -1
# colsample_bytree: 0.8
# subsample: 0.8
# reg_alpha: 0.1
# reg_lambda: 0.1
# my_script.py
import yaml
from models.lgbm_model import build_lgbm_classifier
with open("config/lgbm_params.yaml", "r") as f:
params = yaml.safe_load(f)['classifier_v1']
model = build_lgbm_classifier(params)
model.fit(X_train, y_train)
1.2. Set Reproducibility Flags
Ensure random_state (or seed) and deterministic=True are always set for consistent results.
❌ BAD: Non-deterministic training
model = lgb.LGBMClassifier(n_estimators=100) # Results vary on each run
✅ GOOD: Fully reproducible training
model = lgb.LGBMClassifier(n_estimators=100, random_state=42, deterministic=True)
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 4d ago First seen · 314 lines · 2,840 tokens per session scan A ef0b4b719612
lightgbm is a cursor rule published in the GitHub repository sanjeed5/awesome-cursor-rules-mdc (3,571 stars, last pushed 3mo ago), licensed CC0-1.0. It adds 2,840 tokens to every session, about $0.0142 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-03.
Other cursor rules, from other repositories
pyspark-etl-best-practices-cursorrules-prompt-file
Cursor rules for PySpark ETL development with code style, joins, window functions, map operations, and Iceberg patterns.
python-llm-ml-workflow-cursorrules-prompt-file
Cursor rules for Python LLM & ML development with workflow integration.
automl-hyperparameter-optimization
AutoML and hyperparameter optimization rules for Python ML projects using Ray Tune, Optuna, PyCaret, and time-series AutoML libraries.
fenic
Cursor rule "fenic" from typedef-ai/fenic, covering writing fenic, must-knows and traps fenic check can't catch — get these right by hand.
cursorrules
You are building an AI/ML project with Python. The project uses PyTorch for model training, handles data pipelines with proper validation, tracks experiments systematically, and follows production ML engineering practices. Code is type-hinted, tested, and reproducible.
sygaldry
You are an expert in Python, Mirascope, and the Sygaldry AI framework.