dev-guides-first

A project rule that requires reading relevant documents in docs/dev-guides before making development changes and updating those documents when needed.

In plain words
What is it for?
Use it before developing features that touch the listed systems, then use the guidance to implement the change and keep the documentation current.
Why use it?
It prevents changes from ignoring existing project rules for areas such as streaming, logging, authentication, language support, errors, and file paths.

Cursor rule for Cursor

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add rules/microsoft/data-formulator/dev-guides-first
Clone the repo
git clone --depth 1 https://github.com/microsoft/data-formulator

Made for: Cursor.

Per session 1,028 This file is loaded in full into every session.
When invoked 1,028 The same file — it is already loaded in full.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.01028 $0.01028
Opus 5 $0.00514 $0.00514
Sonnet 5 $0.00206 $0.00206
Haiku 4.5 $0.00103 $0.00103

Measured yesterday against content hash fb84b0b53e2c, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

dev-guides-first scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.cursor/rules/dev-guides-first.mdc · 71 lines

How it starts

The opening of the file, as written. The whole thing — 71 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Dev Guides First

Before Starting Any Development

Before implementing or designing a new feature, module, or significant change, always read the relevant documents in docs/dev-guides/ to understand existing conventions:

Guide When to Read
docs/dev-guides/1-streaming-protocol.md Any work on streaming endpoints or NDJSON protocol
docs/dev-guides/2-log-sanitization.md Any work involving logging, credentials, external services, or DataLoaders
docs/dev-guides/3-data-loader-development.md Any work on ExternalDataLoader, DataConnector, or connector routes
docs/dev-guides/4-authentication-oidc-tokenstore.md Any work on OIDC, TokenStore, AUTH_MODE, or SSO flows
docs/dev-guides/6-i18n-language-injection.md Any work on Agent prompts, Agent routes, backend user-visible messages, or frontend i18n
docs/dev-guides/7-unified-error-handling.md Any work on API errors, frontend API calls, streaming error events, or error tests
docs/dev-guides/8-path-safety.md Any work on backend file access, downloads, Workspace paths, Agent tools, DataLoaders, or sandbox config
docs/dev-guides/10-agent-knowledge-reasoning-log.md Any work on Agent knowledge injection, KnowledgeStore, reasoning logs, or experience distillation
docs/dev-guides/11-catalog-metadata-sync.md Any work on catalog sync, catalog_cache, catalog_annotations, metadata merge, Agent catalog tools, or frontend catalog browsing
docs/dev-guides/12-sandbox-session.md Any work on sandbox execution, Agent tool-calling loops, explore/execute_python code execution, or namespace management
docs/dev-guides/13-unified-row-limits.md Any work on row limits, data loading size caps, frontendRowLimit, MAX_IMPORT_ROWS, or DataLoader size parameter
docs/dev-guides/14-model-capability-runtime-degradation.md Any work on LLM Client calls, Agent LLM invocations, model capability checks, reasoning_effort, or image/vision degradation
docs/dev-guides/15-dataframe-serialization.md Any work on DataFrame→JSON serialization, Agent result rows, DataLoader sample_rows, or table API responses
.cursor/rules/i18n-no-hardcoded-strings.mdc Any work adding or changing user-visible strings in src/

Read the full file on GitHub · 71 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. yesterday First seen · 71 lines · 1,028 tokens per session scan A fb84b0b53e2c

Subscribe to this mod's changes

dev-guides-first is a cursor rule published in the GitHub repository microsoft/data-formulator (17,048 stars, last pushed 3d ago), licensed MIT. It adds 1,028 tokens to every session, about $0.0051 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.