203-analytics-engineering

A set of working rules for building Mirror’s analytics layer with dbt and DuckDB. dbt transforms raw data into organized datasets, while DuckDB is a database that runs locally.

In plain words
What is it for?
Use it when creating or reviewing dbt models, tests, documentation, privacy handling, daily summaries, scores, and anomaly filtering.
Why use it?
It provides checks for privacy, incremental processing, model tests, documentation, and consistent data transformations. This helps prevent unreliable or exposed personal data.

Cursor rule

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add rules/hamzaamjad/cursor-rules/203-analytics-engineering
Clone the repo
git clone --depth 1 https://github.com/hamzaamjad/cursor-rules
Per session 0 Nothing until a file matches its globs; then the whole rule loads.
When invoked 408 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00000 $0.00408
Opus 5 $0.00000 $0.00204
Sonnet 5 $0.00000 $0.00082
Haiku 4.5 $0.00000 $0.00041

Measured yesterday against content hash 58974df28e24, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

203-analytics-engineering scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

rules/200-domain/203-analytics-engineering.mdc · 28 lines

What it actually says

analytics-engineering.mdc

  • Purpose: Guide development of Mirror's analytics engineering layer using dbt+DuckDB for transforming raw personal data into AI-consumable information.
  • Requirements:
    1. Privacy-First: All transformations run locally. Set send_anonymous_usage_stats: false in dbt_project.yml
    2. Incremental Processing: Use {{ is_incremental() }} for efficiency on personal devices
    3. Semantic Layer: Models follow staging → intermediate → marts pattern for DIKW progression
    4. Testing Coverage: Every model needs primary key test + business logic validation
    5. Documentation: Each model requires description explaining its role in the semantic layer
  • Validation:
    • Check: Does model compile with dbt compile?
    • Check: Are PII fields hashed using {{ hash_pii() }} macro?
    • Check: Do incremental models have proper unique_key?
    • Check: Is there a corresponding test in schema.yml?
  • Patterns:
    • Daily Aggregations: Use date_trunc('day', timestamp) for consistent grouping
    • Score Calculations: Normalize to 0-100 range for AI interpretation
    • Anomaly Handling: Filter outliers in staging, don't propagate to marts
    • Mobile Sync: Only sync mart tables, never raw/staging data
  • Wildcard Ideas:
    • Real-time CDC: Use DuckDB's APPENDER API for streaming inserts
    • Graph Analytics: Model social health metrics using DuckDB's recursive CTEs
    • Federated Learning: Pre-aggregate personal models for privacy-preserving ML
  • Source References: Mirror dbt evaluation, DuckDB best practices
Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. yesterday First seen · 28 lines · 0 tokens per session scan A 58974df28e24

Subscribe to this mod's changes

203-analytics-engineering is a cursor rule published in the GitHub repository hamzaamjad/cursor-rules (2 stars, last pushed 1y ago), licensed MIT. It costs nothing until one of its globs matches a file; then it loads 408 tokens. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.