Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add DataSQRL/sqrl --skill data-observationgit clone --depth 1 https://github.com/DataSQRL/sqrlWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/datasqrl/sqrl/data-observation)<a href="https://agentmods.dev/skills/datasqrl/sqrl/data-observation"><img src="https://agentmods.dev/badge/skills/datasqrl/sqrl/data-observation/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/datasqrl/sqrl/data-observation"><img src="https://agentmods.dev/badge/skills/datasqrl/sqrl/data-observation.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00061 | $0.00673 |
| Opus 5.5 | $0.00024 | $0.00269 |
| Sonnet 5.5 | $0.00012 | $0.00135 |
| Haiku 4.5 | $0.00006 | $0.00067 |
Grade A, and why
data-observation scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 10d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
Observe provided data before modeling it
Always inspect provided data (real, sample, or test data) before designing schemas, connectors, or test data. Do not infer the schema from filenames, READMEs, or prior knowledge alone. Always read the provided data (or subsets thereof if too large) first. Always consider provided data as authoritative and flag any inconsistencies with READMEs, requirements, or other provided instructions. Never resolve such inconsistencies without user direction.
For each provided data file:
- Check the size first. Use the following commands:
ls -la <file>ordu -h <file>. The size decides how you inspect it and whether it needs special handling (step 4). - Read a bounded prefix — never the whole file — in a format-aware way:
- Text (
.jsonl,.ndjson,.csv,.tsv): usehead -100 <file>command - Gzipped (
.gz, e.g..csv.gz): usezcat <file> | head -100command — decompress only a prefix, never the whole file - Columnar Parquet: use
duckdb -c "SELECT * FROM '<file>' LIMIT 100"command. For.orc/.avro, DuckDB needs a dedicated reader (e.g.INSTALL avro; LOAD avro;thenread_avro('<file>')) or use another tool — a bareFROM 'file.orc'will not work - Archives (
.zip): useunzip -l <file>command to list members, then inspect a single small member
- Text (
- Understand the data from what you actually saw: column names and order, types, real value formats, null/optional fields, header rows, and edge cases. Base the schema and test data on the observed rows, not on assumptions.
- If a file is large (roughly tens of MB or more), do not wire a connector to it or leave it unchanged in the project. As DataSQRL copies every data file into
build/on each compile/test command, so a large file makes every run extremely slow (andcompile/testcommands used in downstream tasks will be blocked). Invoke the/handle-large-data-filesskill to sample it down or exclude it. - For CSV specifically: confirm the delimiter, quoting, and header, and check that rows have a consistent field count — provided CSV might be ragged (trailing delimiters, extra empty columns, the odd malformed row), which passes compile and fails the Flink job at runtime. Nail the column count from the real data, then configure the source with the
/manage-connectorskill (formats/csv.md, Ragged rows).
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 10d ago First seen · 21 lines · 61 tokens per session scan A 948d5a310cbd
data-observation is a skill published in the GitHub repository DataSQRL/sqrl (230 stars, last pushed today), licensed Apache-2.0. It adds 61 tokens to every session and 673 once invoked, about $0.0002 per session on Opus 5.5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-30.
Other skills, from other repositories
grpc-expert
Expert-level gRPC, Protocol Buffers, microservices communication, and streaming. Use when the user mentions Protocol Buffers, microservices, RPC, or streaming, or when the task involves gRPC Fundamentals, Communication Patterns, or Production Features.
kafka-expert
Expert-level Apache Kafka, event streaming, Kafka Streams, and distributed messaging. Use when the user mentions streaming, messaging, event driven, or Kafka Streams.
data-ingestion-pipeline
Build data ingestion pipelines for batch and streaming data from multiple sources. Covers extraction strategies, format normalization, deduplication, validation gates, and staging patterns. Triggers on data ingestion, ETL pipeline, or data import architecture requests.
openrouter
OpenRouter unified AI API - Access 200+ LLMs through single interface with intelligent routing, streaming, cost optimization, and model fallbacks.
llm-integration
Claude and OpenAI API patterns, prompt design, context management, streaming, tool use, cost control, multi-model routing.
fastapi-patterns
FastAPI production patterns — routing, dependency injection, background tasks, streaming, error handling, and async. Use when building or reviewing a FastAPI service.