Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add VincentChuWaiChow/vanguard-frontier-agentic --skill databricks-genai-evaluation-observabilitygit clone --depth 1 https://github.com/VincentChuWaiChow/vanguard-frontier-agenticWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/vincentchuwaichow/vanguard-frontier-agentic/databricks-genai-evaluation-observability)<a href="https://agentmods.dev/skills/vincentchuwaichow/vanguard-frontier-agentic/databricks-genai-evaluation-observability"><img src="https://agentmods.dev/badge/skills/vincentchuwaichow/vanguard-frontier-agentic/databricks-genai-evaluation-observability/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/vincentchuwaichow/vanguard-frontier-agentic/databricks-genai-evaluation-observability"><img src="https://agentmods.dev/badge/skills/vincentchuwaichow/vanguard-frontier-agentic/databricks-genai-evaluation-observability.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector warn
SkillSpector: 1 finding, up to high
These are SkillSpector’s own severities. On a checked sample its high-severity flags on skills were ~96% false positives — a documented command, a public API, a “never do X” rule — so we show them as a caution to read, not a verdict. Why →
- high Privilege Escalation · line 73 Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.Fix: Remove references to credential paths. Use environment variables or secrets managers. For docs, use placeholder paths (e.g., /path/to/config). Never load .env or token files in production code paths.
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00113 | $0.03583 |
| Opus 5 | $0.00056 | $0.01792 |
| Sonnet 5 | $0.00023 | $0.00717 |
| Haiku 4.5 | $0.00011 | $0.00358 |
Grade A, and why
databricks-genai-evaluation-observability scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 7d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 140 lines — stays where its author put it; the contents beside it link to each section on GitHub.
databricks-genai-evaluation-observability
Purpose
This skill decides whether evaluation and observability are correctly designed for generative AI on Databricks: traces are instrumented with rich spans, trace storage is chosen for governance and durability, evaluation datasets have consistent expectations, judges are validated against human labels before regression claims, judge configuration is held constant across releases, human feedback is bias-checked, and cost/latency are measured accurately. Sound design avoids confounded regression detection, unvalidated judge conclusions, and real-time cost claims from BETA tables.
When to use
- A user is setting up MLflow Tracing instrumentation for an agent and needs to confirm span design and storage choice.
- A user is designing an evaluation run using
mlflow.genai.evaluate()and needs to select judges and scorers. - A user has detected a quality regression between releases and needs to confirm the regression is real and not due to judge variability.
- A user is building a human-feedback loop and needs to confirm annotator agreement and bias-checking practices.
- A user is setting up cost and latency observability for external models and needs to confirm data sources and aggregation cadence.
When NOT to use
- No evaluation dataset or judge selection is stated — ask for the specific dataset schema and judge list before reviewing.
- A regression claim rests only on a single LLM judge without independent validation — refuse and ask for human-label validation or a secondary signal.
- The question is about fixing the identified failing component (agent, retrieval, model) — route to the appropriate specialist.
- The question is about whether a quality change matters in business terms — route to
databricks-value-realization-agent. - The question is about release mechanics implicated in a regression — route to
databricks-developer-platform-agent.
Scope
- MLflow Tracing: instrumentation APIs, span hierarchy, auto-instrumentation frameworks, trace tagging for analysis.
- Trace storage: experiment-based (legacy) versus Unity Catalog OpenTelemetry Delta tables (
system.traces.*); implications for retention, governance, and SQL queryability. - Evaluation harness:
mlflow.genai.evaluate()design, dataset schema, predictions and expectations. - Judges and scorers: the judge-versus-scorer distinction, the ten single-turn judges (RelevanceToQuery, RetrievalRelevance, Safety, RetrievalGroundedness, Correctness, RetrievalSufficiency, Guidelines, ExpectationsGuidelines, ToolCallCorrectness, ToolCallEfficiency), the seven multi-turn judges, custom scorers.
- Judge validation: human-label holdout sets, inter-rater agreement checks, judge-consistency across releases.
- Regression detection: confounding factors, dataset stability, judge configuration constancy, independent corroboration.
- Human feedback and observability: feedback collection, bias-checking, cost and latency measurement.
What ships with it
6 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 7d ago First seen · 140 lines · 113 tokens per session scan A 168e167f5293
databricks-genai-evaluation-observability is a skill published in the GitHub repository VincentChuWaiChow/vanguard-frontier-agentic (22 stars, last pushed 3d ago), licensed Apache-2.0. It adds 113 tokens to every session and 3,583 once invoked, about $0.0006 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-04.
Other skills, from other repositories
r-ml
R machine learning packages. Use for classification, regression, clustering, deep learning, gradient boosting (xgboost, lightgbm), random forests, neural networks, and time series forecasting.
r-parallel
R parallel computing and high performance packages. Use for multi-core processing, distributed computing, Spark integration, and C++ acceleration with Rcpp.
broom
R broom package for tidying model outputs. Use for converting statistical model results to tidy data frames.
r-ml-boosting
R gradient boosting packages. Use for xgboost, lightgbm, gbm, and catboost.
lightgbm
R lightgbm package for gradient boosting. Use for fast, distributed, high-performance gradient boosting.
xgboost
R xgboost package for gradient boosting. Use for high-performance classification, regression, and ranking.