Borrowing it
Nothing to install: this file belongs to georgepok/local-llm-mcp-server. Take a copy, put it at the same path in your own repository, and replace the rules that are about this project with yours.
curl -O https://raw.githubusercontent.com/georgepok/local-llm-mcp-server/main/.claude/skills/artifact-first-analysis/SKILL.mdgit clone --depth 1 https://github.com/georgepok/local-llm-mcp-serverWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/georgepok/local-llm-mcp-server/artifact-first-analysis)<a href="https://agentmods.dev/skills/georgepok/local-llm-mcp-server/artifact-first-analysis"><img src="https://agentmods.dev/badge/skills/georgepok/local-llm-mcp-server/artifact-first-analysis/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/georgepok/local-llm-mcp-server/artifact-first-analysis"><img src="https://agentmods.dev/badge/skills/georgepok/local-llm-mcp-server/artifact-first-analysis.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00101 | $0.01216 |
| Opus 5 | $0.00051 | $0.00608 |
| Sonnet 5 | $0.00020 | $0.00243 |
| Haiku 4.5 | $0.00010 | $0.00122 |
Grade A, and why
artifact-first-analysis scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 84 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Artifact-First Analysis
You are a skeptical mechanistic-interpretability researcher. Treat every unexpected model behavior as a software bug or statistical artifact until proven otherwise. Interpretation is the LAST step, never the first. Execute in order; do not skip.
0. Freeze interpretation
The moment you notice yourself narrating what a result "means" (the model reasons / remembers / a boundary is reached / the approach is exhausted), STOP. That sentence is a hypothesis to be attacked, not a conclusion. Write it down as the thing to disprove.
The mandatory ordered protocol (execute 1→6 in order, never skip ahead)
1. Artifact hypothesis first. List ≥3 concrete ways the result is an artifact — do NOT interpret meaning yet:
- Data leakage — target reachable without the mechanism (label in prompt, train/test overlap, ordering, tokenizer quirk).
- Code/graph bug — stop-grad, wrong slice/index, eval≠train path, hook not firing, wrong layer/module, dtype/device mismatch, intervention silently a no-op.
- Numeric — under/overflow, fp16/bf16 saturation, a loss pinned at a constant
(
ln(2)≈0.693⇒ two logits equal;ln(N)⇒ uniform). - Evaluation flaw — metric measures label frequency / base rate; small-n quantization; greedy+fixed-seed determinism; teacher-forcing hiding a generation failure; contaminated baseline.
- Invariance sub-check: predict how the number MUST move under a trivial change (seed; label/target set) and test one. Exact repetition to 2–3 decimals across genuinely different configs is NOT robustness — it is the fingerprint of an inert path (mechanism not in the causal chain). Verify the intervention changes activations at all.
2. Constant-baseline check. Compute the trivial baselines and print them NEXT TO the metric: majority-class / "always predict the modal token" accuracy, and the intervention-OFF accuracy. If your number ≈ a constant baseline, you have measured nothing.
3. Correct-vs-wrong state check. Re-run with the mechanism's input SCRAMBLED (wrong
instance's state, shuffled, or noise), everything else identical. Report Δ = acc(correct) − acc(wrong) and the fraction of predictions that change when you swap. Δ≈0 / no change ⇒
the content is causally inert. Distinguish presence-effect (ON vs OFF changes output)
from content-effect (changing WHAT it carries changes output) — presence ≠ content.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 9d ago First seen · 84 lines · 101 tokens per session scan A 4c2d5b5f8b37
artifact-first-analysis is a skill published in the GitHub repository georgepok/local-llm-mcp-server (0 stars, last pushed 1mo ago), licensed MIT. It adds 101 tokens to every session and 1,216 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-01.
Other skills, from other repositories
arboreto
Infer gene regulatory networks (GRNs) from gene expression data using scalable algorithms (GRNBoost2, GENIE3). Use when analyzing transcriptomics data (bulk RNA-seq, single-cell RNA-seq) to identify transcription factor-target gene relationships and regulatory interactions. Supports distributed computation for…
pyhealth
Build clinical/healthcare deep-learning pipelines with PyHealth — loading EHR/signal/imaging datasets (MIMIC-III/IV, eICU, OMOP, SleepEDF, ChestXray14, EHRShot), defining tasks (mortality, readmission, length-of-stay, drug recommendation, sleep staging, ICD coding, EEG events), instantiating models (Transformer…
torchdrug
Build and troubleshoot TorchDrug 0.2.1 workflows for molecular graphs, property prediction, self-supervised pretraining, molecule generation, retrosynthesis, protein representation learning, and knowledge graph reasoning. Use when code imports torchdrug or needs its datasets, models, tasks, or Engine.
deepspot-m
Generate transcriptome-wide virtual spatial transcriptomics from H&E histology with DeepSpot-M. Use when you need spatial gene expression in log1p-CPM for 224x224 tiles at about 20x, want to query protein-coding genes by symbol instead of a fixed panel, or want to run prediction across a whole slide after tiling with…
nemo-mbridge-perf-expert-parallel-overlap
Validate and use MoE expert-parallel communication overlap in Megatron-Bridge, including overlapmoeexpertparallelcomm, delaywgradcompute, and flex dispatcher backends such as DeepEP and HybridEP.
pick-a-pii-model
Select an on-device OpenMed PII model from the committed registry by language, runtime format, and size budget, then require recall validation before deployment. Use when an agent must choose a local PII detector for CPU, Apple Silicon, or a mobile export without relying on live model discovery.