Borrowing it
Nothing to install: this file belongs to 45ck/open-genome-agent. Take a copy, put it at the same path in your own repository, and replace the rules that are about this project with yours.
curl -O https://raw.githubusercontent.com/45ck/open-genome-agent/main/.agents/skills/ingest-vcf/SKILL.mdgit clone --depth 1 https://github.com/45ck/open-genome-agentWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/45ck/open-genome-agent/ingest-vcf)<a href="https://agentmods.dev/skills/45ck/open-genome-agent/ingest-vcf"><img src="https://agentmods.dev/badge/skills/45ck/open-genome-agent/ingest-vcf/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/45ck/open-genome-agent/ingest-vcf"><img src="https://agentmods.dev/badge/skills/45ck/open-genome-agent/ingest-vcf.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00036 | $0.00349 |
| Opus 5 | $0.00018 | $0.00175 |
| Sonnet 5 | $0.00007 | $0.00070 |
| Haiku 4.5 | $0.00004 | $0.00035 |
Grade A, and why
ingest-vcf scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
Ingest VCF
Validate and inventory VCF or BCF input files before downstream analysis. Use when a user provides a genome variant file or asks what data is inside it.
When to use
- a VCF, VCF.GZ, BCF, TBI, or CSI file arrives
- you need sample names, header metadata, or file hashes
Do not use when
- the workflow starts from FASTQ only
Expected outputs
run_manifest.jsonsample_summary.json
Goal
Create the first trustworthy manifest for a supplied VCF/BCF input.
Procedure
- Identify the file kind:
vcf,vcf.gz, orbcf. - Check for the expected index (
.tbior.csi) where relevant. - Extract sample names and header metadata without modifying the raw input.
- Hash the input files and record them in
run_manifest.json. - Create a minimal
sample_summary.jsonwith unresolved fields set tonull.
Guardrails
- Never rewrite the source file in place.
- Do not guess the reference build from chromosome names alone if the evidence is weak.
- Multi-sample files require explicit sample-selection logic.
Escalate when
- the file has no samples
- the header is malformed
- the file is compressed but unindexed
References
See references/README.md for durable notes and scripts/ for deterministic helpers.
What ships with it
2 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 9d ago First seen · 48 lines · 36 tokens per session scan A 4c6a8a25595c
ingest-vcf is a skill published in the GitHub repository 45ck/open-genome-agent (4 stars, last pushed 2mo ago), licensed MIT. It adds 36 tokens to every session and 349 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
biomcp
Search and retrieve biomedical data - genes, variants, clinical trials, diagnostic tests, articles, drugs, diseases, pathways, proteins, adverse events, pharmacogenomics, and phenotype-disease matching. Use for gene function, variant pathogenicity, trials, diagnostics, drug safety, pathway context, disease workups…
biomcp-research
Do biomedical literature and variant research with the BioMCP CLI, and file what you learn about the tool itself as issues in the biomcp repo.
cellxgene-census-query
Query CZ CELLxGENE Census (61M+ cells). Filter by cell type/tissue/disease, retrieve expression data, and integrate with scanpy/PyTorch for population-scale single-cell analysis. Use this skill when: (1) Querying single-cell expression data by cell type, tissue, or disease, (2) Exploring available single-cell datasets…
biological-expert
Expert-level biology, biotechnology, genetics, bioinformatics, and computational biology. Use when the user mentions biology, biotechnology, genetics, bioinformatics, or genomics, or when the task involves Molecular Biology, Genomics & Bioinformatics, Systems Biology, or Data Analysis.
genomics-alignment
Load when computing alignment QC metrics (mapping rate, MAPQ distribution, insert size, duplicate rate, proper-pair rate) from a SAM or BAM file produced by any short-/long-read aligner (BWA / Bowtie2 / Minimap2). Skip when running the alignment step itself; only FASTQ-level QC is needed (use genomics-qc).
genomics-assembly
Load when computing genome-assembly QC metrics — N50/N90, L50/L90, total length, contig count, GC content, longest-contig — from a FASTA produced by any assembler (SPAdes / Megahit / Flye / Canu). Skip when running the assembly itself; assessing alignment quality (use genomics-alignment).