Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add ma-compbio-lab/SkillFoundry --skill protein-language-model-function-analysis-startergit clone --depth 1 https://github.com/ma-compbio-lab/SkillFoundryWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/ma-compbio-lab/skillfoundry/protein-language-model-function-analysis-starter)<a href="https://agentmods.dev/skills/ma-compbio-lab/skillfoundry/protein-language-model-function-analysis-starter"><img src="https://agentmods.dev/badge/skills/ma-compbio-lab/skillfoundry/protein-language-model-function-analysis-starter.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00000 | $0.00420 |
| Opus 5 | $0.00000 | $0.00210 |
| Sonnet 5 | $0.00000 | $0.00084 |
| Haiku 4.5 | $0.00000 | $0.00042 |
Grade A, and why
protein-language-model-function-analysis-starter scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 35 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Protein Language Model Function Analysis Starter
Use this skill to validate protein FASTA inputs, extract deterministic smoke-safe embeddings, and run a reusable sequence-to-function triage flow that can later be swapped onto real ESM-2 or ProtT5 backends.
What This Skill Does
- reads protein sequences from FASTA and rejects unsupported residue symbols
- emits per-sequence embeddings to a TSV contract
- runs centroid-based function analysis when labels are supplied
- reports nearest-neighbor structure in embedding space
- preserves a stable CLI for real
transformersbackends such as ESM-2 and ProtT5
When To Use It
- when you need a concrete local skill for
protein-embeddingsandsequence-to-function-modeling - when you want a deterministic smoke path before moving to GPU-backed protein language model inference
- when you need a compact handoff artifact for downstream DeepFRI, annotation, or benchmarking work
Run
python3 skills/proteomics/protein-language-model-function-analysis-starter/scripts/run_protein_language_model_function_analysis.py \
--input skills/proteomics/protein-language-model-function-analysis-starter/examples/toy_sequences.fasta \
--labels skills/proteomics/protein-language-model-function-analysis-starter/examples/toy_labels.tsv \
--config skills/proteomics/protein-language-model-function-analysis-starter/examples/analysis_config.json \
--embeddings-out scratch/protein-lm/toy_embeddings.tsv \
--summary-out scratch/protein-lm/toy_summary.json
Notes
- The default
mockbackend is intentionally deterministic and test-friendly. It preserves the same file contract as a real protein language model run. - For live model inference, switch the backend to
transformersand pointmodel_idat an ESM-2 or ProtT5 checkpoint. Seerefs.mdfor canonical sources and formatting notes. - The bundled toy labels are illustrative and suitable only for smoke testing or pipeline scaffolding, not biological claims.
What ships with it
13 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
- assets/README.md 92 B
- assets/toy_protein_lm_summary.json 6.6 KB
- assets/toy_protein_lm_summary.tsv 2.0 KB
- examples/analysis_config.json 239 B
- examples/README.md 94 B
- examples/resource_context.json 723 B
- examples/toy_labels.tsv 211 B
- examples/toy_sequences.fasta 317 B
- metadata.yaml 2.4 KB
- refs.md 2.0 KB
- scripts/run_frontier_starter.py 1.3 KB runs code
- scripts/run_protein_language_model_function_analysis.py 16 KB runs code
- tests/test_protein_language_model_function_analysis_starter.py 3.7 KB runs code
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 4d ago First seen · 35 lines · 0 tokens per session scan A 6a51e28c2896
protein-language-model-function-analysis-starter is a skill published in the GitHub repository ma-compbio-lab/SkillFoundry (38 stars, last pushed 4mo ago), licensed Apache-2.0. It costs nothing until one of its globs matches a file; then it loads 420 tokens. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-03.
Other skills, from other repositories
esmfold2
Biohub ESMFold2 / ESMFold2-Fast all-atom co-folding (Candido et al. 2026, github.com/Biohub/esm). Single-sequence and MSA modes; protein, DNA, RNA, ligand (CCD/SMILES), modified residues. FoldBench Ab-Ag 50-55%, PPI 70-77% DockQ-pass. Also covers the ESMC-{300M,600M,6B} protein language models from the same release…
scgpt
Embed and annotate single-cell expression data with scGPT, a foundation model for single-cell biology. Use this skill when: (1) Producing cell embeddings from an AnnData for clustering/integration, (2) Zero-shot or fine-tuned cell-type annotation, (3) Gene-level representation for perturbation/GRN tasks. For…
evo2
Score, embed, and generate DNA sequences with Evo 2, a long-context genomic foundation model. Use this skill when: (1) Computing per-nucleotide or per-sequence likelihoods for variant effect scoring, (2) Embedding genomic windows for downstream classification, (3) Generating DNA conditioned on a prefix, (4) Scoring…
cv-classification
Best practices for image classification tasks. Use when working on CIFAR, ImageNet, or other classification benchmarks.
cv-detection
Best practices for object detection tasks. Use when working on COCO, VOC, or detection architectures like YOLO and DETR.
experimental-design
Best practices for designing reproducible ML experiments. Use when planning ablations, baselines, or controlled experiments.