Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/openscientist-io/openscientist/kbase-querynpx skills add openscientist-io/openscientist --skill kbase-querygit clone --depth 1 https://github.com/openscientist-io/openscientistWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/openscientist-io/openscientist/kbase-query)<a href="https://agentmods.dev/skills/openscientist-io/openscientist/kbase-query"><img src="https://agentmods.dev/badge/skills/openscientist-io/openscientist/kbase-query.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00079 | $0.01082 |
| Opus 5 | $0.00039 | $0.00541 |
| Sonnet 5 | $0.00016 | $0.00216 |
| Haiku 4.5 | $0.00008 | $0.00108 |
Grade A, and why
kbase-query scanned grade A with 1 finding against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 5d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Makes network callslowCapability
Not a fault in itself. Listed so you know the mod talks to something, and to what.
resp = requests.post( How it starts
The opening of the file, as written. The whole thing — 150 lines — stays where its author put it; the contents beside it link to each section on GitHub.
KBase Query
Query the KBase/BERDL Datalake MCP Server via REST API using Python.
Setup
The KBASE_TOKEN environment variable must be set. Tokens expire after ~1 week.
Python API Functions
Use execute_code with these helper functions to query the lakehouse:
import os
import requests
import pandas as pd
KBASE_TOKEN = os.environ.get('KBASE_TOKEN')
BASE_URL = 'https://hub.berdl.kbase.us/apis/mcp'
def get_headers():
return {
'Authorization': f'Bearer {KBASE_TOKEN}',
'Content-Type': 'application/json',
'accept': 'application/json'
}
def list_databases():
'''List all available databases'''
resp = requests.post(
f'{BASE_URL}/delta/databases/list',
headers=get_headers(),
json={'use_hms': True, 'filter_by_namespace': True}
)
resp.raise_for_status()
return resp.json()['databases']
def list_tables(database):
'''List tables in a database'''
resp = requests.post(
f'{BASE_URL}/delta/databases/tables/list',
headers=get_headers(),
json={'database': database, 'use_hms': True}
)
resp.raise_for_status()
return resp.json()['tables']
def query(sql, limit=100):
'''Execute SQL query and return DataFrame'''
resp = requests.post(
f'{BASE_URL}/delta/tables/query',
headers=get_headers(),
json={'query': sql, 'limit': limit}
)
resp.raise_for_status()
data = resp.json()
rows = data.get('result', [])
return pd.DataFrame(rows)
Example Workflow
1. Explore available databases
dbs = list_databases()
print(f"Available databases: {dbs}")
# → ['enigma_coral', 'nmdc_core', 'globalusers_kepangenome_parquet_1', ...]
2. List tables in a database
tables = list_tables('nmdc_core')
print(f"Found {len(tables)} tables: {tables[:10]}")
# → ['annotation_terms_unified', 'cog_categories', 'kegg_ko_module', ...]
3. Query data with SQL
# Simple query
df = query("SELECT * FROM nmdc_core.kegg_ko_module LIMIT 10")
print(df)
# Aggregation query
df = query("""
SELECT module_id, COUNT(*) as ko_count
FROM nmdc_core.kegg_ko_module
GROUP BY module_id
ORDER BY ko_count DESC
LIMIT 20
""")
print(df)
# Join query (when needed)
df = query("""
SELECT a.*, b.description
FROM nmdc_core.kegg_ko_module a
JOIN nmdc_core.kegg_modules b ON a.module_id = b.module_id
LIMIT 10
""")
What ships with it
10 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
- references/api_reference.md 4.3 KB
- scripts/kbase_db_structure.sh 527 B runs code
- scripts/kbase_health.sh 308 B runs code
- scripts/kbase_list_databases.sh 420 B runs code
- scripts/kbase_list_tables.sh 527 B runs code
- scripts/kbase_query.sh 544 B runs code
- scripts/kbase_select.sh 676 B runs code
- scripts/kbase_table_count.sh 612 B runs code
- scripts/kbase_table_sample.sh 694 B runs code
- scripts/kbase_table_schema.sh 627 B runs code
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 5d ago First seen · 150 lines · 79 tokens per session scan A e643a4bf16c8
kbase-query is a skill published in the GitHub repository openscientist-io/openscientist (49 stars, last pushed yesterday), licensed Apache-2.0. It adds 79 tokens to every session and 1,082 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 1 finding (makes network calls). No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
lamindb
Use when working with LaminDB, the open-source lineage-native lakehouse for biological datasets and models. Covers setup, artifact registration, query/search, lineage tracking, validation, ontology-backed annotation with Bionty, collections, branches, storage, and workflow integrations.
tiledbvcf
Efficient storage and retrieval of genomic variant data using TileDB. Scalable VCF/BCF ingestion, incremental sample addition, compressed storage, parallel queries, and export capabilities for population genomics.
defining-cohort-phenotypes
Authors computable phenotype and cohort definitions in the OHDSI ATLAS / CIRCE style over the OMOP CDM, combining standard concept sets with NLP-derived features that OpenMed extracts. Use when the user wants to define a patient cohort, write a computable phenotype, reuse PheKB or OHDSI Phenotype Library logic, build…
benchling-integration
Benchling R&D platform integration. Access registry (DNA, proteins), inventory, ELN entries, workflows via API, build Benchling Apps, query Data Warehouse, for lab data management automation.
chembl-database
Query the ChEMBL database for bioactive molecules, drug targets, bioactivity data, approved drugs, and chemical structures. Use when the user asks about compounds, targets, IC50/Ki values, drug mechanisms, or structure searches.
nvalchemi-data-storage
How to write, read, compose, and load atomic data using nvalchemi's composable Zarr-backed storage pipeline (Writer, Reader, Dataset, MultiDataset, DataLoader). Use when saving simulation outputs or trajectories to disk, converting structures (e.g. ASE / extxyz) into Zarr stores, assembling datasets for training or…