Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add obielin/responsible-ai-skills --skill bias-assessmentgit clone --depth 1 https://github.com/obielin/responsible-ai-skillsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/obielin/responsible-ai-skills/bias-assessment)<a href="https://agentmods.dev/skills/obielin/responsible-ai-skills/bias-assessment"><img src="https://agentmods.dev/badge/skills/obielin/responsible-ai-skills/bias-assessment/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/obielin/responsible-ai-skills/bias-assessment"><img src="https://agentmods.dev/badge/skills/obielin/responsible-ai-skills/bias-assessment.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00038 | $0.01151 |
| Opus 5 | $0.00019 | $0.00575 |
| Sonnet 5 | $0.00008 | $0.00230 |
| Haiku 4.5 | $0.00004 | $0.00115 |
Grade A, and why
bias-assessment scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 12d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 151 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Bias Assessment
Bias in AI systems causes real harm. A model that appears accurate overall can systematically disadvantage specific groups. You MUST complete this assessment before training or evaluating any model.
Phase 1: Data Audit (Before Training)
1.1 Check Representation
Run the representation check script:
python skills/bias-assessment/scripts/check_representation.py --data <your_dataset>
If no script applies, manually verify:
# For each protected attribute in your dataset:
for attr in ['age', 'sex', 'ethnicity', 'disability', 'postcode']:
if attr in df.columns:
print(f"\n{attr} distribution:")
print(df[attr].value_counts(normalize=True))
# Flag underrepresented groups (<5% of dataset)
underrepresented = df[attr].value_counts(normalize=True)
flagged = underrepresented[underrepresented < 0.05].index.tolist()
if flagged:
print(f"⚠️ UNDERREPRESENTED: {flagged}")
Stop and fix if: Any group that will be affected by predictions has <5% representation.
1.2 Check for Proxy Variables
Proxy variables appear neutral but encode protected characteristics:
| Proxy Variable | May Encode |
|---|---|
| Postcode / ZIP code | Ethnicity, deprivation |
| Name | Ethnicity, sex |
| School attended | Socioeconomic status, ethnicity |
| Job title history | Sex, disability |
| Device type | Socioeconomic status |
Action: For each proxy variable, decide: remove it, transform it, or document the risk explicitly.
1.3 Check Label Quality
Biased labels produce biased models:
- Were labels assigned by humans? → Check inter-annotator agreement across annotator demographics
- Were labels derived from historical decisions? → Those decisions may contain historical bias
- Are labels consistent across demographic groups? → Run:
df.groupby(protected_attr)['label'].mean()
Phase 2: Model Evaluation (After Training)
2.1 Disaggregated Performance
What ships with it
1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 12d ago First seen · 151 lines · 38 tokens per session scan A 744ac5d52a9f
bias-assessment is a skill published in the GitHub repository obielin/responsible-ai-skills (2 stars, last pushed 5mo ago), licensed MIT. It adds 38 tokens to every session and 1,151 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
embed
Generate, inspect, and use node/text embeddings in Semantica — compute Node2Vec embeddings, find similar nodes, score link predictions, batch similarity, and pairwise similarity. Uses NodeEmbedder, SimilarityCalculator, LinkPredictor, and AgentContext. Sub-commands: compute, similar, similarity, predict-link…
visualize
Visualize the Semantica knowledge graph — topology, centrality, communities, paths, embeddings, decision insights, and temporal evolution. Uses GraphAnalyzer, CentralityCalculator, CommunityDetector, PathFinder, and ContextGraph analytics. Sub-commands: topology, centrality, community, path, decision-graph, insights…
reason
Run reasoning over the Semantica knowledge graph — deductive logic, abductive hypothesis generation, Datalog programs, SPARQL queries, Rete network evaluation. Uses DeductiveReasoner, AbductiveReasoner, DatalogReasoner, SPARQLReasoner, ReteEngine. Sub-commands: deductive, abductive, datalog, sparql, rete, prove…
temporal
Temporal graph operations on Semantica — scoped queries at a point in time, graph snapshots, node change timelines, temporal causal analysis, and graph state reconstruction. Uses AgentContext.findprecedents(asof=), ContextGraph.stateat(), CausalChainAnalyzer.traceattime(), and TemporalQueryRewriter. Sub-commands…
validate
Validate Semantica pipelines, extraction quality, graph schemas, and ontology consistency. Returns structured error/warning checklists. Uses PipelineValidator, PipelineBuilder.validatepipeline(), GraphValidator, and OntologyValidator. Sub-commands: pipeline, step, dependencies, extraction, graph, ontology, performance.
extract
Run the full Semantica semantic extraction pipeline on a file or selected text — NER, relations, events, coreference resolution, triplets, and validation. Clears result cache before each run. Returns Markdown tables with entity/relation/event/triplet results and inline validator warnings.