Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add pangzhenying2025/hermes-automotive-skills --skill automotive-dfm-benchmarkinggit clone --depth 1 https://github.com/pangzhenying2025/hermes-automotive-skillsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/pangzhenying2025/hermes-automotive-skills/automotive-dfm-benchmarking)<a href="https://agentmods.dev/skills/pangzhenying2025/hermes-automotive-skills/automotive-dfm-benchmarking"><img src="https://agentmods.dev/badge/skills/pangzhenying2025/hermes-automotive-skills/automotive-dfm-benchmarking/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/pangzhenying2025/hermes-automotive-skills/automotive-dfm-benchmarking"><img src="https://agentmods.dev/badge/skills/pangzhenying2025/hermes-automotive-skills/automotive-dfm-benchmarking.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00026 | $0.02409 |
| Opus 5 | $0.00013 | $0.01205 |
| Sonnet 5 | $0.00005 | $0.00482 |
| Haiku 4.5 | $0.00003 | $0.00241 |
Grade A, and why
automotive-dfm-benchmarking scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 12d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 317 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Automotive Dfm Benchmarking
Dfm Benchmarking
DFM Benchmarking — Driver Foundation Model Framework for AD Evaluation
Overview
Benchmarking framework based on the Driver Foundation Model (DFM) concept for evaluating autonomous driving systems. DFM uses large-scale naturalistic driving data (NDD) to model human driver behavior distributions, providing a human-performance baseline for AD system evaluation. This skill supports scenario generation, performance benchmarking, and safety argument construction using NDD-derived metrics.
DFM Concept
驾驶员基础模型 (Driver Foundation Model) 概念
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Core Idea:
Human drivers provide a safety baseline:
- Average driver: ~1 fatality per 10^8 km (developed countries)
- Good driver: ~10x safer than average
- AD must be at least as safe as good human driver
DFM Approach:
1. Collect large-scale NDD (7.5M+ aerial trajectories)
2. Model human driving behavior distributions
3. Extract scenario-specific performance baselines
4. Benchmark AD systems against human baselines
5. Quantify relative safety improvement
DFM as Foundation Model:
├── Pre-trained on massive NDD
├── Captures diverse driving styles and conditions
├── Fine-tunable for specific scenarios/regions
├── Provides probabilistic behavior predictions
└── Serves as benchmark generator and evaluator
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Data Foundation
NDD Collection & Processing
# NDD Processing Pipeline for DFM
class NDDProcessor:
"""
Process Naturalistic Driving Data for DFM benchmarking.
Supports aerial trajectory data (drone-based) and fleet data.
"""
def __init__(self, data_source: str):
"""
data_source options:
- "aerial": Drone-based trajectory extraction (7.5M+ trajectories)
- "fleet": Vehicle-mounted sensor data
- "hybrid": Combined aerial + fleet data
"""
self.source = data_source
def extract_driving_primitives(self, trajectories):
"""
Extract fundamental driving behaviors from trajectory data.
Driving primitives:
- Car-following (跟车)
- Lane-changing (换道)
- Merging (汇入)
- Diverging (分流)
- Crossing (交叉)
- Free-driving (自由行驶)
"""
primitives = {
"car_following": self.extract_car_following(trajectories),
"lane_change": self.extract_lane_changes(trajectories),
"merge": self.extract_merges(trajectories),
"diverge": self.extract_diverges(trajectories),
"crossing": self.extract_crossings(trajectories),
"free_driving": self.extract_free_driving(trajectories),
}
return primitives
def build_behavior_distributions(self, primitives):
"""
Build statistical distributions of driving behaviors.
For car-following:
- Time headway distribution: P(THW)
- TTC distribution: P(TTC)
- Speed distribution: P(v | context)
- Acceleration distribution: P(a | context)
- Lane offset distribution: P(offset | context)
"""
distributions = {}
for primitive_type, data in primitives.items():
distributions[primitive_type] = {
"thw": fit_distribution(data.thw_values),
"ttc": fit_distribution(data.ttc_values),
"speed": conditional_distribution(data.speeds, data.contexts),
"acceleration": conditional_distribution(data.accels, data.contexts),
"lateral_offset": fit_distribution(data.offsets),
"jerk": fit_distribution(data.jerks),
}
return distributions
def generate_benchmark_scenarios(self, distributions, n_scenarios=1000):
"""
Generate benchmark scenarios by sampling from behavior distributions.
Importance sampling: over-sample from tail (critical) regions
"""
scenarios = []
for i in range(n_scenarios):
# Sample scenario type based on exposure
scenario_type = sample_weighted(distributions.keys(),
weights=exposure_weights)
# Sample parameters from distribution
params = sample_from_distribution(
distributions[scenario_type],
sampling="importance", # over-sample tails
criticality_weight=2.0
)
scenarios.append(BenchmarkScenario(
type=scenario_type,
parameters=params,
human_baseline=distributions[scenario_type],
))
return scenarios
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 12d ago First seen · 317 lines · 26 tokens per session scan A 904b44e7bd27
automotive-dfm-benchmarking is a skill published in the GitHub repository pangzhenying2025/hermes-automotive-skills (5 stars, last pushed 3mo ago), licensed MIT. It adds 26 tokens to every session and 2,409 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
iso26262
ISO 26262 functional-safety expert that operates in two modes: (1) HARA / ASIL determination — enumerate hazardous events from item malfunctions × driving situations, rate Severity (S0–S3), Exposure (E0–E4), Controllability (C0–C3), look up ASIL from ISO 26262-3:2018 Table 4, and produce a HARA report with Safety…
automotive-syseng
When the user wants to analyze automotive requirements, check INCOSE/EARS compliance, review MISRA-C code, assess ADAS levels, or verify ISO 26262/AUTOSAR/SOTIF conformance. Also use when the user says 'check requirements', 'EARS check', 'INCOSE analysis', 'MISRA check', 'ASIL assessment', 'V-model check'…
automotive-expert
Expert-level automotive systems, connected vehicles, fleet management, telematics, ADAS, and automotive software. Use when the user mentions connected car, fleet, telematics, ADAS, or vehicle, or when the task involves Automotive Systems, Technologies, Standards and Protocols, or Fleet Management.
misra
MISRA C:2025 expert that operates in two modes: (1) Review — scan existing C code for violations across all 223 guidelines (22 directives + 201 rules), report findings with rule IDs, corrected code, and deviation justification templates; (2) Develop — generate new C functions, modules, or data structures that are…
requirements
Requirements-engineering expert that operates in three modes: (1) Elicitation — extract atomic, testable requirements from briefs, meeting notes, or system specs using EARS notation, with full attribute set (ID, type, priority, ASIL, verification method, source), flagging ambiguities as open questions; (2) Refinement…
instrument-data-to-allotrope
Convert laboratory instrument output files (PDF, CSV, Excel, TXT) to Allotrope Simple Model (ASM) JSON format or flattened 2D CSV. Use this skill when scientists need to standardize instrument data for LIMS systems, data lakes, or downstream analysis. Supports auto-detection of instrument types. Outputs include full…