automotive-dfm-benchmarking

automotive-dfm-benchmarking is a skill for Claude Code, Codex from pangzhenying2025/hermes-automotive-skills. It costs 26 tokens per session (2,409 once invoked), scanned A, original, MIT.

A framework for comparing autonomous-driving systems with models of how human drivers behave. It uses naturalistic driving data—records of real-world driving—to create human-performance baselines for different scenarios.

In plain words
What is it for?
Use it to generate driving scenarios, benchmark autonomous-driving performance, compare results with human-driver distributions, and support safety arguments.
Why use it?
It helps teams measure an automated-driving system against a defined human reference instead of judging performance without context.

Skill for Claude CodeCodex

Written for no agent in particular: nothing here depends on one.

Good fit Use it to generate driving scenarios, benchmark autonomous-driving performance, compare results with human-driver distributions, and support safety arguments.

Compare 6 skills from other repositories ↓
Install with agentmods
npx agentmods add skills/pangzhenying2025/hermes-automotive-skills/automotive-dfm-benchmarking
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Any agent
npx skills add pangzhenying2025/hermes-automotive-skills --skill automotive-dfm-benchmarking
Clone the repo
git clone --depth 1 https://github.com/pangzhenying2025/hermes-automotive-skills

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for automotive-dfm-benchmarking

README.md
[![agentmods](https://agentmods.dev/badge/skills/pangzhenying2025/hermes-automotive-skills/automotive-dfm-benchmarking/github.svg)](https://agentmods.dev/skills/pangzhenying2025/hermes-automotive-skills/automotive-dfm-benchmarking)
Your own site
<a href="https://agentmods.dev/skills/pangzhenying2025/hermes-automotive-skills/automotive-dfm-benchmarking"><img src="https://agentmods.dev/badge/skills/pangzhenying2025/hermes-automotive-skills/automotive-dfm-benchmarking/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for automotive-dfm-benchmarking

Your own site · 80×15
<a href="https://agentmods.dev/skills/pangzhenying2025/hermes-automotive-skills/automotive-dfm-benchmarking"><img src="https://agentmods.dev/badge/skills/pangzhenying2025/hermes-automotive-skills/automotive-dfm-benchmarking.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 26 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 2,409 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00026 $0.02409
Opus 5 $0.00013 $0.01205
Sonnet 5 $0.00005 $0.00482
Haiku 4.5 $0.00003 $0.00241

Measured 12d ago against content hash 904b44e7bd27, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-12, from the pricing page.

Security

Grade A, and why

automotive-dfm-benchmarking scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 12d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/automotive-dfm-benchmarking/SKILL.md · 317 lines

How it starts

The opening of the file, as written. The whole thing — 317 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Automotive Dfm Benchmarking

Dfm Benchmarking

DFM Benchmarking — Driver Foundation Model Framework for AD Evaluation

Overview

Benchmarking framework based on the Driver Foundation Model (DFM) concept for evaluating autonomous driving systems. DFM uses large-scale naturalistic driving data (NDD) to model human driver behavior distributions, providing a human-performance baseline for AD system evaluation. This skill supports scenario generation, performance benchmarking, and safety argument construction using NDD-derived metrics.

DFM Concept

驾驶员基础模型 (Driver Foundation Model) 概念
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

Core Idea:
  Human drivers provide a safety baseline:
  - Average driver: ~1 fatality per 10^8 km (developed countries)
  - Good driver: ~10x safer than average
  - AD must be at least as safe as good human driver

DFM Approach:
  1. Collect large-scale NDD (7.5M+ aerial trajectories)
  2. Model human driving behavior distributions
  3. Extract scenario-specific performance baselines
  4. Benchmark AD systems against human baselines
  5. Quantify relative safety improvement

DFM as Foundation Model:
  ├── Pre-trained on massive NDD
  ├── Captures diverse driving styles and conditions
  ├── Fine-tunable for specific scenarios/regions
  ├── Provides probabilistic behavior predictions
  └── Serves as benchmark generator and evaluator
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━

Data Foundation

NDD Collection & Processing

# NDD Processing Pipeline for DFM
class NDDProcessor:
    """
    Process Naturalistic Driving Data for DFM benchmarking.
    Supports aerial trajectory data (drone-based) and fleet data.
    """

    def __init__(self, data_source: str):
        """
        data_source options:
        - "aerial": Drone-based trajectory extraction (7.5M+ trajectories)
        - "fleet": Vehicle-mounted sensor data
        - "hybrid": Combined aerial + fleet data
        """
        self.source = data_source

    def extract_driving_primitives(self, trajectories):
        """
        Extract fundamental driving behaviors from trajectory data.

        Driving primitives:
        - Car-following (跟车)
        - Lane-changing (换道)
        - Merging (汇入)
        - Diverging (分流)
        - Crossing (交叉)
        - Free-driving (自由行驶)
        """
        primitives = {
            "car_following": self.extract_car_following(trajectories),
            "lane_change": self.extract_lane_changes(trajectories),
            "merge": self.extract_merges(trajectories),
            "diverge": self.extract_diverges(trajectories),
            "crossing": self.extract_crossings(trajectories),
            "free_driving": self.extract_free_driving(trajectories),
        }
        return primitives

    def build_behavior_distributions(self, primitives):
        """
        Build statistical distributions of driving behaviors.

        For car-following:
        - Time headway distribution: P(THW)
        - TTC distribution: P(TTC)
        - Speed distribution: P(v | context)
        - Acceleration distribution: P(a | context)
        - Lane offset distribution: P(offset | context)
        """
        distributions = {}
        for primitive_type, data in primitives.items():
            distributions[primitive_type] = {
                "thw": fit_distribution(data.thw_values),
                "ttc": fit_distribution(data.ttc_values),
                "speed": conditional_distribution(data.speeds, data.contexts),
                "acceleration": conditional_distribution(data.accels, data.contexts),
                "lateral_offset": fit_distribution(data.offsets),
                "jerk": fit_distribution(data.jerks),
            }
        return distributions

    def generate_benchmark_scenarios(self, distributions, n_scenarios=1000):
        """
        Generate benchmark scenarios by sampling from behavior distributions.

        Importance sampling: over-sample from tail (critical) regions
        """
        scenarios = []
        for i in range(n_scenarios):
            # Sample scenario type based on exposure
            scenario_type = sample_weighted(distributions.keys(),
                                           weights=exposure_weights)
            # Sample parameters from distribution
            params = sample_from_distribution(
                distributions[scenario_type],
                sampling="importance",  # over-sample tails
                criticality_weight=2.0
            )
            scenarios.append(BenchmarkScenario(
                type=scenario_type,
                parameters=params,
                human_baseline=distributions[scenario_type],
            ))
        return scenarios

Read the full file on GitHub · 317 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 12d ago First seen · 317 lines · 26 tokens per session scan A 904b44e7bd27

Subscribe to this mod's changes

automotive-dfm-benchmarking is a skill published in the GitHub repository pangzhenying2025/hermes-automotive-skills (5 stars, last pushed 3mo ago), licensed MIT. It adds 26 tokens to every session and 2,409 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

iso26262

ISO 26262 functional-safety expert that operates in two modes: (1) HARA / ASIL determination — enumerate hazardous events from item malfunctions × driving situations, rate Severity (S0–S3), Exposure (E0–E4), Controllability (C0–C3), look up ASIL from ISO 26262-3:2018 Table 4, and produce a HARA report with Safety…

ptsilivis/autonomousguy · 197 tokens

automotive-syseng

When the user wants to analyze automotive requirements, check INCOSE/EARS compliance, review MISRA-C code, assess ADAS levels, or verify ISO 26262/AUTOSAR/SOTIF conformance. Also use when the user says 'check requirements', 'EARS check', 'INCOSE analysis', 'MISRA check', 'ASIL assessment', 'V-model check'…

duonghvu/automotive-syseng · 122 tokens

automotive-expert

Expert-level automotive systems, connected vehicles, fleet management, telematics, ADAS, and automotive software. Use when the user mentions connected car, fleet, telematics, ADAS, or vehicle, or when the task involves Automotive Systems, Technologies, Standards and Protocols, or Fleet Management.

personamanagmentlayer/pcl · 64 tokens

misra

MISRA C:2025 expert that operates in two modes: (1) Review — scan existing C code for violations across all 223 guidelines (22 directives + 201 rules), report findings with rule IDs, corrected code, and deviation justification templates; (2) Develop — generate new C functions, modules, or data structures that are…

ptsilivis/autonomousguy · 221 tokens

requirements

Requirements-engineering expert that operates in three modes: (1) Elicitation — extract atomic, testable requirements from briefs, meeting notes, or system specs using EARS notation, with full attribute set (ID, type, priority, ASIL, verification method, source), flagging ambiguities as open questions; (2) Refinement…

ptsilivis/autonomousguy · 221 tokens

instrument-data-to-allotrope

Convert laboratory instrument output files (PDF, CSV, Excel, TXT) to Allotrope Simple Model (ASM) JSON format or flattened 2D CSV. Use this skill when scientists need to standardize instrument data for LIMS systems, data lakes, or downstream analysis. Supports auto-detection of instrument types. Outputs include full…

anthropics/knowledge-work-plugins · 123 tokens