ai-training-lawfulness

ai-training-lawfulness is a skill for Claude Code from mukul975/Privacy-Data-Protection-Skills. It costs 77 tokens per session (2,929 once invoked), scanned A, original, Apache-2.0.

A GDPR assessment guide for deciding whether personal data may legally be used to train AI models. GDPR is the European Union’s data-protection law, and the EDPB is the EU body that helps interpret it.

In plain words
What is it for?
Use it to assess the legal basis for training data, apply a legitimate-interest balancing test, review consent problems, and examine whether scraped or third-party datasets can be used.
Why use it?
It helps separate AI training from other data uses and identify issues with consent, legitimate interest, public datasets, and web-scraped data.

Skill for Claude Code

Written for Claude Code: shipped in a Claude Code plugin.

Part of the ai-privacy-governance-skills plugin — 15 skills shipped together

Good fit Use it to assess the legal basis for training data, apply a legitimate-interest balancing test, review consent problems, and examine whether scraped or third-party datasets can be used.

Compare 6 skills from other repositories ↓
Install with agentmods
npx agentmods add skills/mukul975/privacy-data-protection-skills/ai-training-lawfulness
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Any agent
npx skills add mukul975/Privacy-Data-Protection-Skills --skill ai-training-lawfulness
Clone the repo
git clone --depth 1 https://github.com/mukul975/Privacy-Data-Protection-Skills

Made for: Claude Code.

Or install ai-privacy-governance-skills, the plugin that ships this one along with the rest of its 15 skills.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for ai-training-lawfulness

README.md
[![agentmods](https://agentmods.dev/badge/skills/mukul975/privacy-data-protection-skills/ai-training-lawfulness/github.svg)](https://agentmods.dev/skills/mukul975/privacy-data-protection-skills/ai-training-lawfulness)
Your own site
<a href="https://agentmods.dev/skills/mukul975/privacy-data-protection-skills/ai-training-lawfulness"><img src="https://agentmods.dev/badge/skills/mukul975/privacy-data-protection-skills/ai-training-lawfulness/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for ai-training-lawfulness

Your own site · 80×15
<a href="https://agentmods.dev/skills/mukul975/privacy-data-protection-skills/ai-training-lawfulness"><img src="https://agentmods.dev/badge/skills/mukul975/privacy-data-protection-skills/ai-training-lawfulness.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 77 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 2,929 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00077 $0.02929
Opus 5 $0.00039 $0.01465
Sonnet 5 $0.00015 $0.00586
Haiku 4.5 $0.00008 $0.00293

Measured 13d ago against content hash ffe41e03d667, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-12, from the pricing page.

Security

Grade A, and why

ai-training-lawfulness scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 13d ago.

The scan reads SKILL.md. This mod also ships 1 executable file (scripts/process.py), listed below but not scanned — reading those needs a real analyzer, not pattern matching.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

Origin

Copies of this mod

1 near-identical copy found in the catalogue:

plugins/ai-privacy-governance-skills/skills/ai-training-lawfulness/SKILL.md · 226 lines

How it starts

The opening of the file, as written. The whole thing — 226 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Lawful Basis for AI Training Data

Overview

The processing of personal data for AI model training constitutes a distinct processing operation requiring its own lawful basis under GDPR Art. 6(1). The EDPB Guidelines 04/2025 and the coordinated ChatGPT Taskforce findings establish that AI training creates unique lawful basis challenges: the scale of data collection, the difficulty of obtaining meaningful consent for open-ended AI training purposes, the tension between legitimate interest and data subject expectations, and the complexity of determining lawfulness for web-scraped and third-party datasets. This skill provides the comprehensive lawful basis assessment framework for AI training data processing, addressing each Art. 6(1) basis as applied to ML training contexts.

Fundamental Principles

AI Training as Personal Data Processing

The EDPB has confirmed that AI model training constitutes processing of personal data under Art. 4(2) GDPR when:

  1. Training datasets contain personal data (directly or indirectly identifiable natural persons)
  2. The model is trained on data that includes personal data, even if the intent is to learn general patterns
  3. The resulting model retains the capability to generate or reproduce personal data from training sets
  4. Personal data is used in any pipeline stage: collection, cleaning, annotation, augmentation, validation, testing

The controller cannot avoid GDPR obligations by claiming the model has "learned" rather than "stored" personal data. The processing occurs at the point of training, regardless of whether the model can later reproduce specific records.

Purpose Specification for AI Training

Art. 5(1)(b) requires that personal data be collected for specified, explicit, and legitimate purposes. For AI training, this means:

  • "Training an AI model" is insufficiently specific — the controller must articulate the specific capability being developed
  • "Improving our services" through AI training must be disaggregated into concrete purposes
  • Each purpose must be documented before training begins, not retroactively justified
  • The purpose must be communicated to data subjects in privacy notices per Arts. 13-14

Read the full file on GitHub · 226 lines

Files

What ships with it

4 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 13d ago First seen · 226 lines · 77 tokens per session scan A ffe41e03d667

Subscribe to this mod's changes

ai-training-lawfulness is a skill published in the GitHub repository mukul975/Privacy-Data-Protection-Skills (272 stars, last pushed 5mo ago), licensed Apache-2.0. It adds 77 tokens to every session and 2,929 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

privacy-compliance

Comprehensive global privacy compliance agent skill covering GDPR, CCPA/CPRA, HIPAA Privacy Rule, EU AI Act, LGPD, cross-border data transfer mechanisms (SCCs, BCRs, EU-US DPF), PII identification and classification, data minimization, consent management, privacy-by-design patterns, DPIA workflows, data subject access…

JPeetz/agent-skills · 132 tokens

fedramp

Expert guidance for FedRAMP certification and compliance under CR26 (FedRAMP Consolidated Rules for 2026). Use this skill whenever a user asks about FedRAMP authorization, ATO (Authority to Operate), cloud security for federal government, NIST SP 800-53 controls, CSP compliance, or any of the core FedRAMP document…

Sushegaad/Claude-Skills-Governance-Risk-and-Compliance · 236 tokens

gdpr-compliance

Expert GDPR compliance assistant covering all four core workflows: (1) auditing code and systems for GDPR violations, (2) drafting GDPR-compliant documents such as privacy policies, Data Processing Agreements (DPAs), and consent notices, (3) answering GDPR compliance questions with authoritative article citations, and…

Sushegaad/Claude-Skills-Governance-Risk-and-Compliance · 176 tokens

iso42001

Expert ISO 42001 AI Management System (AIMS) compliance advisor. Use this skill whenever a user asks about ISO/IEC 42001:2023, AI governance, AI management systems, AI risk assessment, AI system impact assessment, Annex A controls for AI, Statement of Applicability for AI systems, AI policy, responsible AI, AI…

Sushegaad/Claude-Skills-Governance-Risk-and-Compliance · 173 tokens

eu-cra

Expert EU Cyber Resilience Act (CRA) advisor for Regulation (EU) 2024/2847 — mandatory cybersecurity and vulnerability handling requirements for all products with digital elements (PDEs) sold in the EU. Use this skill for gap analysis, product classification (Default / Class I / Class II), conformity assessment route…

Sushegaad/Claude-Skills-Governance-Risk-and-Compliance · 133 tokens

nist-800-53

NIST SP 800-53 Rev 5 compliance advisor — all 20 control families (AC, AT, AU, CA, CM, CP, IA, IR, MA, MP, PE, PL, PM, PS, PT, RA, SA, SC, SI, SR), Low/Moderate/High baseline selection, FIPS 199/200 system categorization, control tailoring and overlays, privacy controls (PT family), supply chain risk management (SR…

Sushegaad/Claude-Skills-Governance-Risk-and-Compliance · 177 tokens