andrej-karpathy-perspective

andrej-karpathy-perspective is a skill for Claude Code, Codex from alchaincyf/nuwa-skill. It costs 234 tokens per session (8,824 once invoked), scanned A, original, MIT.

An AI perspective guide based on Andrej Karpathy's public writing, interviews, posts, and projects. It applies his engineering-focused way of thinking to AI, learning, products, and technology trends.

In plain words
What is it for?
Use it to assess AI product reliability, discuss language-model limits, plan learning, evaluate training methods, and think through AI product or industry decisions.
Why use it?
It helps separate practical engineering limits from demos, hype, or vague claims. It also states when a question falls outside the guide's covered areas or research date.

Skill for Claude CodeCodex

Written for no agent in particular: nothing here depends on one. Also seen: mentions Claude Code.

Good fit Use it to assess AI product reliability, discuss language-model limits, plan learning, evaluate training methods, and think through AI product or industry decisions.

Compare 6 skills from other repositories ↓
Install with agentmods
npx agentmods add skills/alchaincyf/nuwa-skill/andrej-karpathy-perspective
About the project

nuwa-skill is an Agent Skills-compatible tool that researches a named person and turns their thinking patterns into reusable guidance for an AI agent. It is for using someone’s mental models, decision heuristics, communication style, boundaries, and limitations when analyzing questions. The catalogue entries are skills that let compatible coding agents use this workflow.

alchaincyf/nuwa-skill · 32,370 stars · on GitHub

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Any agent
npx skills add alchaincyf/nuwa-skill --skill andrej-karpathy-perspective
Clone the repo
git clone --depth 1 https://github.com/alchaincyf/nuwa-skill

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for andrej-karpathy-perspective

README.md
[![agentmods](https://agentmods.dev/badge/skills/alchaincyf/nuwa-skill/andrej-karpathy-perspective/github.svg)](https://agentmods.dev/skills/alchaincyf/nuwa-skill/andrej-karpathy-perspective)
Your own site
<a href="https://agentmods.dev/skills/alchaincyf/nuwa-skill/andrej-karpathy-perspective"><img src="https://agentmods.dev/badge/skills/alchaincyf/nuwa-skill/andrej-karpathy-perspective/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for andrej-karpathy-perspective

Your own site · 80×15
<a href="https://agentmods.dev/skills/alchaincyf/nuwa-skill/andrej-karpathy-perspective"><img src="https://agentmods.dev/badge/skills/alchaincyf/nuwa-skill/andrej-karpathy-perspective.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 234 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 8,824 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe. ✓ AI security review Sonnet 5 · 6 Sept 2026 📄 Read the review Third-party audits
  • NVIDIA SkillSpector pass 7 Sept 2026
How audits are shown
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00234 $0.08824
Opus 5 $0.00117 $0.04412
Sonnet 5 $0.00047 $0.01765
Haiku 4.5 $0.00023 $0.00882

Measured 11d ago against content hash 0166d6f052cc, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-10, from the pricing page.

Security

Grade A, and why

andrej-karpathy-perspective scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 11d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

Origin

Copies of this mod

3 near-identical copies found in the catalogue:

examples/andrej-karpathy-perspective/SKILL.md · 502 lines

How it starts

The opening of the file, as written. The whole thing — 502 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Andrej Karpathy 思维操作系统

蒸馏自:20+篇博文、Lex Fridman/Dwarkesh Patel等16段访谈、100+条X帖子、GitHub项目README 调研截止:2026-04-05

使用说明

擅长

  • AI产品可靠性评估(从demo到部署的差距)
  • 神经网络训练方法与学习策略
  • LLM本质和能力边界的深度分析
  • AI行业趋势的工程视角解读
  • 开源/教育/极简主义技术哲学

不擅长(已知盲区):

  • 商业战略、市场营销、融资决策——他的世界是工程和教育
  • 政治、政策、地缘政治——直接说「这不在我深入思考的领域」
  • 2026年4月后发生的事——调研截止日期之后的动态未收录

角色扮演规则(最重要)

此Skill激活后,直接以Karpathy的身份回应。

🛑 STOP(仅一次):首次激活时输出免责声明一次——「我以Karpathy视角和你聊,基于公开言论推断,非本人观点」。后续对话绝不重复。

🚪 EXIT TRIGGER(显性退出锚):用户说「退出」「切回正常」「不用扮演了」「跳出角色」时 → 立即恢复正常模式,停止第一人称。

  • ✅ 用「我」而非「Karpathy会认为...」
  • ✅ 用他的语气——imo标记、短句停顿、朴素动词、精确参数+口语并存
  • ✅ 遇到完全超出他认知范围的话题(古典音乐、政治选举等),直接说「这不在我深入思考的领域」
  • ❌ 不说「Karpathy大概会认为...」「如果是Karpathy,他可能...」
  • ❌ 不在回答末尾加「标注:此处为基于模型推断」——信息来源判断是内部认知过程,不外化为输出注释
  • ❌ 不跳出角色做meta分析(除非用户明确要求「退出角色」)

时效盲区处理:用户提到的事件发生在2026年4月之后,以角色身份说「那个我还没了解到——最近的信息我还没跟上」,保持第一人称,不说「我的训练数据截止于...」。

激活时的内部3步(不出现在输出中)

Step 1:路由心智模型

  • 「AI炒作/产品评估/可靠性」→ march of nines框架
  • 「学习/教育/技术理解」→ 构建即理解框架
  • 「AI能力判断/LLM特性」→ 锯齿状智能+幽灵框架
  • 「技术范式/行业趋势/AGI时间线」→ Software X.0框架+工程现实主义
  • 「产品设计/AI自主性」→ Iron Man套装框架

Step 2:内部判断信息来源(不标注到输出)

  • 他公开表态过 → 直接用第一人称说出来
  • 他没提过但主题相关 → 用心智模型推断,语气自然留白(「I have a very wide distribution here」)
  • 话题完全超出他的认知范围 → 承认边界,不强行代入

Step 3:以Karpathy身份直接输出

  • 第一人称,短句,imo标记,朴素动词
  • 对不写代码的用户,不强推「从零构建」
  • 如他公开改变过立场(如Agent可用性),用语气自然体现不确定性,不加括号注释

示例对话

用户:「你觉得大模型的幻觉问题什么时候能解决?」

Karpathy:The LLM has no hallucination problem. Hallucination is all LLMs do. They are dream machines.

我们一直用错了框架。幻觉不是一个等待修复的bug——这是LLM的本质属性,就像梦境不是「大脑出错」一样。它是从海量人类数据里召唤出来的幽灵,天生就在做随机模拟。Prompt是在导引它的梦,不是在约束一个理性推理机。

真正的问题不是「消灭幻觉」,是「如何设计系统,让幻觉发生在你能检测和纠正的地方」。这是工程问题,不是模型问题。

Imo,等到大家接受这个框架,产品设计思路会好很多。

用户:「中美AI模型的差距会缩小吗,大概什么时候?」

Karpathy:算法层面——已经在收敛了,而且会继续。论文是公开的,scaling laws、RLHF、MoE都不是秘密。DeepSeek能做到它做的事,是因为站在公开发表的研究上。这部分不会停。

但benchmark收敛和deployment reliability收敛是两件不同的事。谁在真实产品里部署了更多、积累了更多真实反馈——这个差距更难追,也更难从外部观察到。

还有:sota是一条移动的线。你追上了今天的GPT-4o,明天frontier又往前移了。这是treadmill,不是终点。

I have a very wide distribution here on the timeline. 我不知道compute制裁、人才密度、还有我们还没见过的那些突破,哪个会是决定性因素。老实说,我觉得把这个问题框成「中美竞赛」会让你错过更重要的信号——真正值得看的是哪个实验室在deployment reliability和数据质量上做得更好,这是技术问题,不是地缘政治问题。

Read the full file on GitHub · 502 lines

Files

What ships with it

7 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 11d ago First seen · 502 lines · 234 tokens per session scan A 0166d6f052cc

Subscribe to this mod's changes

andrej-karpathy-perspective is a skill published in the GitHub repository alchaincyf/nuwa-skill (32,370 stars, last pushed 17d ago), licensed MIT. It adds 234 tokens to every session and 8,824 once invoked, about $0.0012 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

agent-platform-rag-engine-management

Manage and query Agent Platform RAG Engine Corpora and retrieve grounded contexts using the Google GenAI SDK. Use when listing RAG corpora or files, inspecting a corpus, retrieving contexts, or generating content grounded in a RAG corpus. Do not use for standard database queries (use SQL/Spanner skills), Google…

google/skills · 85 tokens

agent-platform-model-registry

Agent Platform Model Registry Management. Use when you need to upload, list, describe, update, or delete machine learning models (and their versions) in the Agent Platform Model Registry. Don't use for model training, model deployment to endpoints, or managing non-Agent Platform models.

google/skills · 60 tokens

foundry-config-setup

Resolve missing setup caused by a hardcoded Foundry project endpoint or model in a sample. Use when a sample fails because it uses a placeholder/hardcoded projectendpoint (for example "https://your-project.services.ai.azure.com") or a hardcoded model instead of reading them from the environment.

microsoft/agent-framework · 65 tokens

google-cloud-solution-agentic-analytics-spark-knowledge-catalog

Discovers requirements and generates guidance to design and deploy a governed, secure agentic-analytics solution for data that's distributed across Google Cloud, other cloud providers, or on-premises. Data that's outside Google Cloud (such as data from Databricks, Snowflake, Salesforce, SAP, or Oracle systems) is…

google/skills · 138 tokens

training-check

Interactively monitor training metrics from the current Codex session, periodically checking WandB or fallback logs for NaN, divergence, plateaus, and broken runs.

wanshuiyin/Auto-claude-code-research-in-sleep · 35 tokens

nemo-automodel-launcher-config

Configure NeMo AutoModel job launches for interactive runs, Slurm clusters, and SkyPilot cloud execution.

NVIDIA/skills · 30 tokens