Evals for Next.js up to 15.5.6 to test AI model competency at Next.js
vercel/next-evals-oss is a set of evaluations that tests how well AI coding agents complete tasks in Next.js applications. Each evaluation gives an agent a small app in an isolated sandbox and checks the resulting code with assertions that are hidden from the agent. The catalogue entries provide instructions and a skill for running or developing these evaluations.
3 files for Claude Code, Codex and OpenCode: add-eval-model, next-evals-oss AGENTS.md, next-evals-oss CLAUDE.md — 416 tokens loaded in every session.
AGENTS.md A 413 tok CLAUDE.md A 3 tok .agents/skills/add-eval-model/SKILL.md A 91 tok These files are vercel/next-evals-oss's own configuration — they tell Claude Code, Codex and OpenCode how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.