A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.
EvalScope is a framework for evaluating and stress-testing large language, vision-language, embedding, reranking, and generative models, with tools for running benchmarks and reviewing results. It is for developers and researchers comparing model capabilities, agent behavior, and inference performance. The catalogue add-ons provide workflows for running evaluations and interpreting their outputs.
Latest release v1.11.1 · 31 Aug 2026
These files are modelscope/evalscope's own configuration. They tell GitHub Copilot, Codex and OpenCode how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.
.github/copilot-instructions.md A 4 tok AGENTS.md A 2,880 tok