modelscope/evalscope

A streamlined and customizable framework for efficient large model (LLM, VLM, AIGC) evaluation and performance benchmarking.

About the project

EvalScope is a framework for evaluating and stress-testing large language, vision-language, embedding, reranking, and generative models, with tools for running benchmarks and reviewing results. It is for developers and researchers comparing model capabilities, agent behavior, and inference performance. The catalogue add-ons provide workflows for running evaluations and interpreting their outputs.

Latest release v1.11.1 · 31 Aug 2026

These files are modelscope/evalscope's own configuration. They tell GitHub Copilot, Codex and OpenCode how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.

3,373Stars on the repository
2Files it configures its agents with
2,884Tokens loaded in every session
3Agents configured

Instructions