EvolvingLMMs-Lab/LLaVA-OneVision-2

Fully Open Framework for Democratized Multimodal Training

About the project

LLaVA-OneVision-2 is an openly released multimodal AI model and training framework that processes images, long-form video, and spatial information. Researchers use it to train, evaluate, and reproduce vision-language models with the project’s released data, encoders, checkpoints, and training records. The catalogue skills support work with this model and its training resources.

Latest release 2.0 — Release v2.0 · 6 Aug 2026

These files are EvolvingLMMs-Lab/LLaVA-OneVision-2's own configuration. They tell OpenCode how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.

1,197Stars on the repository
8Files it configures its agents with
Tokens loaded in every session
1Agent configured

Skills