Fully Open Framework for Democratized Multimodal Training
LLaVA-OneVision-2 is an openly released multimodal AI model and training framework that processes images, long-form video, and spatial information. Researchers use it to train, evaluate, and reproduce vision-language models with the project’s released data, encoders, checkpoints, and training records. The catalogue skills support work with this model and its training resources.
Latest release 2.0 — Release v2.0 · 6 Aug 2026
These files are EvolvingLMMs-Lab/LLaVA-OneVision-2's own configuration. They tell OpenCode how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.
.opencode/skills/commit-message/SKILL.md A 17 tok .opencode/skills/cu-lengths-attention-flow/SKILL.md A 35 tok .opencode/skills/distributed-offline-packing/SKILL.md C 40 tok .opencode/skills/length-pool-sort-dataset/SKILL.md A 27 tok .opencode/skills/llava-onevision2-consistency/SKILL.md A 32 tok .opencode/skills/megatron-checkpoint-layout/SKILL.md A 33 tok .opencode/skills/merge-ov2/SKILL.md A 29 tok .opencode/skills/offline-packing-env-vars/SKILL.md A 65 tok