vision language skills

6 tagged vision language, measured the same way as everything else here.

Browse within: Multimodal 5

research-blip-2

01

GrayCodeAI/starling

Skill Claude CodeCodex

Vision-language pre-training framework bridging frozen image encoders and LLMs. Use when you need image captioning, visual question answering, image-text retrieval, or multimodal chat with state-of...

2 2d ago A 45 tokens original MIT