Loading lesson page...
LLaVA-OneVision: Single-Image, Multi-Image, Video in One Model
BuildPython (stdlibtoken budget solver + curriculum planner)No prerequisitesBefore LLaVA-OneVision (Li et al., August 2024) the open-VLM world had separate lineages: LLaVA-1.5 for single images, multi-image models like Mantis and VILA, video models like Video-LLaVA and Video-LLaMA. Each won its benchmark and faile...