TwelveLabs Unveils Marengo 3.5, Advancing Video-Native Multimodal AI for Enhanced Retrieval
TwelveLabs has launched Marengo 3.5, its latest video-native multimodal embedding model, designed to improve video classification, question answering, and moment retrieval. Building on Marengo 3.0, this new version focuses on representing diverse signals, aligning context with time, and structuring video around meaningful events. Marengo 3.5 processes video, audio, images, text, and documents within a shared semantic space, enabling cross-modal and composed input queries. The model demonstrates superior performance across various benchmarks, including MMEB-v2 video classification and question answering, and leads in visual documents. It also introduces flexible dimensions, uncertainty estimation, and enhanced temporal segmentation, accurately identifying transition boundaries in video content. The model's ability to fuse timestamped metadata further enriches retrieval capabilities for complex queries.