New Scalable Pipeline Enhances Video Indexing and Person Retrieval with Multimodal AI
Researchers have developed a scalable and semantic pipeline designed for efficient video indexing and person retrieval from large video collections. This new architecture integrates global visual embeddings, text embeddings, and face embeddings into a single vector database, enabling both semantic and identity-level retrieval. The system supports various query types, including image-based, text-based, hybrid, and face-based searches. It leverages advanced technologies such as YOLO11 for real-time person detection and tracking, SigLIP2 for multilingual vision-language embeddings, and InsightFace for face recognition. The pipeline also incorporates metadata-enriched vector indexing, ensuring full traceability from retrieval results back to the original video content. Experimental results on the RSTPReid benchmark indicate that the approach achieves state-of-the-art performance in text-based person retrieval, particularly after fine-tuning the SigLIP2 model on diverse person-centric datasets.