Skip to main content
1
PROPRIETARY
Advanced Video Search -
Leveraging Twelve Labs and
Milvus for Semantic Retrieval
Unstructured Data Meetup - South Bay Edition
August 13, 2024
2
PROPRIETARY
Video most closely resembles the
sensory inputs from the
real-world. We model the world
by shipping next-generation
multimodal foundation models
and pushing the boundaries of
video understanding.
Our Mission
3
PROPRIETARY
What is Video
Understanding?
4
PROPRIETARY
How Video Understanding Has Evolved Over The Years
What is Video Understanding?
Video invented
1878
First speech-to-text
commercialized
First CNN-based
image recognition (LeNet-5)
1996 1997 2022
Keep Binging till eternity!
1 Manual watching 2 Transcripts
Awful to read and
Totally disconnected from visual info
Doesn’t capture meaning or context and can
create huge discrepancies
Tags
3
5
PROPRIETARY
The Past, Present, and Future of Video Understanding
What is Video Understanding?
1. The Past: Solving Low-Level Video Perception Tasks (Object Detection, Object Tracking, Action Recognition,
Instance Segmentation)
2. The Present: Handling High-Level Video Understanding Tasks (Classification, Retrieval, Question Answering,
Captioning)
3. The Future: Going Multimodal with Video Foundation Models -> General-Purpose Video Understanding
6
PROPRIETARY
Video Foundation Models
What is Video Understanding?
Individual users & enterprises
Video-centric applications
Law Enforcement Contextual Ads E-learning Sports Creator Economy
End-users will demand video applications
to be intelligent from inception.
Applications built on top of embedding-based
gateway APIs (Search, Classify, Generate, and
Embed).
VIDEO FOUNDATION MODELS
Multimodal embeddings (Video-text)
SEARCH & CLASSIFY APIs GENERATE API EMBED API
APIs for downstream tasks
Video foundation models generate powerful
video-text embeddings.
Video embedding is a numerical
representation that stores all conversational
and visual semantic information from a video.
7
PROPRIETARY
Video Foundation
Models
8
PROPRIETARY
The Magic of Video Embeddings
Video Foundation Models
Source: The Multimodal Evolution of Vector Embeddings
9
PROPRIETARY
A SOTA Video Foundation Model for Any-to-Any Search
Video Foundation Models
Source: Introducing Marengo-2.6
10
PROPRIETARY
Twelve Labs Search API
Video Foundation Models
11
PROPRIETARY
Twelve Labs Classify API
Video Foundation Models
12
PROPRIETARY
Twelve Labs Embed API
Video Foundation Models
Video
POST Embed
input_type
file
: video, audio, image, text
: video.mp4
GET Embed
task_id
Embeddings
: 61e1127861c43d6d9b736194
[0.6,-0.2,0.3,0.4,...]
Video embeddings (semantic
representation)
GET
GET
Video-level Embeddings
Clip-level Embeddings
[0.6,-0.2,0.3,0.4,...],
[0.6,-0.2,0.3,0.4,...],
…
[0.6,-0.2,0.3,0.4,...]
13
PROPRIETARY
Twelve Labs Research Horizon
Video Foundation Models and Video Language Models
Source: Introducing Video-To-Text and Pegasus-1 (80B)
14
PROPRIETARY
Twelve Labs
Meets Milvus
15
PROPRIETARY
Advanced Multimodal Embeddings meets Efficient Vector Database
Twelve Labs and Milvus
Source: Advanced Video Search: Leveraging Twelve Labs and Milvus for Semantic Retrieval
16
PROPRIETARY
Twelve Labs and Milvus
Connecting to Milvus + Creating a Milvus Collection
17
PROPRIETARY
Twelve Labs and Milvus
Generating Embeddings + Insert Embeddings + Similarity Search
18
PROPRIETARY
Developer
Resources
19
PROPRIETARY
Twelve Labs Video Understanding Platform
Developer Resources
Source: Platform Overview
20
PROPRIETARY
Twelve Labs SDKs
Developer Resources
Source: Twelve Labs SDKs
21
PROPRIETARY
Jockey - A Conversational Video Agent
Developer Resources
Source: Introducing Jockey
22
PROPRIETARY
Developer Resources
Join Our
Discord For
Support
23
PROPRIETARY
James Le
james@twelvelabs.io