Skip to main content
What Real-Time AI
Requires from Your
Database
Dor Laor, CEO & Co-Founder, ScyllaDB
Avi Kivity, CTO & Co-Founder, ScyllaDB
Avi Kivity, CTO & Co-Founder, ScyllaDB
Dor Laor, CEO & Co-Founder, ScyllaDB
Introductions
2
ScyllaDB, Senior agent, ScyllaDB
Powering India's top
social media platform
Video recommendation
management
Real-time fraud
detection
Seamless experiences
across content + devices
Network security
threat detection
Content personalization &
recommendation platform
Mobile Growth &
Monetization Platform
Inventory hub for
retail operations
Property listings
and updates
Cryptocurrency
exchange app
Real-time auctions
advertising platform
Predictable performance
for on sales surges
Online gaming ad
targeting
Media streaming
for 45M+ subscribers
Bridging AI to IT Service
Management
Real-time ML-driven
recommendations
Real-time endpoint threat
detection and security
Real-time personalized
recommendations
World leading beauty
platform behind Avon
Real-time AI decisioning
for digital advertisers
AI-centric customer
research platform
Powering Unreal Engine
real-time asset distribution
Real-time interactions
at massive scale
Always-on e-commerce
platform for millions of fans
ScyllaDB Users
AI Break Databases
Prehistoric Past
Centralized DBs
& early NoSQL
The Iron Age
Close to metal DBs
Real Time AI Age
AI and ML boom
Agentic AI
(in Real Time)
Near Future
+ Huge scale
+ Enormous bursts
+ Low latency expectations
+ Complications: Hybrid DBs - Vector & Text search
AI Break Database Workloads
AI Break Database Workloads
Vector Search
Enables semantic
information retrieval by
indexing high-dimensional
embeddings. Critical for
RAG architectures,
Feature Store
A centralized repository for
managing and serving ML
features. Ensures
consistency between
training and serving while
promoting feature
reusability across teams.
Agentic AI
Autonomous systems
capable of reasoning,
planning, and tool
utilization. Orchestrates
multi-step tasks by shifting
from passive chat to active
goal execution.
Large Scale
High-throughput workloads
requiring massive
distributed clusters.
Optimized for model
pre-training, fine-tuning,
and ultra-low latency
inference at scale.
AI Break Database Workloads
Vector Similarity
Feature Store Highly Scalable Database
within an AI Stack
Dissimilar
Dissimilar
Similar Results
(Matches)
Vector Space
Query
Vector
Fast, Accurate
Matching of Complex Data
Raw Data
Sources
Real-time
Event Streams
ML Training
ML Serving
(Interference)
Central
Feature Store
(ScyllaDB)
Ingestion
Processing
& Training
Serving &
Inference
High Throughput & Low Latency Access
Agentic AI
The Agentic AI Age
+ Agents are the primary database user — Not Humans
+ Human - Chat requests per second; Agents - per milliseconds
+ Human - few users per org; Agents - Unlimited number
+ Human - predictable usage (mostly); Agents - Unpredictable
Brave New Agentic World
LangGraph: Achieving zero agent downtime with ScyllaDB.
More Agentic Resources:
AI Use Case: Feature Store
Driving Tripadvisor ML personalization
How ShareChat built a scalable cost efficient
ML Feature system
https://sharechat.com/blogs/artificial-intelligence/how-sharechat-built-a-scalable-cost-efficient-ml-feature-system
14
Latency, Op/s, Scale Results
The Real Time AI Age
Self-driving Model Training
Use case: Serves a huge GPU farm as a database
for model training. Migration off a well known DB
due to cost
Scale:
+ 180TB per node, compressed
+ 0.5PB per node uncompressed
+ Tens of nodes
+ > 1M op/s
+ 256 cores per node
Scylla Design Decisions
Threads
Shards
1 C++ instead of Java
2 All Things Async
3 Shard per Core
4 Unified Cache
5 I/O Scheduler
6 Autonomous
Scylla Design Decisions
Cassandra Scylla
Key cache
Row cache
On-heap /
Off-heap
Linux page cache
SSTables
Unified cache
SSTables
1 C++ instead of Java
2 All Things Async
3 Shard per Core
4 Unified Cache
5 I/O Scheduler
6 Autonomous
App
thread
Kernel
SSD
Page fault
Suspend thread
Initiate I/O
Context switch
I/O
completes
Interrupt
Context
switch
Map page
Resume
thread
Page fault
Scylla Design Decisions
1 C++ instead of Java
2 All Things Async
3 Shard per Core
4 Unified Cache
5 I/O Scheduler
6 Autonomous
Memtable
Seastar
Scheduler
Compaction
Query
Repair
Commitlog
SSD
Compaction
Backlog Monitor
Memory Monitor
Adjust priority
Adjust priority
WAN
CPU
AI & DB Peaks
AWS DynamoDB Auto Scaling is Not a Magic Bullet
+ Faster Topology Changes
+ Immediate Request Serving
+ Easy Downscale
+ Continuous Auto-Balancing
Tablets – True Elastic Scale
X Cloud Elasticity -> Serverless
Policy Updated
New Nodes Joining
Mixed Instance Size
Latency: Before, During, 1M Scale
+ At rest and in transit Dictionary compression
+ Latest ARM instances
+ Trie Indexes
+ Incremental repair
+ Authentication and authorization Cache
+ Thin Replicas
+ Tiered Storage
TCO - Every ScyllaDB release is more efficient
Every Agent is different
AI wouldn’t break ScyllaDB
Time
Volume
ie3n’s
i8g
Time
Throug
hput
Capacity
Required
Time
Throug
hput
On-demand
Base
Typeless Sizeless limitless
AI Use Case: Vector Search
Vector Similarity & Text Search Architecture
App
App
App
AZ1
AZ2
AZ3
Scylla
DB
Scylla
DB
Scylla
DB
Scylla
DB
Scylla
DB
Scylla
DB
Scylla
DB
Scylla
DB
Scylla
DB
VS
VS
CQL
Novel CDC at Scale
ScyllaDB->CDC->Vector node
INSERT INTO base_table(...)...
CQL
(Opt) preimage read
Vector Similarity Search Resources
Native Vector Search
for the DynamoDB API
Movie RAG
search with
CQL
Local or
Global index
filters
Is your Database ready?
Real-Time Agentic AI is here
Thank you
for joining us today.
@scylladb scylladb/
slack.scylladb.com
@scylladb company/scylladb/
scylladb/