Skip to main content
Digital Transformation Through Data Analytics
AI Powered Transformation
Dell / NVIDIA AI Roadshow
Bill Wong – Dell Technologies Artificial Intelligence and Data Analytics Practice Leader
Adam Shubinsky – NVIDIA Data Science & High Performance Compute Solutions Leader
Ziad Najjar – Dell Technologies Infrastructure Solutions Group Director
Agenda
 Key Business Challenges and Trends
 NVIDIA Update
 AI Transformation Challenges
 NVIDIA Infrastructure Solutions
 Dell Technologies AI Strategy
 Industry Trends
 AI Partner Ecosystem
 Summary
 Dell NVIDIA Partnership
COVID-19 and the New Reality
Supply Chain Disruptions
 COVID-19 is forcing companies to question
the ability of their supply chains to effective
and reliably respond to rapidly evolving risk
factors.
Labour
 The unemployment rate rose 5.2
percentage points in April to 13.0%.
 2.1 million people worked reduced hours.
Revenue Challenges
 Across the country, over half of businesses in
Alberta (57.7%), Ontario (56.3%), B.C.
(54.8%), Newfoundland and Labrador (53.5%),
and Saskatchewan (52.8%) saw declines of
20% or more in revenue.
(Conference Board of Canada - March, 2020)
Transportation Disruptions
 Over 80 countries have issued travel
restrictions related to COVID-19.
Today’s Business Trends
Trend #1: Remote working will be a prevalent way of working
Trend #2: Organizations across industries will increasingly rely on digital
platforms/channels to increase future resilience and growth
Trend #3: Data and analytics become essential in assisting faster and better
decision-making
According to IDC, By 2025, at least 90% of new enterprise apps will embed
artificial intelligence. Most of these will be AI-enabled apps, delivering
incremental improvements to make applications "smarter" and more dynamic.
Support for a variety of
approaches to AI
(product or operational)
• Extensive Portfolio of Offerings
Workstations, servers, networking, storage, software and
services — to create end-to-end solutions that underpin
successful AI, machine and deep learning
implementations
• AI Expertise
Dell EMC experts to help users adapt as AI, machine and
deep learning evolve over time – and work with customers
on systems designs and POCs (HPC and AI Innovations
Lab, Solution Centers)
• AI Partner Solutions
Simplified deployment on an optimized IT infrastructure
Dell Technologies Value Proposition for AI
DGX A100: THE UNIVERSAL AI SYSTEM
Performance meets utility
– analytics, AI training and
inference all in one
One System for
Every AI
Workload
Fast-track AI
transformation with
DGXpert know-how
and experience
Integrated Access to
Unmatched AI
Expertise
Fastest time-to-solution with the
world’s first 5 petaFLOPS AI
system, built on NVIDIA A100
Game-changing
Performance for
Innovators
Build leadership-class
infrastructure that scales
to keep ahead of demand
Unmatched Data
Center Scalability
ONE SYSTEM FOR ALL AI INFRASTRUCTURE
any job | any size | any node | anytime
Analytics  Training  Inference
Flexible AI infrastructure that adapts to the
pace of enterprise
One universal building block for the AI data
center
Uniform, consistent performance across the
data center
Any workload on any node - any time
Limitless capacity planning with predictably
great performance with scale
AI Infrastructure Re-Imagined, Optimized, and Ready for Enterprise AI-at-Scale
TODAY’S AI
DATA CENTER
50 DGX-1 systems for AI
training
600 CPU systems for AI
inference
$11M
25 racks
630 kW
5 DGX A100 systems for
AI training and
inference
$1M
1 rack
28 kW
1/10th
COST
1/20th
POWER
$1M 28 kW
DGX A100
DATA CENTER
Game-changing performance for innovators
9x Mellanox ConnectX-6 200Gb/s Network Interface
8x NVIDIA A100 GPUs with 320GB Total GPU Memory
15TB Gen4 NVME SSD
Dual 64-core AMD Rome CPUs and 1TB RAM
4.8TB/sec Bi-directional Bandwidth
2X More than Previous Generation NVSwitch
6x NVIDIA NVSwitches
12 NVLinks/GPU
600GB/sec GPU-to-GPU Bi-directional
Bandwidth
25GB/sec Peak Bandwidth
2X Faster than Gen3 NVME SSDs
3.2X More Cores to Power the Most Intensive AI Jobs
450GB/sec Peak Bi-directional Bandwidth
Nvidia DGX A100 SYSTEM SPECS
App Focus Components
GPUs 8x NVIDIA A100 Tensor Core GPUs
GPU Memory 320GB Total
NVIDIA NVSwitch 6
Performance
5 petaFLOPS AI
10 petaOPS, INT8
CPU
Dual AMD Rome, 128 cores total, 2.25 GHz
(base), 3.4 GHz (max boost)
System Memory 1TB
Networking
9x Mellanox ConnectX-6 VPI HDR
InfiniBand/200GigE
10th Dual-port ConnectX-6 optional
Storage
OS: 2x 1.92TB M.2 NVME drives
Internal Storage: 15TB (4x 3.84TB) U.2
NVME drives
Power and Physical Dimensions
System Power Usage 6.5 kW Max
System Weight 271 lbs (123 kgs)
System Dimensions
6 Rack Units (RU)
Height: 10.4 in (264.0 mm)
Width: 19.0 in (482.3 mm) Max
Length: 35.3 in (897.1 mm) Max
Operating Temperature 5ºC to 30ºC (41ºF to 86ºF)
Cooling Air
NVIDIA A100
Greatest Generational Leap – 20X Volta
54B XTOR | 826mm2 | TSMC 7N | 40GB Samsung HBM2 | 600 GB/s NVLink
Peak Vs Volta
FP32 TRAINING 312 TFLOPS 20X
INT8 INFERENCE 1,248 TOPS 20X
FP64 HPC 19.5 TFLOPS 2.5X
MULTI INSTANCE GPU 7X GPUs
DGX A100: NEW A100 GPUs AND 2X FASTER NVSWITCH
5 PetaFLOPS AI Performance
Eight new A100 Tensor Core GPUs/320GB total HBM2
Twelve NVLinks per GPU, 2x more than V100
600GB/s bi-directional bandwidth between any GPU pair
~10X PCIe Gen4 bandwidth with next-gen NVLink
All GPUs fully connected with six next-gen NVSwitch
4.8TB/s bi-directional bandwidth
In one second we could transfer 426 hours of HD video
NEW T32 TENSOR CORES ON A100
20X Higher FLOPS for AI, Zero Code Change
20X Faster than Volta FP32 | Works like FP32 for AI with Range of FP32 and Precision of FP16
No Code Change Required for End Users | Supported on PyTorch, TensorFlow and MXNet Frameworks Containers
MOST FLEXIBLE AI PLATFORM WITH MULTI-INSTANCE GPU (MIG)
Optimize GPU Utilization, Expand Access to More Users with Guaranteed Quality of Service
Up To 7 GPU Instances In a Single A100:
Simultaneous Workload Execution With
Guaranteed Quality Of Service:
All MIG instances run in parallel with predictable
throughput & latency
Flexibility to run any type of workload on a MIG
instance
Right Sized GPU Allocation:
Different sized MIG instances based on target
workloads
Amber
GPU Mem
GPU
GPU Mem
GPU
GPU Mem
GPU
GPU Mem
GPU
GPU Mem
GPU
GPU Mem
GPU
GPU Mem
GPU
MULTI-INSTANCE GPU (MIG) ON DGX A100
More Users and Better GPU Utilization
Flexible Utilization
Configure GPUs for vastly different workloads
with GPU instances that are fault-isolated
GPU 32 4 5 6 7 8
GPU
Instance Size
Number of
GPU Instances
Available
GPU Memory
1 GPU Slice 7 5 GB
2 GPU Slice 3 10 GB
3 GPU Slice 2 20 GB
4 GPU Slice 1 20 GB
7 GPU Slice 1 40 GB
1
1 DGX A100
=
56 users
Inference with TensorRTBatch training with NGC containerJupyter Notebook
CONSOLIDATING DIFFERENT WORKLOADS ON DGX A100
One Platform for Training, Inference and Data Analytics
TRT TRT TRT TRT TRT TRT TRT
TRT TRT TRT TRT TRTT TRT TRT
Instance 1 Instance 7
Instance 14Instance 8
2x A100s for inference in MIG mode
Data Analytics
Training
4x A100s
2x A100s
Highest Network Throughput for Data and Clustering
UNMATCHED SCALABILITY WITH MELLANOX NETWORKING
Cluster
Networking
Storage Networking
Single-port
CX-6 NIC
Cluster
Networking
For clustering networking:
Eight Mellanox single-port ConnectX-6
Supporting HDR/HDR100/EDR InfiniBand default or 200GigE
For data/storage networking:
One Mellanox dual-port ConnectX-6
Supporting: 200/100/50/40/25/10Gb Ethernet default or
HDR/HDR100/EDR InfiniBand
One optional Dual-Port CX-6 available as add-on
450GB/sec peak bi-directional bandwidth
All I/O now PCIe Gen4, 2x performance increase over Gen3
Scale up multiple DGX A100 nodes with Mellanox Quantum Switch,
the world’s smartest network switch
DGX A100 PERFORMANCE
8x V100
FP32
DGX A100
TF32
216
Sequences/s
1289
Sequences/s
CPU Server DGX A100
58 TOPS
10 PetaOPS
172
X
CPU Cluster* DGX A100*
52B Graph
Edges/s
13X
Inference
Peak Compute
Analytics
PageRank
Training
NLP: BERT-Large
688B Graph
Edges/s
3000x CPU Servers vs. 4x DGX A100
Published Common Crawl Data Set:
128B Edges, 2.6TB Graph
6X
BERT Pre-Training Throughput using PyTorch including
(2/3)Phase 1 and (1/3)Phase 2 | Phase 1 Seq Len = 128, Phase 2
Seq Len = 512 V100: DGX-1 Server with 8x V100 using FP32
precision
DGX A100: DGX A100 with 8x A100 using TF32 precision
CPU Server: 2x Intel Platinum 8280 using
INT8
DGX A100: DGX A100 with 8x A100 using
INT8 with Structural Sparsity
THE WORLD’s MOST SECURE AI SYSTEM FOR ENTERPRISE
Built-In Security: Multi-layered Defense for AI Infrastructure
Secure boot
Self-Encrypted Drives (SED)
to protect data at rest
Secure
Firmware
Update
CPU
Board
BMC
DGX A100 delivers the most robust security
posture for your AI enterprise
GPU
Board
ELASTIC AI INFRASTRUCTURE WITH DGX A100
DGX A100 with MIG Delivers New Agility for Today’s Enterprise Data Center
DGX A100 Infrastructure is Agile
DGX A100 infrastructure uses MIG to allocate GPU resources to workloads
TRAINING CLUSTER ANALYTICS CLUSTER
INFERENCE
CLUSTER
OVEROPTIMAL UNDER
Infrastructure silos starve AI workloads or waste capacity
ANALYTICS
INFERENCE
TRAINING
Toda
y
Tomorro
w
Next
Week
Traditional Infrastructure is Constrained
DGX A100 LOWERS TCO WITH MAXIMIZED UTILIZATION
Adapt to Changing Business Needs Without Reinvesting
Legacy infrastructure is inflexible
Sits idle when demand drops, unable to scale
when demand increases
Nearly impossible to optimize utilization
BEFORE AFTER
DGX A100 is agile, outperforming legacy for every AI workload:
analytics, training, and inference
Adapts to business demand providing a single elastic
infrastructure that’s more efficient
Better utilization = lower TCO and faster ROI on AI
0
10
20
30
40
50
60
70
80
90
100
Training Cluster
Time
0
10
20
30
40
50
60
70
80
90
100
Combined Workloads on…
Time
Target utilizationTarget utilization
MOST POWERFUL TOOL FOR A DATA SCIENCE TEAM
Using DGX A100 with MIG to Give Every Developer Power to Explore
One DGX A100 delivers:
5 petaFLOPS of AI training power, or
10 petaOPS of AI inference power
With MIG, a team of 25 developers can share a DGX A100
Each developer gets:
Over 180 teraFLOPS for training
= (2) reserved cloud V100 instances
or
Over 357 teraOPS for inference
= (6) dedicated 28-core dual CPU servers
Dell Technologies
24
AI/ML/DL is the fastest growing Datacenter workload
Worldwide AI Spending
~$98 Billion by 2023
Overall CAGR = 28.5%
• H/W CAGR=24.1%
• S/W CAGR=36.7%
• Services CAGR=25.9%
Identify the Game Changer Technologies for Your Organization
Data Lake for Advanced Analytics
Supporting the Digital Transformation at Dell Technologies
Consumption
Zone /
Data Analytics
Raw /
Landing/
Secure Zone/
Data Ingestion
CRM DataMachine
Logs/IoT
Self-Service Dashboards
Advanced Analytics
Sales
Analysts
Consumer Dashboards
Operational Analytics
Data
Scientists
Customers
Marketing
Analysts
Data Governance | Security and Compliance
Enriched /
Discovery Zone /
Data
Transformation
Data Sources
Common Services
Optimized Infrastructure for Advanced Analytics
Social Media
Chatbots
Personas
Tools /
Applications
Data Lake Capabilities
• Provide support for a variety of analytical applications, including self-service, operational, and data science analytics
• Data preparation and integration capabilities to ingest structured and unstructured data, move and transform raw data to
enriched data, and enable data access to for the target user base
• An infrastructure platform optimized for advanced analytics that can perform and scale
ERP Data
AI Accelerators
Flexibility Efficiency
and many more…
GPU Virtualization Economics
Thermal Vision Solutions - Industry Aligned Use Cases
• Health & Safety
– Thermal Evaluation
– Social Distancing
– Facial Recognition
– Mask Verification
• Energy & Utilities
– Transmission Line Monitoring (insulators cracked)
– Pipe corner analysis: erosion detection
– Oil storage tanks level in tank farms
• Manufacturing
– Bearing temperature monitoring on conveyer
systems
– Leak detection
– Stress fractures/hot spots in commercial furnaces,
crude heaters
• Smart Buildings
– Chiller analysis
– Insulation effectiveness/leakage
• …
•Detection of persons/objects
•Display showing temperature differences accurate to 0.1°C
•Alarm in case of exceeding or falling below defined temperature ranges
•Event Triggers (alarm, network message, activation of a switching output)
•Temperature range from -40 to +550 °C
•Face Redaction for privacy
Mobile Version
Dell Workstation
with NVIDIA
Extensive, Validated Partner Surveillance Solutions
• #1 Virtualization
Platform for
Surveillance
• Extend Surveillance
to the Cloud
• Test to Fail Philosophy
• Proven Solutions
• Reduced Complexity
• Unlock Surveillance
Data Value
• Simplify Evidence
Management
Virtualization Video
Surveillance
Mgmt Software
ApplicationsSecurity
• Authorized Access
Protection
• Audit Trails
• Anomaly Detection
• Enterprise Data
Analytics
Analytics
AI Magic Quadrants
Cloud AI developer services are defined as:
• Cloud-hosted services/models that allow development teams to
leverage AI models via APIs without requiring deep data
science expertise
Data Science and Machine Learning Platforms Cloud AI Developer Services
The Marketplace
Continues To
Evolve
Data science and machine-learning platforms are defined as:
• A cohesive software application that offers a mixture of basic building blocks
essential both for creating many kinds of data science solution and incorporating
such solutions into business processes, surrounding infrastructure and products.
Data Analytics and AI Use Cases – Partner Solutions
IOT / Streaming /
Machine Data Analytics
Deliver Near Real-Time
Analytics
• Analyze IOT / Streaming
data
• Improve IT operations and
security leveraging Machine
Data
• Computer vision
applications
Machine / Deep Learning
Transform the business
with analytical insights
• Data Science / Machine
Learning Platform
• Industry-focused AI
platforms
Data Lake/Unstructured
Data Infrastructure
Improving Data Access
and Agility
• Create an enterprise data
platform for structured and
unstructured data
• ETL offload to lower costs
• On-demand deployment of
container-based
environments
Augmented Analytics
and Data Warehouse
Improve Decision
Making
• Support augmented
business analytics
• Create an enterprise data
platform to support
analytics
• Data integration and
Master Data Management
*Note, some products can deliver capabilities that address multiple use cases
H2O.ai DataRobot
AutoML offerings H2O Driverless AI (commercial) and H2O-3 (open source)
• Good adoption of its open source offering
• Machine Learning Interpretability generates the constructs for the data
scientist to use and explain the results of the models
AutoML offerings enables business users and the Citizen Data Scientist
• Easy to use, you do not need to be a data scientist
• Prediction Explanation: Highlights the features that impact each
model’s decision
Driven to be Disruptive – OTTO Motors
Opportunity
Material handing automation (which can account for up to 70% of
the final cost of good) offers the potential for significant
performance improvements and allows facility operators to focus
more resources on core, value-added operations.
Solution
OTTO, the self-driving vehicle, is designed exclusively for material
transport in industrial environments such as manufacturing
facilities and warehouses.
Benefits
OTTO Motors is democratizing robotics technology for their
customers, and making it possible for companies of all sizes to
adopt self-driving vehicles in their work environments.
“We have, over the last few years, evolved
significantly into an autonomous systems and
software company, and with that, really taken the
promise of data collection and data analysis to
heart. I’d say that it’s something we’ve rapidly seen
the benefit from, and are rapidly encouraging the
adoption of.”
- OTTO CTO -
Deep Learning Analytics – GPU, Graphcore
Dell Technologies – AI Compute Platforms
Performance
Inference
Data Analytics
Multi-App HPC / ML / DL
C6420pC6420p
R840
DS8440
8+
4
2 - 3
1
Solution price $
C4140C4140
GPU DB Acceleration, AI/ML R940xa
SDS/VDI R740XDR740XD
1:1 CPU/GPU ratio
Highest density of
CPU and memory
with 2 GPUs
GRAPHCORE IPU
XILINUX FPGA INTEL FPGA
NVIDIA GPU
INTEL CPU AMD CPU
GRAPHCORE IPU
XILINUX FPGA INTEL FPGA
NVIDIA GPU
INTEL CPU AMD CPU
Dell EMC Data Science
Platform
Nauta ClaraAI KubeFlow
NVIDIA
EGX
Domino Cassandra HPCaaS Metropolis Spark Jupyter
Bright
Cluster
Manger
Dell-curated
Ansible/
Terraform
playbooks
CNI MetalLB CoreDNS Prometheus NFS provisioner
Helm
Kubernetes
Linux (RHEL/CentOS) + CRI (Docker/containers)
1 https://infohub.delltechnologies.com/section-assets/h18136-tco-analysis-dell-emc-hpc-ra-for-ai-da-sb
On-premises system for HPC, AI and Data Analytics
AI / Machine learning / Deep
learning
PowerSwitch S3148-ON
S5232F-ON cluster switch
PowerEdge R740
management and
compute nodes
PowerEdge C4140
acceleration nodes
DSS 8440 dense
acceleration nodes
Dell EMC Isilon
Dell EMC Ready Solution for HPC BeeGFS Storage
Dell EMC Ready Solution for HPC NFS Storage
© Copyright 2020 Dell Inc.
HPC AI Ready Architecture
One Platform for AI, Data Analytics, and Simulation Workloads
• Simplified operations and lower
cost while enabling new use
cases for users at the lowest
TCO1
• Allow HPC, DA & AI workloads
to execute on the same cluster;
reducing data movements for
faster results
• Run simulation & modeling,
analytics, visualization, and AI
workloads on a common HPC
infrastructure
Software ecosystem
• Eliminate inefficient islands of storage
– Infrastructure consolidation for both clinical and non-clinical workloads
• Scales as data growth and number of instruments,
modalities, and digital clinical applications
increases
• Enable better information sharing
• Accelerate data analytics to gain new insight
• Extends into the cloud
• Prepared for next generation analytics
Dell EMC
Data Lake
Caffe2
Data Lake Storage Platform
The Digital Future Demands a New Perspective
Cloud First Data First
Infrastructure-centric Business-centric
Takes into consideration:
• Data gravity
• Data velocity
• Data control
• Data privacy and compliance
Driven by:
• Lower infrastructure CapEx
• Offload infrastructure maintenance
• Improve time to market (deployment
time for infrastructure)
Evolve to a Data-Driven Business
Decision Criteria for AI Infrastructure/Solutions
Data Scientist Perspective
IDC 2018
IoT Surveillance LabsAI Innovation Lab
• Global locations
• Dell IoT product development
• Test/validate ISV w/ Dell
hardware
• Round Rock, TX
• 13,000 sq. ft facility
• Supercomputers:
• NVIDIA, Intel, AMD
• No-charge service for customers
Dell Technologies Resources for AI
Industry’s largest, most advanced
surveillance validation labs
The Value of Dell for AI Infrastructure
- Comprehensive and Scalable AI/Analytics Platform Portfolio
- Workstations, Servers, Clusters, Storage, Networking
- Infrastructure and Data Science and Analytics Expertise
- HPC and AI Innovation Lab
- IoT / Intelligent Video Analytics Lab
- Solution-based Offerings
- Pre-configured AI Ready Offerings
- IoT / Safety and Security and
Thermal Vision Solutions
- GPU Virtualization
- ML Platforms
Infrastructure
Scalability
Reduce
Complexity
Address
Demand
Partner
Ecosystem
Cost
Effective
- Appendix -
Dell Technologies
AI and Data Analytics Solutions
Dell Technologies AI and Data Analytics Solutions
AI / Machine Learning / Deep Learning
• Domino Data Science Platform Design Document
• HPC for AI and Data Analytics Ready Architecture
• Retail Loss Prevention Ready Solutions
• DataRobot Reference Architecture
• H2O AI Reference Architecture
• Kubeflow Reference Architecture
• OneConvergence Dkube Reference Architecture
• Iguazio Reference Architecture
• Deep Learning with NVIDIA Ready Solutions
• Isilon with NVIDIA DGX-1 Reference Architecture
• Isilon with NVIDIA DGX-2 Reference Architecture
• Isilon with Dell Precision 7920 Data Science Workstation Reference Architecture
• Isilon with Dell EMC DSS8440 Reference Architecture
• Noodle.ai (OEM) Solution Bundle
IoT / Streaming / Machine Data Analytics
• IntelliSite (OEM) Thermal Detection Solution
• Retail Loss Prevention Ready Solutions
• Dell IoT Safety and Security Portfolio
• Real-Time Data Streaming Ready Architecture
• Splunk Enterprise on Dell EMC Infrastructure
• Streaming Data Platform
• ElasticSearch (OEM) Solution Bundle
© Copyright 2020 Dell Inc.
Augmented Analytics and Data Warehouse
• Spark on Kubernetes
• Kinetica (OEM) Solution Bundle
• ThoughtSpot (OEM) Solution Bundle
• Pivotal Greenplum
• Dell Boomi
Data Lake / Unstructured Data Infrastructure
• Microsoft SQL Server 2019: Big Data Cluster Ready Solution
• Cloudera Hadoop Ready Architecture
• Hortonworks Hadoop Ready Architecture
• Kubernetes Containers with Diamanti (OEM) Solution Bundle
• Grid Dynamics Reference Architecture
• Red Hat OpenShift Reference Architecture
HPC Ready Solutions
• HPC Digital Manufacturing
• HPC Life Sciences
• HPC Research
• HPC BeeGFS Storage
• HPC Lustre Storage
• HPC NFS Storage
• HPC PixStor Storage
*Note, some products can deliver capabilities that address multiple use cases
Product Offerings and Technical Collateral for Analytical Use Cases