Scanning the Internet for External Cloud Exposures via SSL Certs
Hadoop Big data Solution Provider
1. Experts in
&
Enterprise Data Lake
Build Lake on Cloud
A T T U N I T Y
PARTNERS
Innovating and Engineering
High Performance Data Integration
BI & Analytics
Platforms
On-Premise, Cloud or Hybrid.
2. Data Lake is a repository for large quantities and varieties of data,
both structured and unstructured. Data generalists / programmers
can tap the stream data for
real-time analytics.
Data scientists use the
lake for discovery and
ideation.
Data lakes take advantage of commodity cluster
computing techniques for massively scalable,
low-cost storage of data files in any format.
Of�load “Cold” Data From DW to Hadoop
Dramatically lowers the cost per
terabyte to store data - Hadoop
based storage is 30x cheaper
More Information can be
retained and analyzed
Improves performance of the
Data Warehouse
“Cold” data still available to be
queried on-line or interactively
“Cold” data in Hadoop can be
mined for additional insights or
combined with other data
Bene�its
Data
WarehouseETL
Reports / Dashboard /
Queries
“HOT”
Hadoop “COLD”
Ongoing
data load
Initial bulk load of raw or
infrequently used data
Re-factor queries
and reports to
work via HIVE-QL
Translate DW Data
Model to Hive /
HCatalog
For
frequently
used data
AFTERBEFORE
The data lake accepts input from various sources and
can preserve both the original data fidelity and the
lineage of data transformations. Data models emerge
with usage over time rather than being imposed up front.
The lake can serve as a staging area for
the data warehouse, the location of more
carefully "treated" data for reporting and
analysis in batch mode.
What is a Data Lake?
3. Qubole
AWS Data
Pipe Line
FTP
EnterpriseSystems
DATA LAKE
ON CLOUD
AWS - S3
Amazon AWS Cloud
Facebook
Twitter
Google +
iTunes Store
Google Play
You Tube
Amazon MP3
Spotify
VEVO
Amazon Prime
HULU
DATA ARCHIVES
XML
OTHER
EXCEL
TXT
CSV
JSON
EDI
External Business
Partners & Third Party
SAP
MySQL
Product,Customer
&OtherData
CRMOracle
Oracle SQL
Server
MySQL Oracle SQL
Server
MicroStrategy | Business Objects
Dashboard
ETL
Reporting
FTP
Spark
HIVE
Presto
Hadoop
Qubole
Analytics & Data
Scientist
MicroStrategy | TableauHadoop Map
Reduce
Data
Stream’s
to Data
Lake On-Demand Data Flow
Regular Data Flow
Replication
Data Lake
Reference Architecture