SlideShare a Scribd company logo
www.studymafia.org
Submitted To:
www.studymafia.org
Submitted By:
www.studymafia.org
Seminar
On
BIG DATA
INTRODUCTION
 Big Data may well be the Next Big Thing in the IT world.
 Big data burst upon the scene in the first decade of the
21st century.
 The first organizations to embrace it were online and
startup firms. Firms like Google, eBay, LinkedIn, and
Facebook were built around big data from the beginning.
 Like many new information technologies, big data can
bring about dramatic cost reductions, substantial
improvements in the time required to perform a computing
task, or new product and service offerings.
WHAT IS BIG DATA?
 ‘Big Data’ is similar to ‘small data’, but bigger in size
 but having data bigger it requires different
approaches:
Techniques, tools and architecture
 an aim to solve new problems or old problems in a
better way
 Big Data generates value from the storage and
processing of very large quantities of digital
information that cannot be analyzed with traditional
computing techniques.
WHAT IS BIG DATA
 Walmart handles more than 1 million customer
transactions every hour.
• Facebook handles 40 billion photos from its user base.
• Decoding the human genome originally took 10years to
process; now it can be achieved in one week.
THREE CHARACTERISTICS OF BIG DATA
V3S
Volume
• Data
quantity
Velocity
• Data
Speed
Variety
• Data
Types
1ST CHARACTER OF BIG DATA
VOLUME
•A typical PC might have had 10 gigabytes of storage in 2000.
•Today, Facebook ingests 500 terabytes of new data every
day.
•Boeing 737 will generate 240 terabytes of flight data during a
single flight across the US.
2ND CHARACTER OF BIG DATA
VELOCITY
 Clickstreams and ad impressions capture user behavior
at millions of events per second
 high-frequency stock trading algorithms reflect market
changes within microseconds
 machine to machine processes exchange data between
billions of devices
 infrastructure and sensors generate massive log data in
real-time
 on-line gaming systems support millions of concurrent
users, each producing multiple inputs per second.
3RD CHARACTER OF BIG DATA
VARIETY
 Big Data isn't just numbers, dates, and strings.
Big Data is also geospatial data, 3D data, audio
and video, and unstructured text, including log
files and social media.
 Traditional database systems were designed to
address smaller volumes of structured data,
fewer updates or a predictable, consistent data
structure.
 Big Data analysis includes different types of data
STORING BIG DATA
 Analyzing your data characteristics
Selecting data sources for analysis
Eliminating redundant data
Establishing the role of NoSQL
 Overview of Big Data stores
Data models: key value, graph, document, column-family
Hadoop Distributed File System
HBase
Hive
PROCESSING BIG DATA
 Integrating disparate data stores
Mapping data to the programming framework
Connecting and extracting data from storage
Transforming data for processing
Subdividing data in preparation for Hadoop MapReduce
 Employing Hadoop MapReduce
Creating the components of Hadoop MapReduce jobs
Distributing data processing across server farms
Executing Hadoop MapReduce jobs
Monitoring the progress of job flows
THE STRUCTURE OF BIG DATA
 Structured
• Most traditional data
sources
 Semi-structured
• Many sources of big data
 Unstructured
• Video data, audio data
11
WHY BIG DATA
• Growth of Big Data is needed
– Increase of storage capacities
– Increase of processing power
– Availability of data(different data types)
WHY BIG DATA
•FB generates 10TB
daily
•Twitter generates 7TB
of data
Daily
•IBM claims 90% of
today’s
stored data was
generated
in just the last two years.
HOW IS BIG DATA DIFFERENT?
1) Automatically generated by a machine
(e.g. Sensor embedded in an engine)
2) Typically an entirely new source of data
(e.g. Use of the internet)
3) Not designed to be friendly
(e.g. Text streams)
14
BIG DATA SOURCES
Users
Application
Systems
Sensors
Large and growing files
(Big data files)
DATA GENERATION POINTS
EXAMPLES
Mobile Devices
Microphones
Readers/Scanners
Science facilities
Programs/ Software
Social Media
Cameras
BIG DATA ANALYTICS
 Examining large amount of data
 Appropriate information (about data)
 Identification of hidden patterns, unknown correlations
 Better business decisions: strategic and operational
 Effective marketing, customer satisfaction, increased
revenue
TYPES OF TOOLS USED IN BIG-DATA
 Where processing is hosted?
Distributed Servers / Cloud (e.g. Amazon EC2)
 Where data is stored?
Distributed Storage (e.g. Amazon S3)
 What is the programming model?
Distributed Processing (e.g. MapReduce)
 How data is stored & indexed?
High-performance schema-free databases (e.g. MongoDB)
 What operations are performed on data?
Analytic / Semantic Processing
Application Of Big Data analytics
Homeland
Security
Smarter
Healthcare Multi-channel
sales
T
elecom
Manufacturing
TrafficControl Trading
Analytics
Search
Quality
RISKS OF BIG DATA
• Will be so overwhelmed
• Need the right people and solve the right problems
• Costs escalate too fast
• Isn’t necessary to capture 100%
• Many sources of big data
is privacy
• self-regulation (data compression)
• Legal regulation
20
HOW BIG DATA IMPACTS ON IT
• Big data is a troublesome force presenting
opportunities with challenges to IT organizations.
 By 2015 4.4 million IT jobs in Big Data ; 1.9 million
is in US itself
 In 2017, Data scientist’s was No. 1 Job in the
Harvard’s ranking.
BENEFITS OF BIG DATA
•Real-time big data isn’t just a process for storing
petabytes or exabytes of data in a data warehouse,
It’s about the ability to make better decisions and take
meaningful actions at the right time.
•Fast forward to the present and technologies like
Hadoop give you the scale and flexibility to store data
before you know how you are going to process it.
•Technologies such as MapReduce,Hive and Impala
enable you to run queries without changing the data
structures underneath.
THANK YOU.

More Related Content

What's hot

Big Data - Applications and Technologies Overview
Big Data - Applications and Technologies OverviewBig Data - Applications and Technologies Overview
Big Data - Applications and Technologies Overview
Sivashankar Ganapathy
 
Big data
Big dataBig data
Big data
madhavsolanki
 
Big data
Big dataBig data
Big data
SaraRao3
 
Big data storage
Big data storageBig data storage
Big data storage
Vikram Nandini
 
BIG DATA-Seminar Report
BIG DATA-Seminar ReportBIG DATA-Seminar Report
BIG DATA-Seminar Report
josnapv
 
Big data Analytics
Big data AnalyticsBig data Analytics
Big data Analytics
ShivanandaVSeeri
 
Introduction to big data
Introduction to big dataIntroduction to big data
Introduction to big data
Hari Priya
 
Presentation About Big Data (DBMS)
Presentation About Big Data (DBMS)Presentation About Big Data (DBMS)
Presentation About Big Data (DBMS)
SiamAhmed16
 
Chapter 1 big data
Chapter 1 big dataChapter 1 big data
Chapter 1 big data
Prof .Pragati Khade
 
What is big data?
What is big data?What is big data?
What is big data?
David Wellman
 
Big Data, Big Deal: For Future Big Data Scientists
Big Data, Big Deal: For Future Big Data ScientistsBig Data, Big Deal: For Future Big Data Scientists
Big Data, Big Deal: For Future Big Data Scientists
Way-Yen Lin
 
Big Data
Big DataBig Data
Big Data
Rohit Jain
 
Data science
Data scienceData science
Data science
Ranjit Nambisan
 
Big data by Mithlesh sadh
Big data by Mithlesh sadhBig data by Mithlesh sadh
Big data by Mithlesh sadh
Mithlesh Sadh
 
Big Data
Big DataBig Data
Big Data
Vinayak Kamath
 
Internet of Things (IoT) and Big Data
Internet of Things (IoT) and Big DataInternet of Things (IoT) and Big Data
Internet of Things (IoT) and Big Data
Guido Schmutz
 
Introduction to data science.pptx
Introduction to data science.pptxIntroduction to data science.pptx
Introduction to data science.pptx
SadhanaParameswaran
 
Introduction to data science
Introduction to data scienceIntroduction to data science
Introduction to data science
Sampath Kumar
 
Introduction to Big Data
Introduction to Big DataIntroduction to Big Data
Introduction to Big Data
Vipin Batra
 

What's hot (20)

Big Data - Applications and Technologies Overview
Big Data - Applications and Technologies OverviewBig Data - Applications and Technologies Overview
Big Data - Applications and Technologies Overview
 
Big data
Big dataBig data
Big data
 
Big data
Big dataBig data
Big data
 
Big data
Big dataBig data
Big data
 
Big data storage
Big data storageBig data storage
Big data storage
 
BIG DATA-Seminar Report
BIG DATA-Seminar ReportBIG DATA-Seminar Report
BIG DATA-Seminar Report
 
Big data Analytics
Big data AnalyticsBig data Analytics
Big data Analytics
 
Introduction to big data
Introduction to big dataIntroduction to big data
Introduction to big data
 
Presentation About Big Data (DBMS)
Presentation About Big Data (DBMS)Presentation About Big Data (DBMS)
Presentation About Big Data (DBMS)
 
Chapter 1 big data
Chapter 1 big dataChapter 1 big data
Chapter 1 big data
 
What is big data?
What is big data?What is big data?
What is big data?
 
Big Data, Big Deal: For Future Big Data Scientists
Big Data, Big Deal: For Future Big Data ScientistsBig Data, Big Deal: For Future Big Data Scientists
Big Data, Big Deal: For Future Big Data Scientists
 
Big Data
Big DataBig Data
Big Data
 
Data science
Data scienceData science
Data science
 
Big data by Mithlesh sadh
Big data by Mithlesh sadhBig data by Mithlesh sadh
Big data by Mithlesh sadh
 
Big Data
Big DataBig Data
Big Data
 
Internet of Things (IoT) and Big Data
Internet of Things (IoT) and Big DataInternet of Things (IoT) and Big Data
Internet of Things (IoT) and Big Data
 
Introduction to data science.pptx
Introduction to data science.pptxIntroduction to data science.pptx
Introduction to data science.pptx
 
Introduction to data science
Introduction to data scienceIntroduction to data science
Introduction to data science
 
Introduction to Big Data
Introduction to Big DataIntroduction to Big Data
Introduction to Big Data
 

Similar to Big data-ppt-

Big data Analytics
Big data Analytics Big data Analytics
Big data Analytics
Guduru Lakshmi Kiranmai
 
Big data
Big dataBig data
Big data
Mahmudul Alam
 
big-data-8722-m8RQ3h1.pptx
big-data-8722-m8RQ3h1.pptxbig-data-8722-m8RQ3h1.pptx
big-data-8722-m8RQ3h1.pptx
VaishnavGhadge1
 
Presentation on Big Data
Presentation on Big DataPresentation on Big Data
Presentation on Big Data
Md. Salman Ahmed
 
Bigdata " new level"
Bigdata " new level"Bigdata " new level"
Bigdata " new level"
Vamshikrishna Goud
 
Big data seminor
Big data seminorBig data seminor
Big data seminor
berasrujana
 
An Overview of BigData
An Overview of BigDataAn Overview of BigData
An Overview of BigData
Valarmathi V
 
Big-Data-Analytics.8592259.powerpoint.pdf
Big-Data-Analytics.8592259.powerpoint.pdfBig-Data-Analytics.8592259.powerpoint.pdf
Big-Data-Analytics.8592259.powerpoint.pdf
rajsharma159890
 
Big Data Driven Solutions to Combat Covid' 19
Big Data Driven Solutions to Combat Covid' 19Big Data Driven Solutions to Combat Covid' 19
Big Data Driven Solutions to Combat Covid' 19
Prof.Balakrishnan S
 
BIG DATA & DATA ANALYTICS
BIG  DATA & DATA  ANALYTICSBIG  DATA & DATA  ANALYTICS
BIG DATA & DATA ANALYTICS
NAGARAJAGIDDE
 
Big data Ppt
Big data PptBig data Ppt
Big data Ppt
Prashant Navatre
 
In memory big data management and processing
In memory big data management and processingIn memory big data management and processing
In memory big data management and processing
Pranav Gontalwar
 
Our big data
Our big dataOur big data
Our big data
uthrarajan
 
new.pptx
new.pptxnew.pptx
Big data
Big dataBig data
Big data
Abhishek Palo
 
Big data
Big dataBig data
Big data
Abhishek Palo
 
bigdatappt.pptx
bigdatappt.pptxbigdatappt.pptx
bigdatappt.pptx
KrishnaTeja570279
 

Similar to Big data-ppt- (20)

Big data Analytics
Big data Analytics Big data Analytics
Big data Analytics
 
Big data
Big dataBig data
Big data
 
1
11
1
 
big-data-8722-m8RQ3h1.pptx
big-data-8722-m8RQ3h1.pptxbig-data-8722-m8RQ3h1.pptx
big-data-8722-m8RQ3h1.pptx
 
Presentation on Big Data
Presentation on Big DataPresentation on Big Data
Presentation on Big Data
 
Bigdata " new level"
Bigdata " new level"Bigdata " new level"
Bigdata " new level"
 
Big Data ppt
Big Data pptBig Data ppt
Big Data ppt
 
Big data seminor
Big data seminorBig data seminor
Big data seminor
 
An Overview of BigData
An Overview of BigDataAn Overview of BigData
An Overview of BigData
 
Big-Data-Analytics.8592259.powerpoint.pdf
Big-Data-Analytics.8592259.powerpoint.pdfBig-Data-Analytics.8592259.powerpoint.pdf
Big-Data-Analytics.8592259.powerpoint.pdf
 
Big Data Driven Solutions to Combat Covid' 19
Big Data Driven Solutions to Combat Covid' 19Big Data Driven Solutions to Combat Covid' 19
Big Data Driven Solutions to Combat Covid' 19
 
BIG DATA & DATA ANALYTICS
BIG  DATA & DATA  ANALYTICSBIG  DATA & DATA  ANALYTICS
BIG DATA & DATA ANALYTICS
 
A Big Data Concept
A Big Data ConceptA Big Data Concept
A Big Data Concept
 
Big data Ppt
Big data PptBig data Ppt
Big data Ppt
 
In memory big data management and processing
In memory big data management and processingIn memory big data management and processing
In memory big data management and processing
 
Our big data
Our big dataOur big data
Our big data
 
new.pptx
new.pptxnew.pptx
new.pptx
 
Big data
Big dataBig data
Big data
 
Big data
Big dataBig data
Big data
 
bigdatappt.pptx
bigdatappt.pptxbigdatappt.pptx
bigdatappt.pptx
 

Recently uploaded

一比一原版(YU毕业证)约克大学毕业证成绩单
一比一原版(YU毕业证)约克大学毕业证成绩单一比一原版(YU毕业证)约克大学毕业证成绩单
一比一原版(YU毕业证)约克大学毕业证成绩单
enxupq
 
SOCRadar Germany 2024 Threat Landscape Report
SOCRadar Germany 2024 Threat Landscape ReportSOCRadar Germany 2024 Threat Landscape Report
SOCRadar Germany 2024 Threat Landscape Report
SOCRadar
 
一比一原版(TWU毕业证)西三一大学毕业证成绩单
一比一原版(TWU毕业证)西三一大学毕业证成绩单一比一原版(TWU毕业证)西三一大学毕业证成绩单
一比一原版(TWU毕业证)西三一大学毕业证成绩单
ocavb
 
Sample_Global Non-invasive Prenatal Testing (NIPT) Market, 2019-2030.pdf
Sample_Global Non-invasive Prenatal Testing (NIPT) Market, 2019-2030.pdfSample_Global Non-invasive Prenatal Testing (NIPT) Market, 2019-2030.pdf
Sample_Global Non-invasive Prenatal Testing (NIPT) Market, 2019-2030.pdf
Linda486226
 
Tabula.io Cheatsheet: automate your data workflows
Tabula.io Cheatsheet: automate your data workflowsTabula.io Cheatsheet: automate your data workflows
Tabula.io Cheatsheet: automate your data workflows
alex933524
 
FP Growth Algorithm and its Applications
FP Growth Algorithm and its ApplicationsFP Growth Algorithm and its Applications
FP Growth Algorithm and its Applications
MaleehaSheikh2
 
Predicting Product Ad Campaign Performance: A Data Analysis Project Presentation
Predicting Product Ad Campaign Performance: A Data Analysis Project PresentationPredicting Product Ad Campaign Performance: A Data Analysis Project Presentation
Predicting Product Ad Campaign Performance: A Data Analysis Project Presentation
Boston Institute of Analytics
 
Levelwise PageRank with Loop-Based Dead End Handling Strategy : SHORT REPORT ...
Levelwise PageRank with Loop-Based Dead End Handling Strategy : SHORT REPORT ...Levelwise PageRank with Loop-Based Dead End Handling Strategy : SHORT REPORT ...
Levelwise PageRank with Loop-Based Dead End Handling Strategy : SHORT REPORT ...
Subhajit Sahu
 
Chatty Kathy - UNC Bootcamp Final Project Presentation - Final Version - 5.23...
Chatty Kathy - UNC Bootcamp Final Project Presentation - Final Version - 5.23...Chatty Kathy - UNC Bootcamp Final Project Presentation - Final Version - 5.23...
Chatty Kathy - UNC Bootcamp Final Project Presentation - Final Version - 5.23...
John Andrews
 
Ch03-Managing the Object-Oriented Information Systems Project a.pdf
Ch03-Managing the Object-Oriented Information Systems Project a.pdfCh03-Managing the Object-Oriented Information Systems Project a.pdf
Ch03-Managing the Object-Oriented Information Systems Project a.pdf
haila53
 
tapal brand analysis PPT slide for comptetive data
tapal brand analysis PPT slide for comptetive datatapal brand analysis PPT slide for comptetive data
tapal brand analysis PPT slide for comptetive data
theahmadsaood
 
Criminal IP - Threat Hunting Webinar.pdf
Criminal IP - Threat Hunting Webinar.pdfCriminal IP - Threat Hunting Webinar.pdf
Criminal IP - Threat Hunting Webinar.pdf
Criminal IP
 
【社内勉強会資料_Octo: An Open-Source Generalist Robot Policy】
【社内勉強会資料_Octo: An Open-Source Generalist Robot Policy】【社内勉強会資料_Octo: An Open-Source Generalist Robot Policy】
【社内勉強会資料_Octo: An Open-Source Generalist Robot Policy】
NABLAS株式会社
 
一比一原版(CU毕业证)卡尔顿大学毕业证成绩单
一比一原版(CU毕业证)卡尔顿大学毕业证成绩单一比一原版(CU毕业证)卡尔顿大学毕业证成绩单
一比一原版(CU毕业证)卡尔顿大学毕业证成绩单
yhkoc
 
做(mqu毕业证书)麦考瑞大学毕业证硕士文凭证书学费发票原版一模一样
做(mqu毕业证书)麦考瑞大学毕业证硕士文凭证书学费发票原版一模一样做(mqu毕业证书)麦考瑞大学毕业证硕士文凭证书学费发票原版一模一样
做(mqu毕业证书)麦考瑞大学毕业证硕士文凭证书学费发票原版一模一样
axoqas
 
Best best suvichar in gujarati english meaning of this sentence as Silk road ...
Best best suvichar in gujarati english meaning of this sentence as Silk road ...Best best suvichar in gujarati english meaning of this sentence as Silk road ...
Best best suvichar in gujarati english meaning of this sentence as Silk road ...
AbhimanyuSinha9
 
一比一原版(CBU毕业证)不列颠海角大学毕业证成绩单
一比一原版(CBU毕业证)不列颠海角大学毕业证成绩单一比一原版(CBU毕业证)不列颠海角大学毕业证成绩单
一比一原版(CBU毕业证)不列颠海角大学毕业证成绩单
nscud
 
Adjusting primitives for graph : SHORT REPORT / NOTES
Adjusting primitives for graph : SHORT REPORT / NOTESAdjusting primitives for graph : SHORT REPORT / NOTES
Adjusting primitives for graph : SHORT REPORT / NOTES
Subhajit Sahu
 
社内勉強会資料_LLM Agents                              .
社内勉強会資料_LLM Agents                              .社内勉強会資料_LLM Agents                              .
社内勉強会資料_LLM Agents                              .
NABLAS株式会社
 
Innovative Methods in Media and Communication Research by Sebastian Kubitschk...
Innovative Methods in Media and Communication Research by Sebastian Kubitschk...Innovative Methods in Media and Communication Research by Sebastian Kubitschk...
Innovative Methods in Media and Communication Research by Sebastian Kubitschk...
correoyaya
 

Recently uploaded (20)

一比一原版(YU毕业证)约克大学毕业证成绩单
一比一原版(YU毕业证)约克大学毕业证成绩单一比一原版(YU毕业证)约克大学毕业证成绩单
一比一原版(YU毕业证)约克大学毕业证成绩单
 
SOCRadar Germany 2024 Threat Landscape Report
SOCRadar Germany 2024 Threat Landscape ReportSOCRadar Germany 2024 Threat Landscape Report
SOCRadar Germany 2024 Threat Landscape Report
 
一比一原版(TWU毕业证)西三一大学毕业证成绩单
一比一原版(TWU毕业证)西三一大学毕业证成绩单一比一原版(TWU毕业证)西三一大学毕业证成绩单
一比一原版(TWU毕业证)西三一大学毕业证成绩单
 
Sample_Global Non-invasive Prenatal Testing (NIPT) Market, 2019-2030.pdf
Sample_Global Non-invasive Prenatal Testing (NIPT) Market, 2019-2030.pdfSample_Global Non-invasive Prenatal Testing (NIPT) Market, 2019-2030.pdf
Sample_Global Non-invasive Prenatal Testing (NIPT) Market, 2019-2030.pdf
 
Tabula.io Cheatsheet: automate your data workflows
Tabula.io Cheatsheet: automate your data workflowsTabula.io Cheatsheet: automate your data workflows
Tabula.io Cheatsheet: automate your data workflows
 
FP Growth Algorithm and its Applications
FP Growth Algorithm and its ApplicationsFP Growth Algorithm and its Applications
FP Growth Algorithm and its Applications
 
Predicting Product Ad Campaign Performance: A Data Analysis Project Presentation
Predicting Product Ad Campaign Performance: A Data Analysis Project PresentationPredicting Product Ad Campaign Performance: A Data Analysis Project Presentation
Predicting Product Ad Campaign Performance: A Data Analysis Project Presentation
 
Levelwise PageRank with Loop-Based Dead End Handling Strategy : SHORT REPORT ...
Levelwise PageRank with Loop-Based Dead End Handling Strategy : SHORT REPORT ...Levelwise PageRank with Loop-Based Dead End Handling Strategy : SHORT REPORT ...
Levelwise PageRank with Loop-Based Dead End Handling Strategy : SHORT REPORT ...
 
Chatty Kathy - UNC Bootcamp Final Project Presentation - Final Version - 5.23...
Chatty Kathy - UNC Bootcamp Final Project Presentation - Final Version - 5.23...Chatty Kathy - UNC Bootcamp Final Project Presentation - Final Version - 5.23...
Chatty Kathy - UNC Bootcamp Final Project Presentation - Final Version - 5.23...
 
Ch03-Managing the Object-Oriented Information Systems Project a.pdf
Ch03-Managing the Object-Oriented Information Systems Project a.pdfCh03-Managing the Object-Oriented Information Systems Project a.pdf
Ch03-Managing the Object-Oriented Information Systems Project a.pdf
 
tapal brand analysis PPT slide for comptetive data
tapal brand analysis PPT slide for comptetive datatapal brand analysis PPT slide for comptetive data
tapal brand analysis PPT slide for comptetive data
 
Criminal IP - Threat Hunting Webinar.pdf
Criminal IP - Threat Hunting Webinar.pdfCriminal IP - Threat Hunting Webinar.pdf
Criminal IP - Threat Hunting Webinar.pdf
 
【社内勉強会資料_Octo: An Open-Source Generalist Robot Policy】
【社内勉強会資料_Octo: An Open-Source Generalist Robot Policy】【社内勉強会資料_Octo: An Open-Source Generalist Robot Policy】
【社内勉強会資料_Octo: An Open-Source Generalist Robot Policy】
 
一比一原版(CU毕业证)卡尔顿大学毕业证成绩单
一比一原版(CU毕业证)卡尔顿大学毕业证成绩单一比一原版(CU毕业证)卡尔顿大学毕业证成绩单
一比一原版(CU毕业证)卡尔顿大学毕业证成绩单
 
做(mqu毕业证书)麦考瑞大学毕业证硕士文凭证书学费发票原版一模一样
做(mqu毕业证书)麦考瑞大学毕业证硕士文凭证书学费发票原版一模一样做(mqu毕业证书)麦考瑞大学毕业证硕士文凭证书学费发票原版一模一样
做(mqu毕业证书)麦考瑞大学毕业证硕士文凭证书学费发票原版一模一样
 
Best best suvichar in gujarati english meaning of this sentence as Silk road ...
Best best suvichar in gujarati english meaning of this sentence as Silk road ...Best best suvichar in gujarati english meaning of this sentence as Silk road ...
Best best suvichar in gujarati english meaning of this sentence as Silk road ...
 
一比一原版(CBU毕业证)不列颠海角大学毕业证成绩单
一比一原版(CBU毕业证)不列颠海角大学毕业证成绩单一比一原版(CBU毕业证)不列颠海角大学毕业证成绩单
一比一原版(CBU毕业证)不列颠海角大学毕业证成绩单
 
Adjusting primitives for graph : SHORT REPORT / NOTES
Adjusting primitives for graph : SHORT REPORT / NOTESAdjusting primitives for graph : SHORT REPORT / NOTES
Adjusting primitives for graph : SHORT REPORT / NOTES
 
社内勉強会資料_LLM Agents                              .
社内勉強会資料_LLM Agents                              .社内勉強会資料_LLM Agents                              .
社内勉強会資料_LLM Agents                              .
 
Innovative Methods in Media and Communication Research by Sebastian Kubitschk...
Innovative Methods in Media and Communication Research by Sebastian Kubitschk...Innovative Methods in Media and Communication Research by Sebastian Kubitschk...
Innovative Methods in Media and Communication Research by Sebastian Kubitschk...
 

Big data-ppt-

  • 2. INTRODUCTION  Big Data may well be the Next Big Thing in the IT world.  Big data burst upon the scene in the first decade of the 21st century.  The first organizations to embrace it were online and startup firms. Firms like Google, eBay, LinkedIn, and Facebook were built around big data from the beginning.  Like many new information technologies, big data can bring about dramatic cost reductions, substantial improvements in the time required to perform a computing task, or new product and service offerings.
  • 3. WHAT IS BIG DATA?  ‘Big Data’ is similar to ‘small data’, but bigger in size  but having data bigger it requires different approaches: Techniques, tools and architecture  an aim to solve new problems or old problems in a better way  Big Data generates value from the storage and processing of very large quantities of digital information that cannot be analyzed with traditional computing techniques.
  • 4. WHAT IS BIG DATA  Walmart handles more than 1 million customer transactions every hour. • Facebook handles 40 billion photos from its user base. • Decoding the human genome originally took 10years to process; now it can be achieved in one week.
  • 5. THREE CHARACTERISTICS OF BIG DATA V3S Volume • Data quantity Velocity • Data Speed Variety • Data Types
  • 6. 1ST CHARACTER OF BIG DATA VOLUME •A typical PC might have had 10 gigabytes of storage in 2000. •Today, Facebook ingests 500 terabytes of new data every day. •Boeing 737 will generate 240 terabytes of flight data during a single flight across the US.
  • 7. 2ND CHARACTER OF BIG DATA VELOCITY  Clickstreams and ad impressions capture user behavior at millions of events per second  high-frequency stock trading algorithms reflect market changes within microseconds  machine to machine processes exchange data between billions of devices  infrastructure and sensors generate massive log data in real-time  on-line gaming systems support millions of concurrent users, each producing multiple inputs per second.
  • 8. 3RD CHARACTER OF BIG DATA VARIETY  Big Data isn't just numbers, dates, and strings. Big Data is also geospatial data, 3D data, audio and video, and unstructured text, including log files and social media.  Traditional database systems were designed to address smaller volumes of structured data, fewer updates or a predictable, consistent data structure.  Big Data analysis includes different types of data
  • 9. STORING BIG DATA  Analyzing your data characteristics Selecting data sources for analysis Eliminating redundant data Establishing the role of NoSQL  Overview of Big Data stores Data models: key value, graph, document, column-family Hadoop Distributed File System HBase Hive
  • 10. PROCESSING BIG DATA  Integrating disparate data stores Mapping data to the programming framework Connecting and extracting data from storage Transforming data for processing Subdividing data in preparation for Hadoop MapReduce  Employing Hadoop MapReduce Creating the components of Hadoop MapReduce jobs Distributing data processing across server farms Executing Hadoop MapReduce jobs Monitoring the progress of job flows
  • 11. THE STRUCTURE OF BIG DATA  Structured • Most traditional data sources  Semi-structured • Many sources of big data  Unstructured • Video data, audio data 11
  • 12. WHY BIG DATA • Growth of Big Data is needed – Increase of storage capacities – Increase of processing power – Availability of data(different data types)
  • 13. WHY BIG DATA •FB generates 10TB daily •Twitter generates 7TB of data Daily •IBM claims 90% of today’s stored data was generated in just the last two years.
  • 14. HOW IS BIG DATA DIFFERENT? 1) Automatically generated by a machine (e.g. Sensor embedded in an engine) 2) Typically an entirely new source of data (e.g. Use of the internet) 3) Not designed to be friendly (e.g. Text streams) 14
  • 15. BIG DATA SOURCES Users Application Systems Sensors Large and growing files (Big data files)
  • 16. DATA GENERATION POINTS EXAMPLES Mobile Devices Microphones Readers/Scanners Science facilities Programs/ Software Social Media Cameras
  • 17. BIG DATA ANALYTICS  Examining large amount of data  Appropriate information (about data)  Identification of hidden patterns, unknown correlations  Better business decisions: strategic and operational  Effective marketing, customer satisfaction, increased revenue
  • 18. TYPES OF TOOLS USED IN BIG-DATA  Where processing is hosted? Distributed Servers / Cloud (e.g. Amazon EC2)  Where data is stored? Distributed Storage (e.g. Amazon S3)  What is the programming model? Distributed Processing (e.g. MapReduce)  How data is stored & indexed? High-performance schema-free databases (e.g. MongoDB)  What operations are performed on data? Analytic / Semantic Processing
  • 19. Application Of Big Data analytics Homeland Security Smarter Healthcare Multi-channel sales T elecom Manufacturing TrafficControl Trading Analytics Search Quality
  • 20. RISKS OF BIG DATA • Will be so overwhelmed • Need the right people and solve the right problems • Costs escalate too fast • Isn’t necessary to capture 100% • Many sources of big data is privacy • self-regulation (data compression) • Legal regulation 20
  • 21. HOW BIG DATA IMPACTS ON IT • Big data is a troublesome force presenting opportunities with challenges to IT organizations.  By 2015 4.4 million IT jobs in Big Data ; 1.9 million is in US itself  In 2017, Data scientist’s was No. 1 Job in the Harvard’s ranking.
  • 22. BENEFITS OF BIG DATA •Real-time big data isn’t just a process for storing petabytes or exabytes of data in a data warehouse, It’s about the ability to make better decisions and take meaningful actions at the right time. •Fast forward to the present and technologies like Hadoop give you the scale and flexibility to store data before you know how you are going to process it. •Technologies such as MapReduce,Hive and Impala enable you to run queries without changing the data structures underneath.