SlideShare a Scribd company logo
HackCodeX Forum
5.06.2023, Riga, Latvia
DATA QUALITY AS A PREREQUISITE
FOR BUSINESS SUCCESS:
WHEN SHOULD I START
TAKING CARE OF IT?
Anastasija Nikiforova
Assistant Professor of Information Systems, Faculty of Science and Technology,
Institute of Computer Science, Chair of Software Engineering, University of Tartu
European Open Science CLoud (EOSC) Task Force “FAIR metrics and data quality”
PHD IN COMPUTER SCIENCE – DATA PROCESSING SYSTEMS AND DATA NETWORKING
RESEARCH INTERESTS: DATA MANAGEMENT WITH A FOCUS ON DATA QUALITY, OPEN
GOVERNMENT DATA, SMART CITY, SOCIETY 5.0, SUSTAINABLE DEVELOPMENT, IOT, HCI,
DIGITIZATION.
✔ASSISTANT PROFESSOR AT THE UNIVERSITY OF TARTU, FACULTY OF SCIENCE AND TECHNOLOGY, INSTITUTE OF COMPUTER SCIENCE,
CHAIR OF SOFTWARE ENGINEERING
✔EUROPEAN OPEN SCIENCE CLOUD TASK FORCE “FAIR METRICS AND DATA QUALITY”
✔EDSC AMBASSADOR (EUROPEAN DIGITAL SKILLS CERTIFICATE, AS PART OF ACTION 9 OF THE DIGITAL EDUCATION ACTION PLAN (2021- 2027) –
JRC/SVQ/2022/OP/0013)
✔IFIP WG8.5 ON ICT AND PUBLIC ADMINISTRATION MEMBER
✔ASSOCIATE MEMBER OF THE LATVIAN OPEN TECHNOLOGY ASSOCIATION
✔EXPERT OF THE LATVIAN COUNCIL OF SCIENCES IN (1) NATURAL SCIENCES – COMPUTER SCIENCE & INFORMATICS, (2) ENGINEERING & TECHNOLOGY-
ELECTRICAL ENGINEERING, ELECTRONICS, ICT, (3) SOCIAL SCIENCES – ECONOMICS & BUSINESS
✔EXPERT OF THE COST – EUROPEAN COOPERATION IN SCIENCE & TECHNOLOGY
✔ASSISTANT PROFESSOR AT THE UNIVERSITY OF TARTU, FACULTY OF SCIENCE AND TECHNOLOGY, INSTITUTE OF COMPUTER SCIENCE,
CHAIR OF SOFTWARE ENGINEERING
✔EUROPEAN OPEN SCIENCE CLOUD TASK FORCE “FAIR METRICS AND DATA QUALITY”
✔EDSC AMBASSADOR (EUROPEAN DIGITAL SKILLS CERTIFICATE, AS PART OF ACTION 9 OF THE DIGITAL EDUCATION ACTION PLAN (2021- 2027) –
JRC/SVQ/2022/OP/0013)
✔IFIP WG8.5 ON ICT AND PUBLIC ADMINISTRATION MEMBER
✔ASSOCIATE MEMBER OF THE LATVIAN OPEN TECHNOLOGY ASSOCIATION
✔EXPERT OF THE LATVIAN COUNCIL OF SCIENCES IN (1) NATURAL SCIENCES – COMPUTER SCIENCE & INFORMATICS, (2) ENGINEERING & TECHNOLOGY-
ELECTRICAL ENGINEERING, ELECTRONICS, ICT, (3) SOCIAL SCIENCES – ECONOMICS & BUSINESS
✔EXPERT OF THE COST – EUROPEAN COOPERATION IN SCIENCE & TECHNOLOGY
✔VISITING RESEARCHER AT THE DELFT UNIVERSITY OF TEHNOLOGY, FACULTY TECHNOLOGY POLICY AND MANAGEMENT (TPM)
✔ASSISTANT PROFESSOR AT THE FACULTY OF COMPUTING, UNIVERSITY OF LATVIA
✔RESEARCHER IN THE INNOVATION LABORATORY, FACULTY OF COMPUTING, UNIVERSITY OF LATVIA
✔IT-EXPERT AT THE LATVIAN BIOMEDICAL RESEARCH AND STUDY CENTRE, BBMRI-ERIC LV NATIONAL NODE
✔ADVISOR FOR THE INSTITUTE FOR SOCIAL AND POLITICAL STUDIES, UNIVERSITY OF LATVIA
✔DATA SECURITY SOLUTIONS, LATVIA
✔VISITING RESEARCHER AT THE DELFT UNIVERSITY OF TEHNOLOGY, FACULTY TECHNOLOGY POLICY AND MANAGEMENT (TPM)
✔ASSISTANT PROFESSOR AT THE FACULTY OF COMPUTING, UNIVERSITY OF LATVIA
✔RESEARCHER IN THE INNOVATION LABORATORY, FACULTY OF COMPUTING, UNIVERSITY OF LATVIA
✔IT-EXPERT AT THE LATVIAN BIOMEDICAL RESEARCH AND STUDY CENTRE, BBMRI-ERIC LV NATIONAL NODE
✔ADVISOR FOR THE INSTITUTE FOR SOCIAL AND POLITICAL STUDIES, UNIVERSITY OF LATVIA
✔DATA SECURITY SOLUTIONS, LATVIA
MOST RECENT EXPERIENCE
PAST EXPERIENCE
https://www.linkedin.com/posts/georgefirican_data-dataquality-datamanagement-activity-7001229524768108544-v-ne/?originalSubdomain=mv
DATA … DATA ARE EVERYWHERE
Sources: Premium Vector | Artificial intelligence logo, icon. vector symbol ai, deep learning blockchain neural network concept. machine learning, artificial intelligence, ai. (freepik.com), Top 10 Successful Data Science Companies in 2023 - Learn | Hevo (hevodata.com),
How to Use Business Intelligence (BI) to Improve Organizational Alignment | Wyn Enterprise (grapecity.com), Machine learning logo - Wi6Labs, Business Intelligence Icon Gráfico por aimagenarium · Creative Fabrica, Open Data – GEOAFRICA,
https://www.gartner.com/en/articles/4-emerging-technologies-you-need-to-know-about?utm_medium=social&utm_source=linkedin&utm_campaign=SM_GB_YOY_GTR_SOC_SF1_SM-SWG&utm_content=&sf267111387=1
DATA … DATA ARE EVERYWHERE
M-Files on Twitter: "Data is the New Oil – Especially in Oil and Gas! https://t.co/zFlrvQqlMs https://t.co/qE3Q4aLNQy" / Twitter
DATA QUALITY - WHAT, WHY, HOW, 10 BEST PRACTICES & MORE - Enterprise Master Data Management • Profisee
https://dataladder.com/the-impact-of-poor-data-quality-risks-challenges-and-solutions/
https://twitter.com/bright_data/status/1346443370718240768
🤨 "Data is the new oil."​ | LinkedIn
Data is the New Oil - HubMeta
Data is the New Oil - HubMeta
NOT REALLY
“DATA IS THE NEW OIL” WHY IT IS NOT?
BUT!
✓
Source: Here's Why Data Is Not The New Oil (forbes.com), Image sources: Oil well – Wikipedia, How do we get oil and gas out of the ground? (world-petroleum.org), Customized Silos For Effective Storage of Food | Nextech Solutions (nextechagrisolutions.com)
DATA, LIKE OIL is a source of power,
and those, who control them,
are establishing themselves as «masters of the universe»,
just as oil barons did 100 years ago
effectively infinitely durable and reusable
treating like oil –storing in siloes, has little benefit & reduces its usefulness
a finite resource
can be replicated indefinitely & moved around the world at
the speed of light, at low cost, through fiber optic networks
OIL
requires huge amounts of resources to be
transported to where it is needed
when used, its energy being lost as heat or light, or
permanently converted into another form (e.g., plastic)
becomes more useful the more it is used - once
processed, data often reveals further applications
as the world’s oil reserves dwindle, extracting
it becomes increasingly difficult and expensive
becoming increasingly available as computer
technology advances
data mining doesn’t intrinsically involve damage to the
environment & exploitation of finite natural resources
*apart from the electricity used to run the system
oil drilling involve causing damage to the natural
environment and exploitation of finite natural resources
“DATA IS THE NEW OIL” WHY IT IS NOT?
✘
Source: Here's Why Data Is Not The New Oil (forbes.com), Image sources: Oil well – Wikipedia, How do we get oil and gas out of the ground? (world-petroleum.org), Customized Silos For Effective Storage of Food | Nextech Solutions (nextechagrisolutions.com)
DATA
✘
✘
✘
✘
IF WE THINK ABOUT DATA AS A POWER SOURCE OR FUEL,
IT WOULD MAKE MORE SENSE TO COMPARE THEM WITH
RENEWABLE SOURCES LIKE THE
SUN, WIND AND TIDES”
-B. Marr, Forbes
Here's Why Data Is Not The New Oil (forbes.com)
Letter from the Editor: Here comes the sun (medicalnewstoday.com), A healthy wind | MIT News | Massachusetts Institute of Technology, Tidal phenomenon: high and low tides | Ponant Magazine
AMONG OTHER “NUANCES”,
DATA QUALITY IS USE-CASE DEPENDENT AND DYNAMIC IN NATURE
“ABSOLUTE DATA QUALITY”
DATA QUALITY LEVEL AT WHICH THE DATA WOULD SATISFY
ALL POSSIBLE USE CASES - IS IMPOSSIBLE TO ACHIEVE,
BUT IT IS A GOAL TO BE PURSUED
Def. 1: FITNESS-FOR-USE
Def. 2: FITNESS-FOR-PURPOSE
Def. 3: FREE OF ERRORS
Def. 1: FITNESS-FOR-USE
Def. 2: FITNESS-FOR-PURPOSE
Def. 3: FREE OF ERRORS
UTILITY*
WARRANTY*
=
=
According to ITIL® 4: the framework for the management of IT-enabled service
ISO def.: THE DEGREE TO WHICH
DATA SATISFIES THE REQUIREMENTS
OF ITS INTENDED PURPOSE
ISO/IEC 25012
IN SIMPLER TERMS… THINK OF WINE…
INTRINSIC - flavor type & intensity
EXTRINSIC - brand, packaging…
Based on ISO 19157,
Langstaff, S. A. (2010). Sensory quality control in the wine industry.
Lacagnina, C., David, R., Nikiforova, A., Kuusniemi, M. E., Cappiello, C., Biehlmaier, O., Wright, L., Schubert, C., Bertino, A., Thiemann, H., & Dennis, R. (2023). Towards a data quality framework for
NOT ONLY ABOUT WHAT, BUT
ALSO ABOUT HOW?
IT IS A PROCESS
NOT ONLY ABOUT WHAT, BUT
ALSO ABOUT HOW?
IT IS A PROCESS –
DATA QUALITY MANAGEMENT PROCESS
DEFINE
MEASURE
ANALYSE
IMPROVE TDQM
DATA QUALITY MANAGEMENT PROCESS
TOTAL DATA QUALITY MANAGEMENT LIFCYCLE (BY MIT)
DEFINE: IDENTIFY RELEVANT DQ DIMENSIONS
MEASURE: PRODUCE DQ METRICS
ANALYSE: IDENTIFY ROOT CAUSES FOR DQ PROBLEMS AND
DETERMINE THE IMPACT OF POOR DQ
IMPROVE: IDENTIFY AND EMPLOY TECHNIQUES FOR
IMPROVING DQ
•Lacagnina, C., David, R., Nikiforova, A., Kuusniemi, M. E., Cappiello, C., Biehlmaier, O., Wright, L.,
Schubert, C., Bertino, A., Thiemann, H., & Dennis, R. (2023). Towards a data quality framework
for EOSC. Zenodo. https://doi.org/10.5281/zenodo.7515816
Source: https://healthinstitute.illinois.edu/connect/news/berd-tips-dimensions-of-data-quality
AVAILABILITY
INTERNAL CONSISTENCY
EXTERNAL CONSISTENCY
ACCESSIBILITY
COMPREHENSIVENESS
INTEGRITY
SEMANTIC ACCURACY
SYNTACTIC ACCURACY
RELEVANCE
BELIEVABILITY
TRUSTWORTHINESS
UNAMBIGUITY
DQ DIMENSIONS
CURRENCY
VOLATILITY
EASE OF UNDERSTANDING
CREDIBILITY
PORTABILITY
RESPONSIVENESS
OBJECTIVITY
REPUTATION
RELIABILITY
AND MANY MORE…
Relevance
Availability
Internal consistency
External consistency
Accessibility
Comprehensiveness
Believability
Integrity
Trustworthiness
Semantic accuracy
Unambiguity
Syntactic accuracy
Source: https://healthinstitute.illinois.edu/connect/news/berd-tips-dimensions-of-data-quality
THERE ARE MORE THAN 100 DATA QUALITY DIMENSIONS
IS THERE ANY COMMONLY ACCEPTED DQ DIMENSION
CLASSIFICATION?
https://iso25000.com/index.php/en/iso-25000-standards/iso-25012/136-iso-iec-2012
ISO 25012
SOFTWARE ENGINEERING — SOFTWARE
PRODUCT QUALITY REQUIREMENTS
AND EVALUATION (SQUARE) — DATA
QUALITY MODEL
DIMENSIONS VARY IN DEFINITION AND SCOPE
ONE AND THE SAME NOTION CAN REFER TO DIFFERENT DIMENSIONS
ONE AND THE SAME DIMENSION CAN HAVE
DIFFERENT NOTIONS [IN DIFFERENT SOURCES]
DATA QUALITY RULES ARE THEN DEFINED
FOR EACH DIMENSION
METRICS ARE THEN SELECTED FOR THEM
SIMPLER
USER-ORIENTED
APPROACH
BASED ON USER DEFINED DATA
QUALITY REQUIREMENTS
✓ STANDARDIZATION, NORMALIZATION AND PARSING
✓ MATCHING / DEDUPLICATION AND MERGING
✓ DATA CLEANSING
✓ VALIDATION
✓ DATA PROFILING / AUDITING
✓ SOME A FEW OF THEM SUPPORT (SEMI-)AUTOMATED DQ RULE RECOGNITION
BASED ON METADATA, BUILT-IN RULES, OR MACHINE LEARNING
DQ TOOLS FOR (SEMI-)AUTOMATED DQM
SO FAR…
DEFINITION USER TIME
DIMENSION
PROCESS PURPOSE
SO FAR…
DEFINITION USER TIME
DIMENSION
PROCESS PURPOSE
WHAT ELSE?
DATA OBJECT
DATASET
DATABASE DATA REPOSITORY INFORMATION SYSTEM
SOFTWARE
NO ONE-SIZE-FITS-ALL
DATA OBJECT
DATASET
DATABASE DATA REPOSITORY INFORMATION SYSTEM
SOFTWARE
DATA OWNER
KNOWN
THIRD-PARTY
NO ONE-SIZE-FITS-ALL
DATA OBJECT
DATASET
DATABASE DATA REPOSITORY INFORMATION SYSTEM
SOFTWARE
DATA STRUCTURE
NO ONE-SIZE-FITS-ALL
STRUCTURED DATA UNSTRUCTURED DATA
SEMI-STRUCTURED DATA
Image sources: https://monkeylearn.com/blog/semi-structured-data/, https://www.pngitem.com/middle/ioJTTbR_organization-structure-icon-png-download-structures-icon-png/
DATA OBJECT
DATASET
DATABASE DATA REPOSITORY INFORMATION SYSTEM
SOFTWARE
DATA WAREHOUSE DATA LAKE
Maybe even something else?
NO ONE-SIZE-FITS-ALL
DATA OBJECT
DATASET
DATABASE DATA REPOSITORY INFORMATION SYSTEM
SOFTWARE
Running Analytics on the Data Lake - The Databricks Blog
NO ONE-SIZE-FITS-ALL
Image source: https://www.grazitti.com/blog/data-lake-vs-data-warehouse-which-one-should-you-go-for/, https://www.qubole.com/data-lakes-vs-data-warehouses-the-co-existence-argument/
SCHEMA ON READ
SCHEMA ON WRITE
“SINGLE SOURCE
OF TRUTH”
Implementing a Data Lake or Data Warehouse Architecture for Business Intelligence? | by Lan Chu | Towards Data Science
NB: EXTRACT-TRANSFORM-LOAD
IS NOT DQM!!!
https://www.slideteam.net/data-lake-it-avoid-data-swamp-in-a-data-lake.html
HOW TO AVOID DATA SWAMP?
Image source: The abstracted future of data engineering | by Justin Gage | Datalogue | Medium
OR HOW TO AVOID GIGO*?
*“GARBAGE IN, GARBAGE OUT”
DATA LAKE FOR BI
BUSINESS DATA LAKE
https://www.capgemini.com/wp-content/uploads/2017/07/pivotal_data_lake_vs_traditional_bi_20140805.pdf
DATA LAKE
+
DATA WRANGLING
[an asset, not a silver bullet]
✔
Source: https://monkeylearn.com/blog/data-wrangling/, https://www.altair.com/what-is-data-wrangling/ , https://pediaa.com/what-is-the-difference-between-data-wrangling-and-data-cleaning
Image source: https://www.google.com/url?sa=i&url=https%3A%2F%2Ftwitter.com%2Frokar9%2Fstatus%2F1452339921629302784&psig=AOvVaw2IUSKtgUWxeaplk56f7CoK&ust=1668004535620000&source=images&cd=vfe&ved=0CA4QjhxqFwoTCJDHwbjnnvsCFQAAAAAdAAAAABAM
THE DATA WRANGLING PROCESS TO PREPARE DATA AND INTEGRATE IT INTO IS
DEPENDING ON THE IS AND THE DESIRED OR REQUIRED TARGET QUALITY*, INDIVIDUAL STEPS
SHOULD BE CARRIED OUT SEVERAL TIMES ➔ !!! DATA WRANGLING IS A CONTINUOUS PROCESS
!!! THAT REPEATS ITSELF REPEATEDLY AT REGULAR INTERVALS.
Information
System
Azeroual, O., Schöpfel, J., Ivanovic, D., & Nikiforova, A. (2022). Combining data lake and
data wrangling for ensuring data quality in CRIS. Procedia Computer Science, 211, 3-16.
DATA LAKE VS DATA WAREHOUSE
HOW TO TAKE
THE ADVANTAGES OF BOTH?
DATA LAKE VS DATA WAREHOUSE
HOW TO TAKE
THE ADVANTAGES OF BOTH?
DATA LAKEHOUSE
DATA LAKEHOUSE IS SEEN AS A COMBINATION OF DATA WAREHOUSING WORKLOADS & DATA LAKE ECONOMICS
Running Analytics on the Data Lake - The Databricks Blog
Running Analytics on the Data Lake - The Databricks Blog, Build a Lake House Architecture on AWS | AWS Big Data Blog (amazon.com), The Data Lakehouse, the Data Warehouse and a Modern Data platform architecture - Microsoft Community Hub
DATA OBJECT
DATASET
DATABASE DATA REPOSITORY INFORMATION SYSTEM
SOFTWARE
Running Analytics on the Data Lake - The Databricks Blog
DATA QUALITY-AWARE SOFTWARE
DEVELOPMENT
&
DATA QUALITY MODEL-BASED TESTING
THINK DATA QUALITY FIRST!!! OR TOWARDS DATA
QUALITY BY DESIGN
Guerra-García, C., Nikiforova, A., Jiménez, S., Perez-Gonzalez, H. G., Ramírez-Torres, M., & Ontañon-
García, L. (2023). ISO/IEC 25012-based methodology for managing data quality requirements in the
development of information systems: Towards Data Quality by Design. Data & Knowledge
Engineering, 145,
DAQUAVORD - A METHODOLOGY FOR PROJECT MANAGEMENT OF DATA QUALITY REQUIREMENTS
SPECIFICATION - AIMED AT ELICITING DQ REQUIREMENTS ARISING FROM DIFFERENT USERS’ VIEWPOINTS
THESE DQ REQUIREMENTS SERVE AS DATA QUALITY SOFTWARE REQUIREMENT AT THE TIME
OF THE DEVELOPMENT OF SOFTWARE THAT TAKES DATA QUALITY INTO ACCOUNT BY
DEFAULT.
IS BASED ON THE VIEWPOINT-ORIENTED REQUIREMENTS DEFINITION (VORD) METHOD, AND
THE LATEST AND MOST GENERALLY ACCEPTED ISO/IEC 25012 STANDARD.
DATA ARTIFACT
WHAT DQM APPROACH DEPENDS ON?
DEFINITION USER
TIME
DIMENSION
PROCESS PURPOSE
MUSK’S TOP PRIORITY: TO IMPROVE THE
PRODUCT…
Q: HOW DOES ONE ENSURE THE RELIABILITY OF DATA
AND DECISIONS MADE BASED ON SAID DATA?
THE ANSWER LIES NOT IN MANAGING THE DATA ALONE,
BUT ALSO THE INFORMATION AROUND AND ABOUT DATA
ACQUISITION, TRANSFORMATIONS AND VISUALIZATION
TO PROVIDE A BETTER UNDERSTANDING AND SUPPORT
DECISION MAKERS
https://www.gqindia.com/get-smart/content/5-things-elon-musk-did-to-become-one-of-the-richest-men-in-the-world
https://www.gqindia.com/get-smart/content/5-things-elon-musk-did-to-become-one-of-the-richest-men-in-the-world
MUSK’S TOP PRIORITY: TO IMPROVE THE
PRODUCT…
Q: HOW DOES ONE ENSURE THE RELIABILITY OF DATA
AND DECISIONS MADE BASED ON SAID DATA?
THE ANSWER LIES NOT IN MANAGING THE DATA ALONE,
BUT ALSO THE INFORMATION AROUND AND ABOUT DATA
ACQUISITION, TRANSFORMATIONS AND VISUALIZATION
TO PROVIDE A BETTER UNDERSTANDING AND SUPPORT
DECISION MAKERS
BY FOCUSING ON SUSTAINABLE DATA, CLEAR
DATA GOVERNANCE
AND STRONG DATA MANAGEMENT
https://www.softcrylic.com/blogs/data-catalogs-in-data-governance/
https://www.gqindia.com/get-smart/content/5-things-elon-musk-did-to-become-one-of-the-richest-men-in-the-world
DATA GOVERNANCE IS THE ANSWER
https://www.edq.com/blog/data-quality-vs-data-governance/
Azeroual O., Nikiforova A., Sha K. (2023) Overlooked Aspects of Data Governance:
Workflow Framework For Enterprise Data Deduplication
https://www.gqindia.com/get-smart/content/5-things-elon-musk-did-to-become-one-of-the-richest-men-in-the-world
DATA GOVERNANCE IS THE ANSWER
https://www.edq.com/blog/data-quality-vs-data-governance/
https://www.gqindia.com/get-smart/content/5-things-elon-musk-did-to-become-one-of-the-richest-men-in-the-world
DATA QUALITY MANAGEMENT IS A CONTINUOUS PROCESS
https://www.gqindia.com/get-smart/content/5-things-elon-musk-did-to-become-one-of-the-richest-men-in-the-world
THINK DATA QUALITY FIRST!
“1-10-100” RULE
1$ SPENT ON PREVENTION
SAVES 10$ ON APPRAISAL AND
100$ ON FAILURE COSTS
https://twitter.com/bright_data/status/1346443370718240768
https://www.gqindia.com/get-smart/content/5-things-elon-musk-did-to-become-one-of-the-richest-men-in-the-world
DEVELOP DATA QUALITY MANAGEMENT AND
GOVERNANCE STRATEGIES
MANTAIN DQM & DQG STRATEGIES
DEFINE
MEASURE
ANALYSE
IMPROVE
+
=
https://starwars.fandom.com/wiki/Destruction_of_Despayre, https://www.linkedin.com/posts/georgefirican_data-dataquality-datamanagement-activity-7001229524768108544-v-ne/?originalSubdomain=mv, History in Objects: Death Star Plans Datacard • Lucasfilm, Video Analysis of an Exploding Death Star | WIRED
+
=
https://starwars.fandom.com/wiki/Destruction_of_Despayre, https://www.linkedin.com/posts/georgefirican_data-dataquality-datamanagement-activity-7001229524768108544-v-ne/?originalSubdomain=mv, History in Objects: Death Star Plans Datacard • Lucasfilm, Video Analysis of an Exploding Death Star | WIRED
For more information, see ResearchGate,
anastasijanikiforova.com
For questions or any queries, contact me via
Nikiforova.Anastasija@gmail.com,

More Related Content

What's hot

Introduction to Data Governance
Introduction to Data GovernanceIntroduction to Data Governance
Introduction to Data Governance
John Bao Vuu
 
Data Architecture Strategies: Data Architecture for Digital Transformation
Data Architecture Strategies: Data Architecture for Digital TransformationData Architecture Strategies: Data Architecture for Digital Transformation
Data Architecture Strategies: Data Architecture for Digital Transformation
DATAVERSITY
 
Strategy For Data Quality
Strategy For Data QualityStrategy For Data Quality
Strategy For Data Quality
Database Answers Ltd.
 
Data-Ed Webinar: Data Quality Success Stories
Data-Ed Webinar: Data Quality Success StoriesData-Ed Webinar: Data Quality Success Stories
Data-Ed Webinar: Data Quality Success Stories
DATAVERSITY
 
Formalize Data Governance with Policies and Procedures
Formalize Data Governance with Policies and ProceduresFormalize Data Governance with Policies and Procedures
Formalize Data Governance with Policies and Procedures
DATAVERSITY
 
Data Quality
Data QualityData Quality
Data Quality
Michael Collins
 
Glossaries, Dictionaries, and Catalogs Result in Data Governance
Glossaries, Dictionaries, and Catalogs Result in Data GovernanceGlossaries, Dictionaries, and Catalogs Result in Data Governance
Glossaries, Dictionaries, and Catalogs Result in Data Governance
DATAVERSITY
 
DAS Slides: Data Governance - Combining Data Management with Organizational ...
DAS Slides: Data Governance -  Combining Data Management with Organizational ...DAS Slides: Data Governance -  Combining Data Management with Organizational ...
DAS Slides: Data Governance - Combining Data Management with Organizational ...
DATAVERSITY
 
Data Governance Takes a Village (So Why is Everyone Hiding?)
Data Governance Takes a Village (So Why is Everyone Hiding?)Data Governance Takes a Village (So Why is Everyone Hiding?)
Data Governance Takes a Village (So Why is Everyone Hiding?)
DATAVERSITY
 
Data Governance Best Practices
Data Governance Best PracticesData Governance Best Practices
Data Governance Best Practices
DATAVERSITY
 
Improving Data Literacy Around Data Architecture
Improving Data Literacy Around Data ArchitectureImproving Data Literacy Around Data Architecture
Improving Data Literacy Around Data Architecture
DATAVERSITY
 
The Importance of Metadata
The Importance of MetadataThe Importance of Metadata
The Importance of Metadata
DATAVERSITY
 
How to Strengthen Enterprise Data Governance with Data Quality
How to Strengthen Enterprise Data Governance with Data QualityHow to Strengthen Enterprise Data Governance with Data Quality
How to Strengthen Enterprise Data Governance with Data Quality
DATAVERSITY
 
Data Quality
Data QualityData Quality
Data Quality
jerdeb
 
Enterprise Architecture vs. Data Architecture
Enterprise Architecture vs. Data ArchitectureEnterprise Architecture vs. Data Architecture
Enterprise Architecture vs. Data Architecture
DATAVERSITY
 
Data Governance and Metadata Management
Data Governance and Metadata ManagementData Governance and Metadata Management
Data Governance and Metadata Management
DATAVERSITY
 
The Role of Data Governance in a Data Strategy
The Role of Data Governance in a Data StrategyThe Role of Data Governance in a Data Strategy
The Role of Data Governance in a Data Strategy
DATAVERSITY
 
Data Management, Metadata Management, and Data Governance – Working Together
Data Management, Metadata Management, and Data Governance – Working TogetherData Management, Metadata Management, and Data Governance – Working Together
Data Management, Metadata Management, and Data Governance – Working Together
DATAVERSITY
 
State of Data Governance in 2021
State of Data Governance in 2021State of Data Governance in 2021
State of Data Governance in 2021
DATAVERSITY
 
Data Governance
Data GovernanceData Governance
Data Governance
Rob Lux
 

What's hot (20)

Introduction to Data Governance
Introduction to Data GovernanceIntroduction to Data Governance
Introduction to Data Governance
 
Data Architecture Strategies: Data Architecture for Digital Transformation
Data Architecture Strategies: Data Architecture for Digital TransformationData Architecture Strategies: Data Architecture for Digital Transformation
Data Architecture Strategies: Data Architecture for Digital Transformation
 
Strategy For Data Quality
Strategy For Data QualityStrategy For Data Quality
Strategy For Data Quality
 
Data-Ed Webinar: Data Quality Success Stories
Data-Ed Webinar: Data Quality Success StoriesData-Ed Webinar: Data Quality Success Stories
Data-Ed Webinar: Data Quality Success Stories
 
Formalize Data Governance with Policies and Procedures
Formalize Data Governance with Policies and ProceduresFormalize Data Governance with Policies and Procedures
Formalize Data Governance with Policies and Procedures
 
Data Quality
Data QualityData Quality
Data Quality
 
Glossaries, Dictionaries, and Catalogs Result in Data Governance
Glossaries, Dictionaries, and Catalogs Result in Data GovernanceGlossaries, Dictionaries, and Catalogs Result in Data Governance
Glossaries, Dictionaries, and Catalogs Result in Data Governance
 
DAS Slides: Data Governance - Combining Data Management with Organizational ...
DAS Slides: Data Governance -  Combining Data Management with Organizational ...DAS Slides: Data Governance -  Combining Data Management with Organizational ...
DAS Slides: Data Governance - Combining Data Management with Organizational ...
 
Data Governance Takes a Village (So Why is Everyone Hiding?)
Data Governance Takes a Village (So Why is Everyone Hiding?)Data Governance Takes a Village (So Why is Everyone Hiding?)
Data Governance Takes a Village (So Why is Everyone Hiding?)
 
Data Governance Best Practices
Data Governance Best PracticesData Governance Best Practices
Data Governance Best Practices
 
Improving Data Literacy Around Data Architecture
Improving Data Literacy Around Data ArchitectureImproving Data Literacy Around Data Architecture
Improving Data Literacy Around Data Architecture
 
The Importance of Metadata
The Importance of MetadataThe Importance of Metadata
The Importance of Metadata
 
How to Strengthen Enterprise Data Governance with Data Quality
How to Strengthen Enterprise Data Governance with Data QualityHow to Strengthen Enterprise Data Governance with Data Quality
How to Strengthen Enterprise Data Governance with Data Quality
 
Data Quality
Data QualityData Quality
Data Quality
 
Enterprise Architecture vs. Data Architecture
Enterprise Architecture vs. Data ArchitectureEnterprise Architecture vs. Data Architecture
Enterprise Architecture vs. Data Architecture
 
Data Governance and Metadata Management
Data Governance and Metadata ManagementData Governance and Metadata Management
Data Governance and Metadata Management
 
The Role of Data Governance in a Data Strategy
The Role of Data Governance in a Data StrategyThe Role of Data Governance in a Data Strategy
The Role of Data Governance in a Data Strategy
 
Data Management, Metadata Management, and Data Governance – Working Together
Data Management, Metadata Management, and Data Governance – Working TogetherData Management, Metadata Management, and Data Governance – Working Together
Data Management, Metadata Management, and Data Governance – Working Together
 
State of Data Governance in 2021
State of Data Governance in 2021State of Data Governance in 2021
State of Data Governance in 2021
 
Data Governance
Data GovernanceData Governance
Data Governance
 

Similar to Data Quality as a prerequisite for you business success: when should I start taking care of it?

Data Quality for AI or AI for Data quality: advances in Data Quality Manageme...
Data Quality for AI or AI for Data quality: advances in Data Quality Manageme...Data Quality for AI or AI for Data quality: advances in Data Quality Manageme...
Data Quality for AI or AI for Data quality: advances in Data Quality Manageme...
Anastasija Nikiforova
 
Public data ecosystems in and for smart cities: how to make open / Big / smar...
Public data ecosystems in and for smart cities: how to make open / Big / smar...Public data ecosystems in and for smart cities: how to make open / Big / smar...
Public data ecosystems in and for smart cities: how to make open / Big / smar...
Anastasija Nikiforova
 
OPEN DATA: ECOSYSTEM, CURRENT AND FUTURE TRENDS, SUCCESS STORIES AND BARRIERS
OPEN DATA: ECOSYSTEM, CURRENT AND FUTURE TRENDS, SUCCESS STORIES AND BARRIERSOPEN DATA: ECOSYSTEM, CURRENT AND FUTURE TRENDS, SUCCESS STORIES AND BARRIERS
OPEN DATA: ECOSYSTEM, CURRENT AND FUTURE TRENDS, SUCCESS STORIES AND BARRIERS
Anastasija Nikiforova
 
BIMCV: The Perfect "Big Data" Storm.
BIMCV: The Perfect "Big Data" Storm. BIMCV: The Perfect "Big Data" Storm.
BIMCV: The Perfect "Big Data" Storm. maigva
 
Smart Data Module 1 introduction to big and smart data
Smart Data Module 1 introduction to big and smart dataSmart Data Module 1 introduction to big and smart data
Smart Data Module 1 introduction to big and smart data
caniceconsulting
 
BIMCV, Banco de Imagen Medica de la Comunidad Valenciana. María de la Iglesia
BIMCV, Banco de Imagen Medica de la Comunidad Valenciana. María de la IglesiaBIMCV, Banco de Imagen Medica de la Comunidad Valenciana. María de la Iglesia
BIMCV, Banco de Imagen Medica de la Comunidad Valenciana. María de la Iglesia
Maria de la Iglesia
 
Data Science and AI in Biomedicine: The World has Changed
Data Science and AI in Biomedicine: The World has ChangedData Science and AI in Biomedicine: The World has Changed
Data Science and AI in Biomedicine: The World has Changed
Philip Bourne
 
2017 11 cascd
2017 11 cascd2017 11 cascd
2017 11 cascd
Johannes Keizer
 
Introduction to Data Science and Analytics
Introduction to Data Science and AnalyticsIntroduction to Data Science and Analytics
Introduction to Data Science and Analytics
Dhruv Saxena
 
Causal networks, learning and inference - Introduction
Causal networks, learning and inference - IntroductionCausal networks, learning and inference - Introduction
Causal networks, learning and inference - Introduction
Fabio Stella
 
Data Science - An emerging Stream of Science with its Spreading Reach & Impact
Data Science - An emerging Stream of Science with its Spreading Reach & ImpactData Science - An emerging Stream of Science with its Spreading Reach & Impact
Data Science - An emerging Stream of Science with its Spreading Reach & Impact
Dr. Sunil Kr. Pandey
 
Challenges and outlook with Big Data
Challenges and outlook with Big Data Challenges and outlook with Big Data
Challenges and outlook with Big Data
IJCERT JOURNAL
 
Mining Big Data using Genetic Algorithm
Mining Big Data using Genetic AlgorithmMining Big Data using Genetic Algorithm
Mining Big Data using Genetic Algorithm
IRJET Journal
 
How Can Public Data Help Your Organization? An Introduction to DataCommons.org
How Can Public Data Help Your Organization? An Introduction to DataCommons.orgHow Can Public Data Help Your Organization? An Introduction to DataCommons.org
How Can Public Data Help Your Organization? An Introduction to DataCommons.org
TechSoup
 
Supervised Multi Attribute Gene Manipulation For Cancer
Supervised Multi Attribute Gene Manipulation For CancerSupervised Multi Attribute Gene Manipulation For Cancer
Supervised Multi Attribute Gene Manipulation For Cancer
paperpublications3
 
Data Science Intro.pptx
Data Science Intro.pptxData Science Intro.pptx
Data Science Intro.pptx
PerumalPitchandi
 
Introduction to Data Science.pptx
Introduction to Data Science.pptxIntroduction to Data Science.pptx
Introduction to Data Science.pptx
PerumalPitchandi
 
Mining Social Media Data for Understanding Drugs Usage
Mining Social Media Data for Understanding Drugs  UsageMining Social Media Data for Understanding Drugs  Usage
Mining Social Media Data for Understanding Drugs Usage
IRJET Journal
 
Introduction to Data Science
Introduction to Data ScienceIntroduction to Data Science
Introduction to Data Science
Laguna State Polytechnic University
 
Cisco service innovation 20110418 v2
Cisco service innovation 20110418 v2Cisco service innovation 20110418 v2
Cisco service innovation 20110418 v2
ISSIP
 

Similar to Data Quality as a prerequisite for you business success: when should I start taking care of it? (20)

Data Quality for AI or AI for Data quality: advances in Data Quality Manageme...
Data Quality for AI or AI for Data quality: advances in Data Quality Manageme...Data Quality for AI or AI for Data quality: advances in Data Quality Manageme...
Data Quality for AI or AI for Data quality: advances in Data Quality Manageme...
 
Public data ecosystems in and for smart cities: how to make open / Big / smar...
Public data ecosystems in and for smart cities: how to make open / Big / smar...Public data ecosystems in and for smart cities: how to make open / Big / smar...
Public data ecosystems in and for smart cities: how to make open / Big / smar...
 
OPEN DATA: ECOSYSTEM, CURRENT AND FUTURE TRENDS, SUCCESS STORIES AND BARRIERS
OPEN DATA: ECOSYSTEM, CURRENT AND FUTURE TRENDS, SUCCESS STORIES AND BARRIERSOPEN DATA: ECOSYSTEM, CURRENT AND FUTURE TRENDS, SUCCESS STORIES AND BARRIERS
OPEN DATA: ECOSYSTEM, CURRENT AND FUTURE TRENDS, SUCCESS STORIES AND BARRIERS
 
BIMCV: The Perfect "Big Data" Storm.
BIMCV: The Perfect "Big Data" Storm. BIMCV: The Perfect "Big Data" Storm.
BIMCV: The Perfect "Big Data" Storm.
 
Smart Data Module 1 introduction to big and smart data
Smart Data Module 1 introduction to big and smart dataSmart Data Module 1 introduction to big and smart data
Smart Data Module 1 introduction to big and smart data
 
BIMCV, Banco de Imagen Medica de la Comunidad Valenciana. María de la Iglesia
BIMCV, Banco de Imagen Medica de la Comunidad Valenciana. María de la IglesiaBIMCV, Banco de Imagen Medica de la Comunidad Valenciana. María de la Iglesia
BIMCV, Banco de Imagen Medica de la Comunidad Valenciana. María de la Iglesia
 
Data Science and AI in Biomedicine: The World has Changed
Data Science and AI in Biomedicine: The World has ChangedData Science and AI in Biomedicine: The World has Changed
Data Science and AI in Biomedicine: The World has Changed
 
2017 11 cascd
2017 11 cascd2017 11 cascd
2017 11 cascd
 
Introduction to Data Science and Analytics
Introduction to Data Science and AnalyticsIntroduction to Data Science and Analytics
Introduction to Data Science and Analytics
 
Causal networks, learning and inference - Introduction
Causal networks, learning and inference - IntroductionCausal networks, learning and inference - Introduction
Causal networks, learning and inference - Introduction
 
Data Science - An emerging Stream of Science with its Spreading Reach & Impact
Data Science - An emerging Stream of Science with its Spreading Reach & ImpactData Science - An emerging Stream of Science with its Spreading Reach & Impact
Data Science - An emerging Stream of Science with its Spreading Reach & Impact
 
Challenges and outlook with Big Data
Challenges and outlook with Big Data Challenges and outlook with Big Data
Challenges and outlook with Big Data
 
Mining Big Data using Genetic Algorithm
Mining Big Data using Genetic AlgorithmMining Big Data using Genetic Algorithm
Mining Big Data using Genetic Algorithm
 
How Can Public Data Help Your Organization? An Introduction to DataCommons.org
How Can Public Data Help Your Organization? An Introduction to DataCommons.orgHow Can Public Data Help Your Organization? An Introduction to DataCommons.org
How Can Public Data Help Your Organization? An Introduction to DataCommons.org
 
Supervised Multi Attribute Gene Manipulation For Cancer
Supervised Multi Attribute Gene Manipulation For CancerSupervised Multi Attribute Gene Manipulation For Cancer
Supervised Multi Attribute Gene Manipulation For Cancer
 
Data Science Intro.pptx
Data Science Intro.pptxData Science Intro.pptx
Data Science Intro.pptx
 
Introduction to Data Science.pptx
Introduction to Data Science.pptxIntroduction to Data Science.pptx
Introduction to Data Science.pptx
 
Mining Social Media Data for Understanding Drugs Usage
Mining Social Media Data for Understanding Drugs  UsageMining Social Media Data for Understanding Drugs  Usage
Mining Social Media Data for Understanding Drugs Usage
 
Introduction to Data Science
Introduction to Data ScienceIntroduction to Data Science
Introduction to Data Science
 
Cisco service innovation 20110418 v2
Cisco service innovation 20110418 v2Cisco service innovation 20110418 v2
Cisco service innovation 20110418 v2
 

More from Anastasija Nikiforova

Towards High-Value Datasets determination for data-driven development: a syst...
Towards High-Value Datasets determination for data-driven development: a syst...Towards High-Value Datasets determination for data-driven development: a syst...
Towards High-Value Datasets determination for data-driven development: a syst...
Anastasija Nikiforova
 
Artificial Intelligence for open data or open data for artificial intelligence?
Artificial Intelligence for open data or open data for artificial intelligence?Artificial Intelligence for open data or open data for artificial intelligence?
Artificial Intelligence for open data or open data for artificial intelligence?
Anastasija Nikiforova
 
Overlooked aspects of data governance: workflow framework for enterprise data...
Overlooked aspects of data governance: workflow framework for enterprise data...Overlooked aspects of data governance: workflow framework for enterprise data...
Overlooked aspects of data governance: workflow framework for enterprise data...
Anastasija Nikiforova
 
Framework for understanding quantum computing use cases from a multidisciplin...
Framework for understanding quantum computing use cases from a multidisciplin...Framework for understanding quantum computing use cases from a multidisciplin...
Framework for understanding quantum computing use cases from a multidisciplin...
Anastasija Nikiforova
 
Data Lake or Data Warehouse? Data Cleaning or Data Wrangling? How to Ensure t...
Data Lake or Data Warehouse? Data Cleaning or Data Wrangling? How to Ensure t...Data Lake or Data Warehouse? Data Cleaning or Data Wrangling? How to Ensure t...
Data Lake or Data Warehouse? Data Cleaning or Data Wrangling? How to Ensure t...
Anastasija Nikiforova
 
Putting FAIR Principles in the Context of Research Information: FAIRness for ...
Putting FAIR Principles in the Context of Research Information: FAIRness for ...Putting FAIR Principles in the Context of Research Information: FAIRness for ...
Putting FAIR Principles in the Context of Research Information: FAIRness for ...
Anastasija Nikiforova
 
Open data hackathon as a tool for increased engagement of Generation Z: to h...
Open data hackathon as a tool for increased engagement of Generation Z:  to h...Open data hackathon as a tool for increased engagement of Generation Z:  to h...
Open data hackathon as a tool for increased engagement of Generation Z: to h...
Anastasija Nikiforova
 
Barriers to Openly Sharing Government Data: Towards an Open Data-adapted Inno...
Barriers to Openly Sharing Government Data: Towards an Open Data-adapted Inno...Barriers to Openly Sharing Government Data: Towards an Open Data-adapted Inno...
Barriers to Openly Sharing Government Data: Towards an Open Data-adapted Inno...
Anastasija Nikiforova
 
Combining Data Lake and Data Wrangling for Ensuring Data Quality in CRIS
Combining Data Lake and Data Wrangling for Ensuring Data Quality in CRISCombining Data Lake and Data Wrangling for Ensuring Data Quality in CRIS
Combining Data Lake and Data Wrangling for Ensuring Data Quality in CRIS
Anastasija Nikiforova
 
The role of open data in the development of sustainable smart cities and smar...
The role of open data in the development of sustainable smart cities and smar...The role of open data in the development of sustainable smart cities and smar...
The role of open data in the development of sustainable smart cities and smar...
Anastasija Nikiforova
 
Data security as a top priority in the digital world: preserve data value by ...
Data security as a top priority in the digital world: preserve data value by ...Data security as a top priority in the digital world: preserve data value by ...
Data security as a top priority in the digital world: preserve data value by ...
Anastasija Nikiforova
 
IoTSE-based Open Database Vulnerability inspection in three Baltic Countries:...
IoTSE-based Open Database Vulnerability inspection in three Baltic Countries:...IoTSE-based Open Database Vulnerability inspection in three Baltic Countries:...
IoTSE-based Open Database Vulnerability inspection in three Baltic Countries:...
Anastasija Nikiforova
 
Stakeholder-centred Identification of Data Quality Issues: Knowledge that Can...
Stakeholder-centred Identification of Data Quality Issues: Knowledge that Can...Stakeholder-centred Identification of Data Quality Issues: Knowledge that Can...
Stakeholder-centred Identification of Data Quality Issues: Knowledge that Can...
Anastasija Nikiforova
 
ShoBeVODSDT: Shodan and Binary Edge based vulnerable open data sources detect...
ShoBeVODSDT: Shodan and Binary Edge based vulnerable open data sources detect...ShoBeVODSDT: Shodan and Binary Edge based vulnerable open data sources detect...
ShoBeVODSDT: Shodan and Binary Edge based vulnerable open data sources detect...
Anastasija Nikiforova
 
Invited talk "Open Data as a driver of Society 5.0: how you and your scientif...
Invited talk "Open Data as a driver of Society 5.0: how you and your scientif...Invited talk "Open Data as a driver of Society 5.0: how you and your scientif...
Invited talk "Open Data as a driver of Society 5.0: how you and your scientif...
Anastasija Nikiforova
 
Towards enrichment of the open government data: a stakeholder-centered determ...
Towards enrichment of the open government data: a stakeholder-centered determ...Towards enrichment of the open government data: a stakeholder-centered determ...
Towards enrichment of the open government data: a stakeholder-centered determ...
Anastasija Nikiforova
 
Atvērto datu potenciāls
Atvērto datu potenciālsAtvērto datu potenciāls
Atvērto datu potenciāls
Anastasija Nikiforova
 
TIMELINESS OF OPEN DATA IN OPEN GOVERNMENT DATA PORTALS THROUGH PANDEMIC-RELA...
TIMELINESS OF OPEN DATA IN OPEN GOVERNMENT DATA PORTALS THROUGH PANDEMIC-RELA...TIMELINESS OF OPEN DATA IN OPEN GOVERNMENT DATA PORTALS THROUGH PANDEMIC-RELA...
TIMELINESS OF OPEN DATA IN OPEN GOVERNMENT DATA PORTALS THROUGH PANDEMIC-RELA...
Anastasija Nikiforova
 
ATVĒRTO DATU SAVLAICĪGUMS NACIONĀLAJOS ATVĒRTO DATU PORTĀLOS AR PANDĒMIJU SAI...
ATVĒRTO DATU SAVLAICĪGUMS NACIONĀLAJOS ATVĒRTO DATU PORTĀLOS AR PANDĒMIJU SAI...ATVĒRTO DATU SAVLAICĪGUMS NACIONĀLAJOS ATVĒRTO DATU PORTĀLOS AR PANDĒMIJU SAI...
ATVĒRTO DATU SAVLAICĪGUMS NACIONĀLAJOS ATVĒRTO DATU PORTĀLOS AR PANDĒMIJU SAI...
Anastasija Nikiforova
 
Towards a Concurrence Analysis in Business Processes
Towards a Concurrence Analysis in Business ProcessesTowards a Concurrence Analysis in Business Processes
Towards a Concurrence Analysis in Business Processes
Anastasija Nikiforova
 

More from Anastasija Nikiforova (20)

Towards High-Value Datasets determination for data-driven development: a syst...
Towards High-Value Datasets determination for data-driven development: a syst...Towards High-Value Datasets determination for data-driven development: a syst...
Towards High-Value Datasets determination for data-driven development: a syst...
 
Artificial Intelligence for open data or open data for artificial intelligence?
Artificial Intelligence for open data or open data for artificial intelligence?Artificial Intelligence for open data or open data for artificial intelligence?
Artificial Intelligence for open data or open data for artificial intelligence?
 
Overlooked aspects of data governance: workflow framework for enterprise data...
Overlooked aspects of data governance: workflow framework for enterprise data...Overlooked aspects of data governance: workflow framework for enterprise data...
Overlooked aspects of data governance: workflow framework for enterprise data...
 
Framework for understanding quantum computing use cases from a multidisciplin...
Framework for understanding quantum computing use cases from a multidisciplin...Framework for understanding quantum computing use cases from a multidisciplin...
Framework for understanding quantum computing use cases from a multidisciplin...
 
Data Lake or Data Warehouse? Data Cleaning or Data Wrangling? How to Ensure t...
Data Lake or Data Warehouse? Data Cleaning or Data Wrangling? How to Ensure t...Data Lake or Data Warehouse? Data Cleaning or Data Wrangling? How to Ensure t...
Data Lake or Data Warehouse? Data Cleaning or Data Wrangling? How to Ensure t...
 
Putting FAIR Principles in the Context of Research Information: FAIRness for ...
Putting FAIR Principles in the Context of Research Information: FAIRness for ...Putting FAIR Principles in the Context of Research Information: FAIRness for ...
Putting FAIR Principles in the Context of Research Information: FAIRness for ...
 
Open data hackathon as a tool for increased engagement of Generation Z: to h...
Open data hackathon as a tool for increased engagement of Generation Z:  to h...Open data hackathon as a tool for increased engagement of Generation Z:  to h...
Open data hackathon as a tool for increased engagement of Generation Z: to h...
 
Barriers to Openly Sharing Government Data: Towards an Open Data-adapted Inno...
Barriers to Openly Sharing Government Data: Towards an Open Data-adapted Inno...Barriers to Openly Sharing Government Data: Towards an Open Data-adapted Inno...
Barriers to Openly Sharing Government Data: Towards an Open Data-adapted Inno...
 
Combining Data Lake and Data Wrangling for Ensuring Data Quality in CRIS
Combining Data Lake and Data Wrangling for Ensuring Data Quality in CRISCombining Data Lake and Data Wrangling for Ensuring Data Quality in CRIS
Combining Data Lake and Data Wrangling for Ensuring Data Quality in CRIS
 
The role of open data in the development of sustainable smart cities and smar...
The role of open data in the development of sustainable smart cities and smar...The role of open data in the development of sustainable smart cities and smar...
The role of open data in the development of sustainable smart cities and smar...
 
Data security as a top priority in the digital world: preserve data value by ...
Data security as a top priority in the digital world: preserve data value by ...Data security as a top priority in the digital world: preserve data value by ...
Data security as a top priority in the digital world: preserve data value by ...
 
IoTSE-based Open Database Vulnerability inspection in three Baltic Countries:...
IoTSE-based Open Database Vulnerability inspection in three Baltic Countries:...IoTSE-based Open Database Vulnerability inspection in three Baltic Countries:...
IoTSE-based Open Database Vulnerability inspection in three Baltic Countries:...
 
Stakeholder-centred Identification of Data Quality Issues: Knowledge that Can...
Stakeholder-centred Identification of Data Quality Issues: Knowledge that Can...Stakeholder-centred Identification of Data Quality Issues: Knowledge that Can...
Stakeholder-centred Identification of Data Quality Issues: Knowledge that Can...
 
ShoBeVODSDT: Shodan and Binary Edge based vulnerable open data sources detect...
ShoBeVODSDT: Shodan and Binary Edge based vulnerable open data sources detect...ShoBeVODSDT: Shodan and Binary Edge based vulnerable open data sources detect...
ShoBeVODSDT: Shodan and Binary Edge based vulnerable open data sources detect...
 
Invited talk "Open Data as a driver of Society 5.0: how you and your scientif...
Invited talk "Open Data as a driver of Society 5.0: how you and your scientif...Invited talk "Open Data as a driver of Society 5.0: how you and your scientif...
Invited talk "Open Data as a driver of Society 5.0: how you and your scientif...
 
Towards enrichment of the open government data: a stakeholder-centered determ...
Towards enrichment of the open government data: a stakeholder-centered determ...Towards enrichment of the open government data: a stakeholder-centered determ...
Towards enrichment of the open government data: a stakeholder-centered determ...
 
Atvērto datu potenciāls
Atvērto datu potenciālsAtvērto datu potenciāls
Atvērto datu potenciāls
 
TIMELINESS OF OPEN DATA IN OPEN GOVERNMENT DATA PORTALS THROUGH PANDEMIC-RELA...
TIMELINESS OF OPEN DATA IN OPEN GOVERNMENT DATA PORTALS THROUGH PANDEMIC-RELA...TIMELINESS OF OPEN DATA IN OPEN GOVERNMENT DATA PORTALS THROUGH PANDEMIC-RELA...
TIMELINESS OF OPEN DATA IN OPEN GOVERNMENT DATA PORTALS THROUGH PANDEMIC-RELA...
 
ATVĒRTO DATU SAVLAICĪGUMS NACIONĀLAJOS ATVĒRTO DATU PORTĀLOS AR PANDĒMIJU SAI...
ATVĒRTO DATU SAVLAICĪGUMS NACIONĀLAJOS ATVĒRTO DATU PORTĀLOS AR PANDĒMIJU SAI...ATVĒRTO DATU SAVLAICĪGUMS NACIONĀLAJOS ATVĒRTO DATU PORTĀLOS AR PANDĒMIJU SAI...
ATVĒRTO DATU SAVLAICĪGUMS NACIONĀLAJOS ATVĒRTO DATU PORTĀLOS AR PANDĒMIJU SAI...
 
Towards a Concurrence Analysis in Business Processes
Towards a Concurrence Analysis in Business ProcessesTowards a Concurrence Analysis in Business Processes
Towards a Concurrence Analysis in Business Processes
 

Recently uploaded

Quality defects in TMT Bars, Possible causes and Potential Solutions.
Quality defects in TMT Bars, Possible causes and Potential Solutions.Quality defects in TMT Bars, Possible causes and Potential Solutions.
Quality defects in TMT Bars, Possible causes and Potential Solutions.
PrashantGoswami42
 
Pile Foundation by Venkatesh Taduvai (Sub Geotechnical Engineering II)-conver...
Pile Foundation by Venkatesh Taduvai (Sub Geotechnical Engineering II)-conver...Pile Foundation by Venkatesh Taduvai (Sub Geotechnical Engineering II)-conver...
Pile Foundation by Venkatesh Taduvai (Sub Geotechnical Engineering II)-conver...
AJAYKUMARPUND1
 
Final project report on grocery store management system..pdf
Final project report on grocery store management system..pdfFinal project report on grocery store management system..pdf
Final project report on grocery store management system..pdf
Kamal Acharya
 
weather web application report.pdf
weather web application report.pdfweather web application report.pdf
weather web application report.pdf
Pratik Pawar
 
Railway Signalling Principles Edition 3.pdf
Railway Signalling Principles Edition 3.pdfRailway Signalling Principles Edition 3.pdf
Railway Signalling Principles Edition 3.pdf
TeeVichai
 
Courier management system project report.pdf
Courier management system project report.pdfCourier management system project report.pdf
Courier management system project report.pdf
Kamal Acharya
 
Vaccine management system project report documentation..pdf
Vaccine management system project report documentation..pdfVaccine management system project report documentation..pdf
Vaccine management system project report documentation..pdf
Kamal Acharya
 
ASME IX(9) 2007 Full Version .pdf
ASME IX(9)  2007 Full Version       .pdfASME IX(9)  2007 Full Version       .pdf
ASME IX(9) 2007 Full Version .pdf
AhmedHussein950959
 
Democratizing Fuzzing at Scale by Abhishek Arya
Democratizing Fuzzing at Scale by Abhishek AryaDemocratizing Fuzzing at Scale by Abhishek Arya
Democratizing Fuzzing at Scale by Abhishek Arya
abh.arya
 
Hybrid optimization of pumped hydro system and solar- Engr. Abdul-Azeez.pdf
Hybrid optimization of pumped hydro system and solar- Engr. Abdul-Azeez.pdfHybrid optimization of pumped hydro system and solar- Engr. Abdul-Azeez.pdf
Hybrid optimization of pumped hydro system and solar- Engr. Abdul-Azeez.pdf
fxintegritypublishin
 
ethical hacking-mobile hacking methods.ppt
ethical hacking-mobile hacking methods.pptethical hacking-mobile hacking methods.ppt
ethical hacking-mobile hacking methods.ppt
Jayaprasanna4
 
The role of big data in decision making.
The role of big data in decision making.The role of big data in decision making.
The role of big data in decision making.
ankuprajapati0525
 
The Benefits and Techniques of Trenchless Pipe Repair.pdf
The Benefits and Techniques of Trenchless Pipe Repair.pdfThe Benefits and Techniques of Trenchless Pipe Repair.pdf
The Benefits and Techniques of Trenchless Pipe Repair.pdf
Pipe Restoration Solutions
 
CFD Simulation of By-pass Flow in a HRSG module by R&R Consult.pptx
CFD Simulation of By-pass Flow in a HRSG module by R&R Consult.pptxCFD Simulation of By-pass Flow in a HRSG module by R&R Consult.pptx
CFD Simulation of By-pass Flow in a HRSG module by R&R Consult.pptx
R&R Consult
 
LIGA(E)11111111111111111111111111111111111111111.ppt
LIGA(E)11111111111111111111111111111111111111111.pptLIGA(E)11111111111111111111111111111111111111111.ppt
LIGA(E)11111111111111111111111111111111111111111.ppt
ssuser9bd3ba
 
HYDROPOWER - Hydroelectric power generation
HYDROPOWER - Hydroelectric power generationHYDROPOWER - Hydroelectric power generation
HYDROPOWER - Hydroelectric power generation
Robbie Edward Sayers
 
Architectural Portfolio Sean Lockwood
Architectural Portfolio Sean LockwoodArchitectural Portfolio Sean Lockwood
Architectural Portfolio Sean Lockwood
seandesed
 
DESIGN A COTTON SEED SEPARATION MACHINE.docx
DESIGN A COTTON SEED SEPARATION MACHINE.docxDESIGN A COTTON SEED SEPARATION MACHINE.docx
DESIGN A COTTON SEED SEPARATION MACHINE.docx
FluxPrime1
 
CME397 Surface Engineering- Professional Elective
CME397 Surface Engineering- Professional ElectiveCME397 Surface Engineering- Professional Elective
CME397 Surface Engineering- Professional Elective
karthi keyan
 
Nuclear Power Economics and Structuring 2024
Nuclear Power Economics and Structuring 2024Nuclear Power Economics and Structuring 2024
Nuclear Power Economics and Structuring 2024
Massimo Talia
 

Recently uploaded (20)

Quality defects in TMT Bars, Possible causes and Potential Solutions.
Quality defects in TMT Bars, Possible causes and Potential Solutions.Quality defects in TMT Bars, Possible causes and Potential Solutions.
Quality defects in TMT Bars, Possible causes and Potential Solutions.
 
Pile Foundation by Venkatesh Taduvai (Sub Geotechnical Engineering II)-conver...
Pile Foundation by Venkatesh Taduvai (Sub Geotechnical Engineering II)-conver...Pile Foundation by Venkatesh Taduvai (Sub Geotechnical Engineering II)-conver...
Pile Foundation by Venkatesh Taduvai (Sub Geotechnical Engineering II)-conver...
 
Final project report on grocery store management system..pdf
Final project report on grocery store management system..pdfFinal project report on grocery store management system..pdf
Final project report on grocery store management system..pdf
 
weather web application report.pdf
weather web application report.pdfweather web application report.pdf
weather web application report.pdf
 
Railway Signalling Principles Edition 3.pdf
Railway Signalling Principles Edition 3.pdfRailway Signalling Principles Edition 3.pdf
Railway Signalling Principles Edition 3.pdf
 
Courier management system project report.pdf
Courier management system project report.pdfCourier management system project report.pdf
Courier management system project report.pdf
 
Vaccine management system project report documentation..pdf
Vaccine management system project report documentation..pdfVaccine management system project report documentation..pdf
Vaccine management system project report documentation..pdf
 
ASME IX(9) 2007 Full Version .pdf
ASME IX(9)  2007 Full Version       .pdfASME IX(9)  2007 Full Version       .pdf
ASME IX(9) 2007 Full Version .pdf
 
Democratizing Fuzzing at Scale by Abhishek Arya
Democratizing Fuzzing at Scale by Abhishek AryaDemocratizing Fuzzing at Scale by Abhishek Arya
Democratizing Fuzzing at Scale by Abhishek Arya
 
Hybrid optimization of pumped hydro system and solar- Engr. Abdul-Azeez.pdf
Hybrid optimization of pumped hydro system and solar- Engr. Abdul-Azeez.pdfHybrid optimization of pumped hydro system and solar- Engr. Abdul-Azeez.pdf
Hybrid optimization of pumped hydro system and solar- Engr. Abdul-Azeez.pdf
 
ethical hacking-mobile hacking methods.ppt
ethical hacking-mobile hacking methods.pptethical hacking-mobile hacking methods.ppt
ethical hacking-mobile hacking methods.ppt
 
The role of big data in decision making.
The role of big data in decision making.The role of big data in decision making.
The role of big data in decision making.
 
The Benefits and Techniques of Trenchless Pipe Repair.pdf
The Benefits and Techniques of Trenchless Pipe Repair.pdfThe Benefits and Techniques of Trenchless Pipe Repair.pdf
The Benefits and Techniques of Trenchless Pipe Repair.pdf
 
CFD Simulation of By-pass Flow in a HRSG module by R&R Consult.pptx
CFD Simulation of By-pass Flow in a HRSG module by R&R Consult.pptxCFD Simulation of By-pass Flow in a HRSG module by R&R Consult.pptx
CFD Simulation of By-pass Flow in a HRSG module by R&R Consult.pptx
 
LIGA(E)11111111111111111111111111111111111111111.ppt
LIGA(E)11111111111111111111111111111111111111111.pptLIGA(E)11111111111111111111111111111111111111111.ppt
LIGA(E)11111111111111111111111111111111111111111.ppt
 
HYDROPOWER - Hydroelectric power generation
HYDROPOWER - Hydroelectric power generationHYDROPOWER - Hydroelectric power generation
HYDROPOWER - Hydroelectric power generation
 
Architectural Portfolio Sean Lockwood
Architectural Portfolio Sean LockwoodArchitectural Portfolio Sean Lockwood
Architectural Portfolio Sean Lockwood
 
DESIGN A COTTON SEED SEPARATION MACHINE.docx
DESIGN A COTTON SEED SEPARATION MACHINE.docxDESIGN A COTTON SEED SEPARATION MACHINE.docx
DESIGN A COTTON SEED SEPARATION MACHINE.docx
 
CME397 Surface Engineering- Professional Elective
CME397 Surface Engineering- Professional ElectiveCME397 Surface Engineering- Professional Elective
CME397 Surface Engineering- Professional Elective
 
Nuclear Power Economics and Structuring 2024
Nuclear Power Economics and Structuring 2024Nuclear Power Economics and Structuring 2024
Nuclear Power Economics and Structuring 2024
 

Data Quality as a prerequisite for you business success: when should I start taking care of it?

  • 1. HackCodeX Forum 5.06.2023, Riga, Latvia DATA QUALITY AS A PREREQUISITE FOR BUSINESS SUCCESS: WHEN SHOULD I START TAKING CARE OF IT? Anastasija Nikiforova Assistant Professor of Information Systems, Faculty of Science and Technology, Institute of Computer Science, Chair of Software Engineering, University of Tartu European Open Science CLoud (EOSC) Task Force “FAIR metrics and data quality”
  • 2. PHD IN COMPUTER SCIENCE – DATA PROCESSING SYSTEMS AND DATA NETWORKING RESEARCH INTERESTS: DATA MANAGEMENT WITH A FOCUS ON DATA QUALITY, OPEN GOVERNMENT DATA, SMART CITY, SOCIETY 5.0, SUSTAINABLE DEVELOPMENT, IOT, HCI, DIGITIZATION. ✔ASSISTANT PROFESSOR AT THE UNIVERSITY OF TARTU, FACULTY OF SCIENCE AND TECHNOLOGY, INSTITUTE OF COMPUTER SCIENCE, CHAIR OF SOFTWARE ENGINEERING ✔EUROPEAN OPEN SCIENCE CLOUD TASK FORCE “FAIR METRICS AND DATA QUALITY” ✔EDSC AMBASSADOR (EUROPEAN DIGITAL SKILLS CERTIFICATE, AS PART OF ACTION 9 OF THE DIGITAL EDUCATION ACTION PLAN (2021- 2027) – JRC/SVQ/2022/OP/0013) ✔IFIP WG8.5 ON ICT AND PUBLIC ADMINISTRATION MEMBER ✔ASSOCIATE MEMBER OF THE LATVIAN OPEN TECHNOLOGY ASSOCIATION ✔EXPERT OF THE LATVIAN COUNCIL OF SCIENCES IN (1) NATURAL SCIENCES – COMPUTER SCIENCE & INFORMATICS, (2) ENGINEERING & TECHNOLOGY- ELECTRICAL ENGINEERING, ELECTRONICS, ICT, (3) SOCIAL SCIENCES – ECONOMICS & BUSINESS ✔EXPERT OF THE COST – EUROPEAN COOPERATION IN SCIENCE & TECHNOLOGY ✔ASSISTANT PROFESSOR AT THE UNIVERSITY OF TARTU, FACULTY OF SCIENCE AND TECHNOLOGY, INSTITUTE OF COMPUTER SCIENCE, CHAIR OF SOFTWARE ENGINEERING ✔EUROPEAN OPEN SCIENCE CLOUD TASK FORCE “FAIR METRICS AND DATA QUALITY” ✔EDSC AMBASSADOR (EUROPEAN DIGITAL SKILLS CERTIFICATE, AS PART OF ACTION 9 OF THE DIGITAL EDUCATION ACTION PLAN (2021- 2027) – JRC/SVQ/2022/OP/0013) ✔IFIP WG8.5 ON ICT AND PUBLIC ADMINISTRATION MEMBER ✔ASSOCIATE MEMBER OF THE LATVIAN OPEN TECHNOLOGY ASSOCIATION ✔EXPERT OF THE LATVIAN COUNCIL OF SCIENCES IN (1) NATURAL SCIENCES – COMPUTER SCIENCE & INFORMATICS, (2) ENGINEERING & TECHNOLOGY- ELECTRICAL ENGINEERING, ELECTRONICS, ICT, (3) SOCIAL SCIENCES – ECONOMICS & BUSINESS ✔EXPERT OF THE COST – EUROPEAN COOPERATION IN SCIENCE & TECHNOLOGY ✔VISITING RESEARCHER AT THE DELFT UNIVERSITY OF TEHNOLOGY, FACULTY TECHNOLOGY POLICY AND MANAGEMENT (TPM) ✔ASSISTANT PROFESSOR AT THE FACULTY OF COMPUTING, UNIVERSITY OF LATVIA ✔RESEARCHER IN THE INNOVATION LABORATORY, FACULTY OF COMPUTING, UNIVERSITY OF LATVIA ✔IT-EXPERT AT THE LATVIAN BIOMEDICAL RESEARCH AND STUDY CENTRE, BBMRI-ERIC LV NATIONAL NODE ✔ADVISOR FOR THE INSTITUTE FOR SOCIAL AND POLITICAL STUDIES, UNIVERSITY OF LATVIA ✔DATA SECURITY SOLUTIONS, LATVIA ✔VISITING RESEARCHER AT THE DELFT UNIVERSITY OF TEHNOLOGY, FACULTY TECHNOLOGY POLICY AND MANAGEMENT (TPM) ✔ASSISTANT PROFESSOR AT THE FACULTY OF COMPUTING, UNIVERSITY OF LATVIA ✔RESEARCHER IN THE INNOVATION LABORATORY, FACULTY OF COMPUTING, UNIVERSITY OF LATVIA ✔IT-EXPERT AT THE LATVIAN BIOMEDICAL RESEARCH AND STUDY CENTRE, BBMRI-ERIC LV NATIONAL NODE ✔ADVISOR FOR THE INSTITUTE FOR SOCIAL AND POLITICAL STUDIES, UNIVERSITY OF LATVIA ✔DATA SECURITY SOLUTIONS, LATVIA MOST RECENT EXPERIENCE PAST EXPERIENCE
  • 4.
  • 5.
  • 6. DATA … DATA ARE EVERYWHERE Sources: Premium Vector | Artificial intelligence logo, icon. vector symbol ai, deep learning blockchain neural network concept. machine learning, artificial intelligence, ai. (freepik.com), Top 10 Successful Data Science Companies in 2023 - Learn | Hevo (hevodata.com), How to Use Business Intelligence (BI) to Improve Organizational Alignment | Wyn Enterprise (grapecity.com), Machine learning logo - Wi6Labs, Business Intelligence Icon Gráfico por aimagenarium · Creative Fabrica, Open Data – GEOAFRICA, https://www.gartner.com/en/articles/4-emerging-technologies-you-need-to-know-about?utm_medium=social&utm_source=linkedin&utm_campaign=SM_GB_YOY_GTR_SOC_SF1_SM-SWG&utm_content=&sf267111387=1
  • 7. DATA … DATA ARE EVERYWHERE M-Files on Twitter: "Data is the New Oil – Especially in Oil and Gas! https://t.co/zFlrvQqlMs https://t.co/qE3Q4aLNQy" / Twitter
  • 8. DATA QUALITY - WHAT, WHY, HOW, 10 BEST PRACTICES & MORE - Enterprise Master Data Management • Profisee
  • 11. 🤨 "Data is the new oil."​ | LinkedIn
  • 12. Data is the New Oil - HubMeta
  • 13. Data is the New Oil - HubMeta NOT REALLY
  • 14. “DATA IS THE NEW OIL” WHY IT IS NOT? BUT! ✓ Source: Here's Why Data Is Not The New Oil (forbes.com), Image sources: Oil well – Wikipedia, How do we get oil and gas out of the ground? (world-petroleum.org), Customized Silos For Effective Storage of Food | Nextech Solutions (nextechagrisolutions.com) DATA, LIKE OIL is a source of power, and those, who control them, are establishing themselves as «masters of the universe», just as oil barons did 100 years ago
  • 15. effectively infinitely durable and reusable treating like oil –storing in siloes, has little benefit & reduces its usefulness a finite resource can be replicated indefinitely & moved around the world at the speed of light, at low cost, through fiber optic networks OIL requires huge amounts of resources to be transported to where it is needed when used, its energy being lost as heat or light, or permanently converted into another form (e.g., plastic) becomes more useful the more it is used - once processed, data often reveals further applications as the world’s oil reserves dwindle, extracting it becomes increasingly difficult and expensive becoming increasingly available as computer technology advances data mining doesn’t intrinsically involve damage to the environment & exploitation of finite natural resources *apart from the electricity used to run the system oil drilling involve causing damage to the natural environment and exploitation of finite natural resources “DATA IS THE NEW OIL” WHY IT IS NOT? ✘ Source: Here's Why Data Is Not The New Oil (forbes.com), Image sources: Oil well – Wikipedia, How do we get oil and gas out of the ground? (world-petroleum.org), Customized Silos For Effective Storage of Food | Nextech Solutions (nextechagrisolutions.com) DATA ✘ ✘ ✘ ✘
  • 16. IF WE THINK ABOUT DATA AS A POWER SOURCE OR FUEL, IT WOULD MAKE MORE SENSE TO COMPARE THEM WITH RENEWABLE SOURCES LIKE THE SUN, WIND AND TIDES” -B. Marr, Forbes Here's Why Data Is Not The New Oil (forbes.com) Letter from the Editor: Here comes the sun (medicalnewstoday.com), A healthy wind | MIT News | Massachusetts Institute of Technology, Tidal phenomenon: high and low tides | Ponant Magazine
  • 17. AMONG OTHER “NUANCES”, DATA QUALITY IS USE-CASE DEPENDENT AND DYNAMIC IN NATURE “ABSOLUTE DATA QUALITY” DATA QUALITY LEVEL AT WHICH THE DATA WOULD SATISFY ALL POSSIBLE USE CASES - IS IMPOSSIBLE TO ACHIEVE, BUT IT IS A GOAL TO BE PURSUED
  • 18.
  • 19. Def. 1: FITNESS-FOR-USE Def. 2: FITNESS-FOR-PURPOSE Def. 3: FREE OF ERRORS
  • 20. Def. 1: FITNESS-FOR-USE Def. 2: FITNESS-FOR-PURPOSE Def. 3: FREE OF ERRORS UTILITY* WARRANTY* = = According to ITIL® 4: the framework for the management of IT-enabled service
  • 21. ISO def.: THE DEGREE TO WHICH DATA SATISFIES THE REQUIREMENTS OF ITS INTENDED PURPOSE ISO/IEC 25012
  • 22. IN SIMPLER TERMS… THINK OF WINE… INTRINSIC - flavor type & intensity EXTRINSIC - brand, packaging… Based on ISO 19157, Langstaff, S. A. (2010). Sensory quality control in the wine industry. Lacagnina, C., David, R., Nikiforova, A., Kuusniemi, M. E., Cappiello, C., Biehlmaier, O., Wright, L., Schubert, C., Bertino, A., Thiemann, H., & Dennis, R. (2023). Towards a data quality framework for
  • 23.
  • 24. NOT ONLY ABOUT WHAT, BUT ALSO ABOUT HOW? IT IS A PROCESS
  • 25. NOT ONLY ABOUT WHAT, BUT ALSO ABOUT HOW? IT IS A PROCESS – DATA QUALITY MANAGEMENT PROCESS
  • 26.
  • 27. DEFINE MEASURE ANALYSE IMPROVE TDQM DATA QUALITY MANAGEMENT PROCESS TOTAL DATA QUALITY MANAGEMENT LIFCYCLE (BY MIT) DEFINE: IDENTIFY RELEVANT DQ DIMENSIONS MEASURE: PRODUCE DQ METRICS ANALYSE: IDENTIFY ROOT CAUSES FOR DQ PROBLEMS AND DETERMINE THE IMPACT OF POOR DQ IMPROVE: IDENTIFY AND EMPLOY TECHNIQUES FOR IMPROVING DQ
  • 28. •Lacagnina, C., David, R., Nikiforova, A., Kuusniemi, M. E., Cappiello, C., Biehlmaier, O., Wright, L., Schubert, C., Bertino, A., Thiemann, H., & Dennis, R. (2023). Towards a data quality framework for EOSC. Zenodo. https://doi.org/10.5281/zenodo.7515816
  • 29. Source: https://healthinstitute.illinois.edu/connect/news/berd-tips-dimensions-of-data-quality AVAILABILITY INTERNAL CONSISTENCY EXTERNAL CONSISTENCY ACCESSIBILITY COMPREHENSIVENESS INTEGRITY SEMANTIC ACCURACY SYNTACTIC ACCURACY RELEVANCE BELIEVABILITY TRUSTWORTHINESS UNAMBIGUITY DQ DIMENSIONS CURRENCY VOLATILITY EASE OF UNDERSTANDING CREDIBILITY PORTABILITY RESPONSIVENESS OBJECTIVITY REPUTATION RELIABILITY AND MANY MORE…
  • 30. Relevance Availability Internal consistency External consistency Accessibility Comprehensiveness Believability Integrity Trustworthiness Semantic accuracy Unambiguity Syntactic accuracy Source: https://healthinstitute.illinois.edu/connect/news/berd-tips-dimensions-of-data-quality THERE ARE MORE THAN 100 DATA QUALITY DIMENSIONS
  • 31. IS THERE ANY COMMONLY ACCEPTED DQ DIMENSION CLASSIFICATION? https://iso25000.com/index.php/en/iso-25000-standards/iso-25012/136-iso-iec-2012 ISO 25012 SOFTWARE ENGINEERING — SOFTWARE PRODUCT QUALITY REQUIREMENTS AND EVALUATION (SQUARE) — DATA QUALITY MODEL
  • 32. DIMENSIONS VARY IN DEFINITION AND SCOPE ONE AND THE SAME NOTION CAN REFER TO DIFFERENT DIMENSIONS ONE AND THE SAME DIMENSION CAN HAVE DIFFERENT NOTIONS [IN DIFFERENT SOURCES] DATA QUALITY RULES ARE THEN DEFINED FOR EACH DIMENSION METRICS ARE THEN SELECTED FOR THEM
  • 33. SIMPLER USER-ORIENTED APPROACH BASED ON USER DEFINED DATA QUALITY REQUIREMENTS
  • 34. ✓ STANDARDIZATION, NORMALIZATION AND PARSING ✓ MATCHING / DEDUPLICATION AND MERGING ✓ DATA CLEANSING ✓ VALIDATION ✓ DATA PROFILING / AUDITING ✓ SOME A FEW OF THEM SUPPORT (SEMI-)AUTOMATED DQ RULE RECOGNITION BASED ON METADATA, BUILT-IN RULES, OR MACHINE LEARNING DQ TOOLS FOR (SEMI-)AUTOMATED DQM
  • 35.
  • 36. SO FAR… DEFINITION USER TIME DIMENSION PROCESS PURPOSE
  • 37. SO FAR… DEFINITION USER TIME DIMENSION PROCESS PURPOSE WHAT ELSE?
  • 38. DATA OBJECT DATASET DATABASE DATA REPOSITORY INFORMATION SYSTEM SOFTWARE NO ONE-SIZE-FITS-ALL
  • 39. DATA OBJECT DATASET DATABASE DATA REPOSITORY INFORMATION SYSTEM SOFTWARE DATA OWNER KNOWN THIRD-PARTY NO ONE-SIZE-FITS-ALL
  • 40. DATA OBJECT DATASET DATABASE DATA REPOSITORY INFORMATION SYSTEM SOFTWARE DATA STRUCTURE NO ONE-SIZE-FITS-ALL STRUCTURED DATA UNSTRUCTURED DATA SEMI-STRUCTURED DATA Image sources: https://monkeylearn.com/blog/semi-structured-data/, https://www.pngitem.com/middle/ioJTTbR_organization-structure-icon-png-download-structures-icon-png/
  • 41. DATA OBJECT DATASET DATABASE DATA REPOSITORY INFORMATION SYSTEM SOFTWARE DATA WAREHOUSE DATA LAKE Maybe even something else? NO ONE-SIZE-FITS-ALL
  • 42. DATA OBJECT DATASET DATABASE DATA REPOSITORY INFORMATION SYSTEM SOFTWARE Running Analytics on the Data Lake - The Databricks Blog NO ONE-SIZE-FITS-ALL
  • 43. Image source: https://www.grazitti.com/blog/data-lake-vs-data-warehouse-which-one-should-you-go-for/, https://www.qubole.com/data-lakes-vs-data-warehouses-the-co-existence-argument/ SCHEMA ON READ SCHEMA ON WRITE “SINGLE SOURCE OF TRUTH”
  • 44. Implementing a Data Lake or Data Warehouse Architecture for Business Intelligence? | by Lan Chu | Towards Data Science NB: EXTRACT-TRANSFORM-LOAD IS NOT DQM!!!
  • 46.
  • 47. Image source: The abstracted future of data engineering | by Justin Gage | Datalogue | Medium OR HOW TO AVOID GIGO*? *“GARBAGE IN, GARBAGE OUT”
  • 48. DATA LAKE FOR BI BUSINESS DATA LAKE https://www.capgemini.com/wp-content/uploads/2017/07/pivotal_data_lake_vs_traditional_bi_20140805.pdf
  • 49. DATA LAKE + DATA WRANGLING [an asset, not a silver bullet] ✔ Source: https://monkeylearn.com/blog/data-wrangling/, https://www.altair.com/what-is-data-wrangling/ , https://pediaa.com/what-is-the-difference-between-data-wrangling-and-data-cleaning
  • 51. THE DATA WRANGLING PROCESS TO PREPARE DATA AND INTEGRATE IT INTO IS DEPENDING ON THE IS AND THE DESIRED OR REQUIRED TARGET QUALITY*, INDIVIDUAL STEPS SHOULD BE CARRIED OUT SEVERAL TIMES ➔ !!! DATA WRANGLING IS A CONTINUOUS PROCESS !!! THAT REPEATS ITSELF REPEATEDLY AT REGULAR INTERVALS. Information System Azeroual, O., Schöpfel, J., Ivanovic, D., & Nikiforova, A. (2022). Combining data lake and data wrangling for ensuring data quality in CRIS. Procedia Computer Science, 211, 3-16.
  • 52. DATA LAKE VS DATA WAREHOUSE HOW TO TAKE THE ADVANTAGES OF BOTH?
  • 53. DATA LAKE VS DATA WAREHOUSE HOW TO TAKE THE ADVANTAGES OF BOTH? DATA LAKEHOUSE
  • 54. DATA LAKEHOUSE IS SEEN AS A COMBINATION OF DATA WAREHOUSING WORKLOADS & DATA LAKE ECONOMICS Running Analytics on the Data Lake - The Databricks Blog
  • 55. Running Analytics on the Data Lake - The Databricks Blog, Build a Lake House Architecture on AWS | AWS Big Data Blog (amazon.com), The Data Lakehouse, the Data Warehouse and a Modern Data platform architecture - Microsoft Community Hub
  • 56. DATA OBJECT DATASET DATABASE DATA REPOSITORY INFORMATION SYSTEM SOFTWARE Running Analytics on the Data Lake - The Databricks Blog
  • 57. DATA QUALITY-AWARE SOFTWARE DEVELOPMENT & DATA QUALITY MODEL-BASED TESTING
  • 58. THINK DATA QUALITY FIRST!!! OR TOWARDS DATA QUALITY BY DESIGN Guerra-García, C., Nikiforova, A., Jiménez, S., Perez-Gonzalez, H. G., Ramírez-Torres, M., & Ontañon- García, L. (2023). ISO/IEC 25012-based methodology for managing data quality requirements in the development of information systems: Towards Data Quality by Design. Data & Knowledge Engineering, 145, DAQUAVORD - A METHODOLOGY FOR PROJECT MANAGEMENT OF DATA QUALITY REQUIREMENTS SPECIFICATION - AIMED AT ELICITING DQ REQUIREMENTS ARISING FROM DIFFERENT USERS’ VIEWPOINTS THESE DQ REQUIREMENTS SERVE AS DATA QUALITY SOFTWARE REQUIREMENT AT THE TIME OF THE DEVELOPMENT OF SOFTWARE THAT TAKES DATA QUALITY INTO ACCOUNT BY DEFAULT. IS BASED ON THE VIEWPOINT-ORIENTED REQUIREMENTS DEFINITION (VORD) METHOD, AND THE LATEST AND MOST GENERALLY ACCEPTED ISO/IEC 25012 STANDARD.
  • 59. DATA ARTIFACT WHAT DQM APPROACH DEPENDS ON? DEFINITION USER TIME DIMENSION PROCESS PURPOSE
  • 60.
  • 61. MUSK’S TOP PRIORITY: TO IMPROVE THE PRODUCT… Q: HOW DOES ONE ENSURE THE RELIABILITY OF DATA AND DECISIONS MADE BASED ON SAID DATA? THE ANSWER LIES NOT IN MANAGING THE DATA ALONE, BUT ALSO THE INFORMATION AROUND AND ABOUT DATA ACQUISITION, TRANSFORMATIONS AND VISUALIZATION TO PROVIDE A BETTER UNDERSTANDING AND SUPPORT DECISION MAKERS https://www.gqindia.com/get-smart/content/5-things-elon-musk-did-to-become-one-of-the-richest-men-in-the-world
  • 62. https://www.gqindia.com/get-smart/content/5-things-elon-musk-did-to-become-one-of-the-richest-men-in-the-world MUSK’S TOP PRIORITY: TO IMPROVE THE PRODUCT… Q: HOW DOES ONE ENSURE THE RELIABILITY OF DATA AND DECISIONS MADE BASED ON SAID DATA? THE ANSWER LIES NOT IN MANAGING THE DATA ALONE, BUT ALSO THE INFORMATION AROUND AND ABOUT DATA ACQUISITION, TRANSFORMATIONS AND VISUALIZATION TO PROVIDE A BETTER UNDERSTANDING AND SUPPORT DECISION MAKERS BY FOCUSING ON SUSTAINABLE DATA, CLEAR DATA GOVERNANCE AND STRONG DATA MANAGEMENT
  • 64. https://www.gqindia.com/get-smart/content/5-things-elon-musk-did-to-become-one-of-the-richest-men-in-the-world DATA GOVERNANCE IS THE ANSWER https://www.edq.com/blog/data-quality-vs-data-governance/ Azeroual O., Nikiforova A., Sha K. (2023) Overlooked Aspects of Data Governance: Workflow Framework For Enterprise Data Deduplication
  • 66.
  • 68. https://www.gqindia.com/get-smart/content/5-things-elon-musk-did-to-become-one-of-the-richest-men-in-the-world THINK DATA QUALITY FIRST! “1-10-100” RULE 1$ SPENT ON PREVENTION SAVES 10$ ON APPRAISAL AND 100$ ON FAILURE COSTS https://twitter.com/bright_data/status/1346443370718240768
  • 69. https://www.gqindia.com/get-smart/content/5-things-elon-musk-did-to-become-one-of-the-richest-men-in-the-world DEVELOP DATA QUALITY MANAGEMENT AND GOVERNANCE STRATEGIES MANTAIN DQM & DQG STRATEGIES DEFINE MEASURE ANALYSE IMPROVE
  • 70.
  • 73. For more information, see ResearchGate, anastasijanikiforova.com For questions or any queries, contact me via Nikiforova.Anastasija@gmail.com,