All certifications / Data+ / Lessons
Data+ DA0-002 lessons
Study Data+ for free
A week-by-week plan with every lesson, quizzes, checkpoint tests, a practice exam and hands-on labs.
Open the Data+ study planA week-by-week plan with every lesson, quizzes, checkpoint tests, a practice exam and hands-on labs.
Domain 1: Data concepts and environments
- Structured, semi-structured and unstructured data, with examples of each
- Relational databases: tables, primary and foreign keys, normalization and relationships
- Non-relational databases: document, key-value, column-family and graph stores
- Data types: strings, integers, decimals, dates and times, Booleans and type conversion
- File formats: CSV, TSV, JSON, XML, Parquet, spreadsheets and flat files
- Data warehouses, data marts, data lakes and lakehouses; OLTP vs OLAP
- Dimensional modeling: fact and dimension tables, star and snowflake schemas, slowly changing dimensions
- Data environments: on-premises vs cloud, and tools such as spreadsheets, SQL clients, notebooks and BI platforms
- Languages for analysis: SQL, Python and R, and when to use each
- AI and automation concepts for analysts: machine learning, generative AI, large language models, NLP and RPA
Domain 2: Data acquisition and preparation
- Data acquisition methods: database extracts, APIs, web scraping, file exports, surveys and sampling
- ETL vs ELT and data pipelines: batch vs streaming, full vs incremental loads
- Writing SQL queries: SELECT, WHERE, ORDER BY, GROUP BY and HAVING
- Combining data: inner, left, right and full joins, unions and appending
- Data quality problems: duplicates, missing values, invalid values, outliers, inconsistent formats and redundancy
- Handling missing data: deletion, imputation and flagging
- Data transformation: parsing, splitting and concatenating fields, type conversion, recoding and derived variables
- Scaling and grouping: normalization, standardization, binning and aggregation
- Reshaping data: pivoting and unpivoting, wide vs long format, filtering and sorting
- Query optimization: indexing, filtering early, avoiding SELECT *, subsets and temporary tables
Domain 3: Data analysis
- Measures of central tendency: mean, median and mode, and how skew affects them
- Measures of dispersion: range, variance, standard deviation, interquartile range and percentiles
- Distributions: normal distribution, skewness, the empirical rule and z-scores
- Descriptive statistics in practice: counts, frequencies, percentages, percent change and ratios
- Inferential statistics: samples vs populations, confidence intervals, hypothesis testing and p-values
- Type I and Type II errors, statistical significance and sample size
- Correlation vs causation, and simple linear regression
- Types of analysis: exploratory, descriptive, diagnostic, predictive, prescriptive and trend analysis
- Performance analysis: KPIs, metrics, targets and variance to plan
- Choosing analysis tools and functions: spreadsheet formulas, SQL aggregates and window functions, Python and R libraries
- Checking and troubleshooting results: sanity checks, reconciliation, and common calculation errors
Domain 4: Visualization and reporting
- Choosing a chart: bar, line, pie, scatter, histogram, box plot, heat map, map and table
- Dashboard design: layout, audience, KPIs, filters, drill-down and interactivity
- Design principles: color, labels, titles, scales and axes, accessibility and avoiding misleading charts
- Report types: static vs dynamic, ad hoc vs recurring, self-service and executive summaries
- Communicating findings: knowing the audience, storytelling with data and stating limitations
- Report elements: cover information, methodology, data sources, refresh dates, disclaimers and appendices
- Delivery and refresh: scheduled refresh, real-time vs snapshot data, subscriptions and distribution
- Report versioning, style guides and corporate branding
- Troubleshooting reports and dashboards: stale data, broken filters, wrong totals and slow performance
Domain 5: Data governance
- Data governance roles: data owner, data steward, data custodian and data consumer
- Metadata, data dictionaries, data catalogs and data lineage
- Data quality dimensions: accuracy, completeness, consistency, validity, timeliness and uniqueness
- Data quality control: validation rules, profiling, quality metrics and monitoring
- Master data management and a single source of truth
- Sensitive data: PII, PHI and payment card data; data classification levels
- Privacy and compliance: GDPR, HIPAA, PCI DSS, data sovereignty and retention policies
- Protecting data: access control and least privilege, masking, anonymization, pseudonymization and encryption
- Data life cycle: collection, storage, use, sharing, archiving and secure disposal