Marsof Academy
Course 1 · Unit 4

Data for AI

Part of Foundations: programming and data

  • 7 lessons
  • ≈ 23 h of study
  • Level: beginner

The tools used to prepare any dataset before training anything. NumPy in depth, pandas, SQL, visualisation and a cleaning and exploratory analysis workflow that avoids information leaks.

Topics covered

  • NumPy
  • Arrays
  • dtypes
  • Broadcasting
  • Vectorisation
  • Views and copies
  • Advanced indexing
  • pandas
  • DataFrame
  • groupby
  • merge
  • Missing values
  • Time series
  • SQL
  • SQLite
  • JOIN
  • Window functions
  • Indexes
  • Visualisation
  • Histograms
  • ECDF
  • Misleading charts
  • Data cleaning
  • EDA
  • Outliers
  • Information leakage

Lessons in this unit

  1. NumPy in depth 60 min
    Arrays, dtypes, broadcasting, views and copies, and advanced indexing; why vectorised code is tens of times faster than a loop.
  2. pandas (I): selecting, filtering and grouping 75 min
    Series, DataFrame and index, a first look at any table, selection with columns, loc, iloc and masks, new columns, sorting, counting and summarising by group with groupby, always thinking in columns rather than rows.
  3. pandas (II): joining tables, missing values and time series 105 min
    merge without multiplying rows, transform for per-group computations, missing values and dtypes, and time series with resample, rolling and shift, checking every step with counts.
  4. SQL with sqlite3 (I): querying, grouping and joining 85 min
    From scratch — tables, SELECT and WHERE, ORDER BY and LIMIT, aggregations with GROUP BY and HAVING, NULL, INNER and LEFT JOIN, and how to run it all from Python with sqlite3 and parameters.
  5. SQL with sqlite3 (II): window functions, CTEs and indexes 105 min
    Window functions (ROW_NUMBER, SUM OVER, LAG) to compute per group without losing rows, CTEs with WITH to write queries step by step, and indexes and EXPLAIN QUERY PLAN to make them fast.
  6. Visualisation — seeing the data before modelling it 50 min
    Choosing the chart for the question, reading distributions with histograms, ECDFs and quantiles, and recognising charts that mislead.
  7. Data cleaning and exploratory analysis 70 min
    A reproducible EDA workflow — types, missing values, duplicates, impossible values and outliers — and the information leaks that make a model shine in the notebook and fail in production.

Prerequisites

Before this unit it helps to have done:

The full explanations, auto-graded exercises, exams and projects are inside the academy.

Every lesson you complete gives you 10 yang, the academy's currency, and every unit exam you pass, 50.

Shall we start?

Create your account and activate your subscription: you get the whole syllabus from day one.