PySpark Starter Kit
Build a complete local PySpark pipeline that reads CSV data with explicit schemas, removes incomplete and duplicate sales, joins customer details, calculates totals, and writes Parquet.
The ZIP includes:
• PySpark for Beginners: Mastering the Basics
• PySpark for Beginners: Beyond the Basics
• A tested CSV-to-Parquet pipeline
• Sample sales and customer data
• A step-by-step Jupyter notebook
• Automated tests
• Five exercises with working solutions
• A schema and joins cheat sheet
Tested with PySpark 4.2.0. Windows users can run the supplied Docker command or use WSL. You need basic Python, Python 3.10 or newer, and Java 17 or newer.
One payment. No subscription.