I build reliable, reproducible data and machine-learning projects where results can be traced back to their source.
Based in Christchurch, New Zealand. This portfolio covers data pipelines, analytics, distributed processing and applied machine learning, with an emphasis on data quality and practical delivery.
| Project | What stands out |
|---|---|
| retail-ai-pipeline | End-to-end Airflow + dbt retail pipeline with atomic versioned publishing; 66 Python tests, £10.25M reconciled across 19,773 orders, and 522,566 loaded rows + 19,343 quarantined rows accounted for exactly. |
| aerial-small-object-detection | YOLO11 small-object detection and tracking on VisDrone2019; 130 tests, Dockerized GPU workflow, ONNX export and measured accuracy/latency benchmarks over 38,759 validation boxes. |
| nz-attraction-pageviews | Incremental Wikimedia API ingestion into DuckDB with idempotent recovery and data-quality checks; 187 offline tests across Python 3.10–3.13. |
| online-retail-analysis-r | Reproducible R + SQL analysis of 541,909 invoice lines; cancellation-aware netting produces £9.88M in clean sales, backed by a SHA-256-pinned input and reconciliation audit. |
| Million-Song-Dataset-Analysis-with-Spark | PySpark ML over 48.4M listening events and 12.2 GB of audio features; genre classification plus implicit-feedback ALS recommendations, with sanitized source notebooks and CI. |
| GHCN-Daily-Climate-Analysis-with-PySpark | Distributed climate-data engineering for the 13+ GB GHCN-Daily archive; fixed-width station enrichment, New Zealand temperature trends and global precipitation outputs. |
- Reproducibility: pinned or fingerprinted inputs, one-command rebuilds and committed evidence.
- Auditability: explicit quality gates, row-level reconciliation and immutable/atomic publication.
- Engineering discipline: CI, automated tests, documented limitations and measured—not assumed—performance.
Python · SQL · R · Airflow · dbt · DuckDB · pandas · PySpark · Docker · Power BI · PyTorch · ONNX