Data Engineer · AWS · Spark · Kafka · Modern Data Platforms
Building scalable batch, streaming and cloud-native data pipelines.
I'm a Data Engineer focused on building reliable, scalable and observable data platforms using Python, SQL, AWS and modern data engineering technologies.
My professional background in production and data operations gave me hands-on experience troubleshooting SQL, APIs, batch jobs, data flows, Linux environments and production systems. I'm now applying that foundation to engineering end-to-end data pipelines and analytics platforms.
My current focus includes:
- 🔄 Building ETL / ELT pipelines for batch and near-real-time workloads
- ⚡ Designing event-driven and CDC architectures using Kafka and Debezium
- 🔥 Processing large-scale datasets with Apache Spark & PySpark
- ☁️ Building cloud-native data platforms with AWS S3, Glue, Redshift, EMR & Lambda
- 🏗️ Implementing Bronze / Silver / Gold lakehouse architectures
- 🧪 Building data-quality, reconciliation and observability frameworks
- 📊 Designing analytics-ready data warehouses and dimensional models
- 🚀 Automating pipelines with Airflow, dbt, Docker, Terraform and CI/CD
I enjoy understanding how data moves through systems — from ingestion and transformation to storage, validation and analytics.
PySpark Debezium CDC Delta Lake Parquet
Built a near-real-time logistics platform for order-to-delivery SLA and route analytics.
Architecture
PostgreSQL → Debezium → Kafka → PySpark → Delta Lake → AWS S3 / Redshift
Engineering Highlights
- Implemented CDC-based event ingestion from PostgreSQL
- Streamed events through Apache Kafka
- Built PySpark transformations for logistics and delivery analytics
- Designed Bronze, Silver and Gold processing layers
- Added automated data-quality validation
- Orchestrated downstream workflows with Apache Airflow
- Served curated analytical datasets through AWS
Built a batch-processing platform for reconciling payments, refunds and settlement records.
Tech
Python · Spark · Airflow · dbt · Cloud Object Storage
Engineering Highlights
- Built automated settlement and revenue reconciliation workflows
- Implemented source-to-target reconciliation
- Added debit-credit balancing controls
- Designed SCD Type 2 models for historical tracking
- Built automated dbt tests for data validation
- Integrated pipelines with CI/CD workflows
- Improved auditability of financial data transformations
Built an event-driven AWS analytics pipeline for orders, inventory and operational reporting.
Tech
AWS S3 · Glue · PySpark · Redshift · CloudWatch
Engineering Highlights
- Built automated ingestion and transformation pipelines using AWS Glue
- Processed and modelled large transactional datasets
- Optimised Redshift queries across 1M+ record fact tables
- Designed distribution and sort keys
- Added automated data-quality validation
- Implemented CloudWatch monitoring and alerts for pipeline failures
- Improved end-to-end pipeline observability
ETL / ELT
Batch Processing
Streaming
Change Data Capture
Data Warehousing
Data Lakehouse
Dimensional Modelling
SCD Type 1 / Type 2
Data Quality
Data Reconciliation
Data Observability
Partitioning
Distributed Processing
CI/CD for Data Pipelines
- ☁️ Advanced AWS Data Engineering
- 🔥 Distributed processing with Apache Spark
- 🌊 Real-time architectures with Kafka
- 🏗️ Production-grade Data Lakehouse architecture
- ⚙️ Data pipeline orchestration and observability
- 🧱 Infrastructure as Code with Terraform
| Certification | Status |
|---|---|
| AWS Certified Data Engineer – Associate | 🔄 In Progress |
| Microsoft DP-700: Fabric Data Engineer Associate | 🔄 In Progress |
| HackerRank SQL Advanced | ✅ Completed |
| Databricks Fundamentals | ✅ Completed |
MSc Computing — Data Analytics
Dublin City University, Ireland
B.Tech Information Technology
SASTRA University, India
I'm interested in opportunities involving:
Data Engineering · Cloud Data Platforms · ETL/ELT · Streaming · Data Infrastructure
Building reliable pipelines, one dataset at a time.