Skip to content
View RahulCodes98's full-sized avatar

Block or report RahulCodes98

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
RahulCodes98/README.md

Hi, I'm Rahul Shankar 👋

Data Engineer  ·  AWS  ·  Spark  ·  Kafka  ·  Modern Data Platforms

Building scalable batch, streaming and cloud-native data pipelines.

   


👨‍💻 About Me

I'm a Data Engineer focused on building reliable, scalable and observable data platforms using Python, SQL, AWS and modern data engineering technologies.

My professional background in production and data operations gave me hands-on experience troubleshooting SQL, APIs, batch jobs, data flows, Linux environments and production systems. I'm now applying that foundation to engineering end-to-end data pipelines and analytics platforms.

My current focus includes:

  • 🔄 Building ETL / ELT pipelines for batch and near-real-time workloads
  • ⚡ Designing event-driven and CDC architectures using Kafka and Debezium
  • 🔥 Processing large-scale datasets with Apache Spark & PySpark
  • ☁️ Building cloud-native data platforms with AWS S3, Glue, Redshift, EMR & Lambda
  • 🏗️ Implementing Bronze / Silver / Gold lakehouse architectures
  • 🧪 Building data-quality, reconciliation and observability frameworks
  • 📊 Designing analytics-ready data warehouses and dimensional models
  • 🚀 Automating pipelines with Airflow, dbt, Docker, Terraform and CI/CD

I enjoy understanding how data moves through systems — from ingestion and transformation to storage, validation and analytics.


🛠️ Tech Stack

Languages

Python SQL Bash

Cloud & Data Platforms

AWS Amazon S3 AWS Glue Amazon Redshift Snowflake

Big Data & Streaming

Apache Spark Kafka Databricks

PySpark   Debezium   CDC   Delta Lake   Parquet

Orchestration & Transformation

Apache Airflow dbt

Databases

PostgreSQL MySQL SQL Server

DevOps & Engineering

Docker Terraform Git GitHub Actions


🚀 Featured Data Engineering Projects

🚚 Real-Time Logistics Event Lakehouse

Built a near-real-time logistics platform for order-to-delivery SLA and route analytics.

Architecture

PostgreSQL → Debezium → Kafka → PySpark → Delta Lake → AWS S3 / Redshift

Engineering Highlights

  • Implemented CDC-based event ingestion from PostgreSQL
  • Streamed events through Apache Kafka
  • Built PySpark transformations for logistics and delivery analytics
  • Designed Bronze, Silver and Gold processing layers
  • Added automated data-quality validation
  • Orchestrated downstream workflows with Apache Airflow
  • Served curated analytical datasets through AWS

💳 Payment Settlement & Revenue Reconciliation Platform

Built a batch-processing platform for reconciling payments, refunds and settlement records.

Tech

Python · Spark · Airflow · dbt · Cloud Object Storage

Engineering Highlights

  • Built automated settlement and revenue reconciliation workflows
  • Implemented source-to-target reconciliation
  • Added debit-credit balancing controls
  • Designed SCD Type 2 models for historical tracking
  • Built automated dbt tests for data validation
  • Integrated pipelines with CI/CD workflows
  • Improved auditability of financial data transformations

🛒 E-Commerce Orders & Inventory Analytics Platform

Built an event-driven AWS analytics pipeline for orders, inventory and operational reporting.

Tech

AWS S3 · Glue · PySpark · Redshift · CloudWatch

Engineering Highlights

  • Built automated ingestion and transformation pipelines using AWS Glue
  • Processed and modelled large transactional datasets
  • Optimised Redshift queries across 1M+ record fact tables
  • Designed distribution and sort keys
  • Added automated data-quality validation
  • Implemented CloudWatch monitoring and alerts for pipeline failures
  • Improved end-to-end pipeline observability

🧠 Data Engineering Concepts

ETL / ELT

Batch Processing

Streaming

Change Data Capture

Data Warehousing

Data Lakehouse

Dimensional Modelling

SCD Type 1 / Type 2

Data Quality

Data Reconciliation

Data Observability

Partitioning

Distributed Processing

CI/CD for Data Pipelines


📚 Currently Learning

  • ☁️ Advanced AWS Data Engineering
  • 🔥 Distributed processing with Apache Spark
  • 🌊 Real-time architectures with Kafka
  • 🏗️ Production-grade Data Lakehouse architecture
  • ⚙️ Data pipeline orchestration and observability
  • 🧱 Infrastructure as Code with Terraform

🏆 Certifications

Certification Status
AWS Certified Data Engineer – Associate 🔄 In Progress
Microsoft DP-700: Fabric Data Engineer Associate 🔄 In Progress
HackerRank SQL Advanced ✅ Completed
Databricks Fundamentals ✅ Completed

🎓 Education

MSc Computing — Data Analytics
Dublin City University, Ireland

B.Tech Information Technology
SASTRA University, India


📊 GitHub Stats


🤝 Let's Connect

I'm interested in opportunities involving:

Data Engineering · Cloud Data Platforms · ETL/ELT · Streaming · Data Infrastructure


Building reliable pipelines, one dataset at a time.

Popular repositories Loading

  1. Real-time-logistics-lakehouse Real-time-logistics-lakehouse Public

    Real-time logistics event lakehouse for order-to-delivery tracking, built with PostgreSQL CDC, Debezium, Kafka, PySpark Structured Streaming, Delta Lake, Airflow, dbt, and AWS

    Python 1

  2. Payment-Settlement-Revenue-Reconciliation-Platform Payment-Settlement-Revenue-Reconciliation-Platform Public

    Python 1

  3. AWS-Data-Governance-Platform- AWS-Data-Governance-Platform- Public

    AWS data lake platform for retail data ingestion, transformation, governance, and Athena analytics.

    Python 1

  4. Aspect-based-Sentiment-Analysis Aspect-based-Sentiment-Analysis Public

    This is my university research project where I used BERT for sentiment analysis and GPT for aspect classification

    Jupyter Notebook

  5. ZomatoDataAnalysisProject ZomatoDataAnalysisProject Public

    Jupyter Notebook

  6. TextSummarisationNLP TextSummarisationNLP Public

    This is a Text Summarisation Project using Natural Language Processing

    Jupyter Notebook