Cloud Data Engineer building reliable batch, streaming, and lakehouse platforms that turn source data into analytics-ready models.
I design and build end-to-end cloud data platforms for reliable batch and real-time analytics using Azure, AWS, Google Cloud, Databricks, Spark, Kafka, Airflow, and dbt—from data ingestion and orchestration to governed lakehouse layers and analytics-ready models.
- Metadata-driven ingestion from databases, files, APIs, and event streams
- Batch and streaming transformations with PySpark and Structured Streaming
- Governed Bronze, Silver, and Gold lakehouse layers using Delta Lake and Unity Catalog
- Incremental processing, CDC, and SCD Type 1/2 history management
- Cross-platform orchestration with Azure Data Factory and Apache Airflow
- Analytics-ready facts, dimensions, and serverless SQL datasets
| Project | Engineering highlights | Stack |
|---|---|---|
| ☁️ Real-Time AWS E-commerce Lakehouse | Streams e-commerce events through MSK Serverless and MSK Connect into S3; incrementally builds Bronze, Silver, and Gold layers with Glue and Delta Lake; publishes an SCD-managed star schema to Redshift Serverless | MSK Serverless, MSK Connect, S3, AWS Glue, PySpark, Delta Lake, Redshift Serverless |
| 🎧 Spotify Azure Incremental Lakehouse | Metadata-driven incremental ingestion from Azure SQL to ADLS Gen2; Auto Loader processing; SCD Type 2 dimensions and a Type 1 streaming fact | ADF, ADLS Gen2, Databricks, PySpark, Delta Lake |
| 🚕 Real-Time Uber Ride Lakehouse | Streams FastAPI-generated ride events through Event Hubs; unifies historical and real-time data; publishes an enriched OBT and SCD-managed dimensional model | FastAPI, Event Hubs, Structured Streaming, Lakeflow, Unity Catalog |
| 🌬️ Airflow + dbt + Databricks | Airflow 3 orchestrates Databricks ingestion and a dependency-aware dbt graph using deferrable tasks and distributed Celery workers backed by Redis and PostgreSQL | Airflow, dbt, Databricks, Celery, Redis, PostgreSQL, Docker |
| 🛰️ NASA GCN Fermi Streaming Pipeline | OAuth-authenticated Kafka ingestion of NASA gamma-ray burst notices; native Spark parsing; governed analytical snowflake schema | Kafka, OAuth 2.0, PySpark, Databricks, Delta Lake |
- E-commerce Databricks Pipeline — Six data domains processed through Bronze, Silver, and Gold layers with data-quality expectations, CDC, SCD Type 2, and dimensional modeling.
- Olist Azure Big Data Platform — Approximately 1.56 million marketplace records ingested from HTTP, SQL, and MongoDB, transformed with PySpark, and served through Synapse Serverless SQL.
- Automated data testing, pipeline observability, and operational reliability
- CI/CD and environment-based deployment for Databricks and cloud data platforms
- Spark and Delta Lake performance tuning
Designing reliable paths from source systems to analytics-ready data.
