I build data pipelines, quality frameworks, and the platforms that sit underneath analytics. Currently working in the energy sector on portfolio data, engineering and reporting products, while shipping data products personally.
- Currently learning: the Modern Data Stack (dbt, PySpark, Airbyte, data fabric) and deepening Azure (already worked with AWS and GCP)
- Currently building: better-architected data systems, and writing about it as I go
- Ask me about: data quality, pipelines, Data Mesh, or breaking into data careers
- Open to collaborating on data projects and open source contributions
- I share data engineering and tech career content on YouTube: codewithIB
- Off the keyboard: go-kart driver and Formula 1 fan
| Project | What it shows |
|---|---|
| Bank Transaction Pipeline | Personal finance ELT: bank API to scheduled ingestion to SQLite/Postgres, with a Streamlit dashboard |
| Data Quality Checks, no frameworks | Production-grade data quality checks in plain Python and DuckDB. No Great Expectations needed |
| Kafka AWS Streaming Pipeline | Real-time stock market events: Kafka on EC2 to S3, Glue, and Athena, queryable in seconds |
| UK Land Registry Pipeline | Apache Beam and FastAPI, scaling to GCP with BigQuery, Pub/Sub, Terraform, K8s and CircleCI |
| Fashion Trend Prediction | Deep learning with EfficientNet transfer learning and K-means clustering (TensorFlow/Keras) |
Exploring next: dbt, PySpark, Airbyte, DuckDB


