Pipelines that don’t wake anyone at 3am.
I’m Sohaib, a data engineer who builds ETL, streaming and governance systems for healthcare, fintech and classifieds platforms. Most of my work is the unglamorous kind: pipelines that hold up when the schema changes at 2am and nobody notices, because nothing broke.
Lahore, PK · remote-friendly · usually replies within a day · Databricks Certified
01 / About
Sohaib Tanveer · Lahore
Data Engineer with a backend engineer’s instincts. I care about correctness, cost and the pager as much as the diagram.
Databricks Certified Data Engineer Associate, BS in Data Science. I’ve built medallion-architecture pipelines on Delta Lake, tag-driven masking in Snowflake, CDC that survives schema drift, and Django REST services before that. These days I am at Dubizzle Labs on the Classifieds, Motors and Properties data platform.
02 / Skills
A technical inventory, not a wish list.
Data Engineering & Streaming
where I live
Daily driver
Also fluent
Software Engineering
the craft underneath
Daily driver
Also fluent
Cloud, AWS
primary platform
Daily driver
Also fluent
Agentic Engineering
LLMs doing real work
Daily driver
Also fluent
Cloud, Azure
secondary platform
Daily driver
Also fluent
Data Science Collaboration
picking up their tasks
Daily driver
Also fluent
03 / Projects
Pipelines in production, and the hard parts.
04 / Notes
Field notes from the pipeline.
Your CDC pipeline will break. Design for the binlog, not the schema
What WAL parsing taught me about schema evolution, and why drift should be an event you handle rather than an outage you page for.
Tags are the only governance that survives a sprint
Binding Snowflake masking policies to object tags, then generating the tags from dbt macros so nobody has to remember.
Synthea defaults lie: tuning synthetic patients toward reality
Module probability weights, disease progression, and the unbillable codes hiding in your claims data.
05 / Contact
Tell me what’s breaking.
A slow warehouse, a migration nobody wants to own, a pipeline held together with cron. Send the messy version. I read every message myself and reply within a day.
Also open to full-time data engineering and backend roles.