Fast queries, scalable pipelines and data you can trust.
Data systems fail quietly: a query fast at a thousand rows crawls at a million, a pipeline that worked in a demo drops records under load. Keeping data systems fast and correct as volume grows is its own discipline — one I practise across production pipelines and analytics workloads.
Articles in this hub
9 articles
IntermediatePostgres vs MySQL in 2026: Performance, Syntax and the Real Differences
Where Postgres and MySQL genuinely differ — storage engines, MVCC, connection models, SQL syntax and operations — plus where SQLite and MariaDB fit, and how to actually choose.
Read article
BeginnerIceberg vs Delta Lake in 2026: A Plain Guide to Picking One
What a table format actually is, what Iceberg and Delta Lake really do differently in 2026, and why the choice is decided by your catalog and your query engine rather than by the format itself.
Read article
IntermediateDo You Need a Graph Database? What BigQuery Graph Going GA Actually Changes
BigQuery Graph reached GA on 1 September 2026. Before you model anything as a graph, here is the honest test for whether you need one — plus the line in the pricing page that decides whether you can even run GQL.
Read article
IntermediateDo You Need a Vector Database? Postgres vs Dedicated in 2026
Most teams adding semantic search do not need a separate vector database. A practical 2026 guide to what pgvector actually does, the four mechanisms that make it run out, and what you are really buying when you pay for a dedicated engine.
Read article
BeginnerSQL vs NoSQL in 2026: Database Types, ACID vs BASE, and How to Actually Choose
"SQL vs NoSQL" is the wrong first question. A practical 2026 guide to database types, what relational really means, ACID vs BASE, and a decision framework to choose by your access patterns — not by hype.
Read article
IntermediateData Governance on GCP in 2026: Who Can See This Data?
A practical 2026 guide to data governance on Google Cloud: the four questions every program must answer, and the exact services that answer them — Knowledge Catalog (Dataplex), Sensitive Data Protection, BigQuery policy tags, and VPC Service Controls — with a comparison table and a decision guide.
Read article
IntermediateCloud Bigtable in 2026: When Wide-Column Beats BigQuery, Spanner, and Cassandra
A practical 2026 guide to Cloud Bigtable: what wide-column storage actually is, how Bigtable works under the hood, when it beats BigQuery, Spanner, Cassandra and DynamoDB, and the one thing — row-key design — that makes or breaks it.
Read article
IntermediateBigQuery vs Snowflake in 2026: An Honest Comparison From a GCP Engineer
A practical, no-marketing comparison of BigQuery and Snowflake in 2026 — how their architecture, billing, performance and AI stories really differ, and how to choose the right one for your workload.
Read article
IntermediateSQL Query Optimization in 2026: 7 Simple Techniques for Faster Database Performance
Seven practical SQL query optimization techniques for faster database performance, with PostgreSQL-focused examples for joins, IN lists, EXISTS, date ranges, aggregates, and deduplication.
Read article
FAQ
What is your data engineering background?
Are you available to hire?
How do we start working together?
Is your data layer keeping up?
From slow queries to fragile pipelines and large-scale collection, I help teams build data systems that stay fast and correct at scale.
See data engineering services →