Practical data engineering knowledge
In-depth guides on modeling, ETL, Azure, Databricks and data architecture — written from real-world projects.
SSIS 2025 with no password in the package: Entra ID, TLS 1.3 and Fabric with Microsoft SqlClient
How to use the new Microsoft SqlClient Data Provider in the SSIS 2025 ADO.NET connection manager to take passwords out of packages, enable TLS 1.3 and connect to Fabric Warehouse with Entra ID.
Read articleAuto Partitioning in the Fabric Data Factory Copy job: move giant tables in minutes, with no partition setup
The Fabric Data Factory Copy job now partitions large tables automatically — it picks the column, computes the boundaries and runs parallel reads from a single toggle. What changes, how to enable it and where it works.
Read articleSSIS + PostgreSQL via ODBC: the fetch options that avoid Out of memory — and the -- comment trap
How UseDeclareFetch/Fetch (a server-side cursor) avoid Out of memory when extracting PostgreSQL in SSIS via ODBC, how to reveal the generic error with CommLog, the -- comment bug that swallows the FETCH, and the Python-cursor fallback.
Read articleDelta Lake in pure Python: write Delta tables without Spark using the deltalake package (delta-rs)
How the deltalake package (delta-rs, a Rust/Arrow core) writes legitimate Delta tables without Spark or a JVM — writing, MERGE, OPTIMIZE and time travel in Python, and where Spark still wins.
Read articleTime-series forecasting in a single SQL call: Databricks' ai_forecast()
How Databricks' ai_forecast() delivers time-series forecasting in a single SQL query — the syntax, per-group forecasting, what the function returns, and when (not) to use it.
Read articleCLUSTER BY AUTO in Databricks: stop picking clustering keys by hand
How CLUSTER BY AUTO (Automatic Liquid Clustering) makes Databricks choose and maintain the clustering keys on its own — what changes versus partitioning/ZORDER, how to enable it, requirements, and when (not) to use it.
Read articleNarwhals: write DataFrame code once and run it on pandas, Polars and PyArrow
What Narwhals is — the lightweight compatibility layer that lets you write DataFrame logic once (Polars-style API) and run it on pandas, Polars, PyArrow and more, with no lock-in and no required dependencies.
Read articleai_parse_document(): turn a PDF into a governed table with a single SQL statement
How Databricks' ai_parse_document() collapses OCR, parsing and table reconstruction into a single SQL statement — landing the result as a governed table in Unity Catalog.
Read articleExpectations in Lakeflow: data quality as code, governed in Unity Catalog
How to declare quality rules next to the transformation, choose between logging, dropping or failing, and govern it all through Unity Catalog with versioned, auditable rules.
Read articleIncremental loads in Azure Data Factory: the watermark pattern step by step
How to do incremental loads in Azure Data Factory using the watermark pattern: Lookup the last value, Copy Data only for the new window, and a Stored Procedure that updates the control table. A practical guide.
Read articleSSIS Data Flow up to 3× faster: tuning the buffer and Fast Load
How to use AutoAdjustBufferSize, DefaultBufferMaxRows and Fast Load to speed up large loads in the SSIS Data Flow. A practical guide with real numbers.
Read articleCheckpoints in SSIS: resume a package from the exact point of failure
How to use SSIS Checkpoints to resume a long package from the exact point of failure — the 3 configuration properties, the pitfalls with Data Flow and loops, and when (or when not) to use them in 2026.
Read articlePolars streaming: process data larger than RAM — without Spark
How the Polars streaming engine processes datasets that don't fit in memory using scan_parquet + sink_parquet, keeping RAM usage constant and doing away with a cluster.
Read articleMetric Views in Unity Catalog: define the KPI once, use it everywhere
How Databricks Unity Catalog Metric Views turn business KPIs into governed, reusable objects — with the YAML walkthrough, the MEASURE() function, the query pattern, and when (or when not) to use them in 2026.
Read articleCDC in SSIS: incremental loads without scanning the whole table
How to use Change Data Capture with the CDC Control Task to turn full loads into incremental loads in SSIS — with code, the LSN state pattern, and when (or when not) to use it in 2026.
Read articleuv: the Python manager every Data Engineer should know
How uv, from Astral, replaced pip, venv and pyenv in our Databricks and Azure projects — with a 10–100× speed boost.
Read articleIncremental ingestion: stop reloading everything every night
Watermarking, change data capture and the patterns that cut cost and processing windows in ETL pipelines.
Read articleSlowly Changing Dimensions Type 2, without the headache
The essential pattern for tracking history in dimensions — explained with a concrete example and the most common mistakes.
Read articleMedallion Architecture: the pattern that organizes your Lakehouse
How the Bronze, Silver and Gold layers turn a chaotic data lake into a reliable, auditable platform.
Read article