Refresh drift is a contract break, not a flaky pipeline
Large facts can "succeed" on every run and still lie. Row counts look fine. Pipelines are green. Measures drift because deletes, late arrivals, and corrections were never part of the incremental contract. Analysts discover it weekly. Engineers call it flaky. It is usually a missing agreement.
Treat refresh drift as a contract break. Align watermark semantics with business expectations for corrections. Fast materialization without drift tests just fails faster. Publish a refresh SLA per fact grain. Make drift checks part of the pipeline, not a weekly analyst surprise. Sales-line style facts are the classic patient; the disease is mismatched notebooks, watermarks, and semantic consumers.

Figure 1. Green refreshes are not the same as honest facts. Contracts must cover late arrivals and corrections. Source: Microsoft Learn: data refresh in Power BI.
I learned this on facts where daily loads looked healthy while finance measures wandered. The pipeline was not random. The watermark told a narrower story than the business believed. Once we wrote the contract and tested drift in-pipeline, "flaky" became a closed ticket class.
The core idea
Incremental success means the contract held, not that the job exited zero.
Row-count success can hide deletes that never landed, late arrivals outside the watermark, and corrections applied upstream but ignored downstream. Semantic consumers assume a grain and a freshness story. If notebooks disagree, drift is the symptom. Naming it flaky delays the fix.
A model that stays explainable
1. Write the incremental contract in one page
State grain, watermark column meaning, how deletes appear, how late arrivals are handled, and how corrections reopen history. If two engineers disagree after reading it, the page is not done. Contracts beat tribal notebook comments.
2. Align watermarks with business correction rules
If business expects same-day corrections to flow, a strict append-only watermark will drift measures. If business expects immutable daily snapshots, soft corrections need a different path. Pick deliberately and test both happy and correction paths.
3. Put drift checks in the pipeline
Compare key measures or hash aggregates against a prior trusted mark or a reconciliation source within tolerance. Fail the run or quarantine the slice when drift exceeds policy. Weekly analyst surprise is not a control.
4. Publish a refresh SLA per fact grain
Say how fresh the fact should be, for whom, and what "good" means beyond green checks. SLAs without drift tests are uptime theatre. SLAs with owners create escalation paths.
5. Separate speed from honesty
Materialize quickly after the contract is testable. Faster wrong facts only spread wrong decisions sooner. Performance work follows correctness gates, not the reverse.
6. Trace notebook, watermark, and semantic assumptions
Inventory where each assumption lives. If the semantic model expects SCD behavior the notebook does not provide, document the mismatch as a defect. Drift often sits in the seams between teams.
Failure modes I design against
Green equals true. Exit codes replace reconciliation.
Undefined deletes. Rows that should vanish stay forever.
Late arrivals ignored. Business thinks they are in; watermark says no.
Analyst-as-monitor. Drift found in meetings, not pipelines.
SLA as marketing. Freshness claims without tests.
Speed-first rewrites. Faster materialization, same contract hole.
What a drift test can look like without boiling the ocean
You do not need perfect financial close automation on day one. Start with stable keys: count of distinct document lines in a closed window, sum of amounts for a locked period, or null rates on critical foreign keys. Compare to the previous successful mark and to a source extract for a sample day. Tolerances should be explicit. Zero tolerance on volatile open periods is how you create false failures; infinite tolerance is how you create false calm.
Communicating drift to stakeholders
Say: the pipeline succeeded mechanically; the incremental contract failed for this class of correction; here is the business rule we will encode; here is the test that will block release if it regresses. Avoid "Fabric was flaky." Avoid blaming analysts for noticing. Precision builds trust faster than reassurance.
Rebuild versus repair
Some drift requires a bounded rebuild of a date range. Some requires a forward fix only. Encode which situations demand rebuild in the runbook. Leaving that choice to 2 a.m. judgment guarantees inconsistent history. Pair rebuild playbooks with capacity awareness so honesty does not become an accidental denial-of-service on the workspace.
Trade-offs
Contracts and drift tests add design time and CI minutes. Undefined incremental behavior adds silent financial arguments. Tighter watermarks are simpler and righter for some domains; correction-friendly watermarks are harder and necessary for others. Quarantine paths complicate consumer wiring; silent publish of drifted slices complicates trust. SLAs create accountability and surface underfunding early.
What I would put on an ADR
- Large facts publish an incremental contract covering deletes, late arrivals, and corrections.
- Watermark semantics match documented business expectations.
- Drift checks run in-pipeline with explicit tolerances and failure policy.
- Each fact grain has a refresh SLA with an owner.
- Performance work on materialization follows honesty gates.
- Rebuild versus forward-fix rules live in a versioned runbook.
Primary references: Fabric pipeline and semantic model refresh guidance on Microsoft Learn, plus your fact-grain ownership docs. Pair SLA text with the drift test names so operations and analytics share vocabulary.
Closing
Drift is rarely a ghost in the machine. It is a contract nobody wrote.
Define incremental meaning. Test it in the pipeline. Publish the SLA. Align watermarks with corrections. Prefer honest speed. Stop calling contract breaks flaky.
