⟨d⟩ dealwithdata

Patterns

Never recompute everything: incremental metadata feeds

A feed that reprocesses millions of datasets every run stops scaling. Recompute only what changed, plus what changed rules touch, and reconcile on a schedule.

draft v1 Updated 2026-10-07 businessenterprise

The situation: a feed assigns every harvested dataset to an owning application. At a few thousand datasets it runs in minutes. At millions of datasets and tens of millions of elements, it runs for hours, because it recomputes and reloads everything, every run.

Find where the time goes

Split a run into three phases and time each:

  1. Extract: read datasets and elements out of the catalog.
  2. Compute: match each dataset to an application (paths, hosts, naming rules).
  3. Load: write assignments back.

Load is usually the worst. Catalog imports tend to be row by row with versioning and audit, so rewriting a million unchanged rows costs nearly as much as writing changed ones.

Fixes, biggest first

  1. Go incremental. Drive the feed from a harvest diff: only added or changed datasets are reassigned. If 1–2% change per cycle, that's a 50–100× reduction.
  2. Assign datasets; let elements inherit. Touching every element individually multiplies the work by the elements-per-dataset ratio.
  3. Skip no-op writes. Compare computed to current, and load only the differences. This alone often cuts load time by 90%.
  4. Compute in sets. Put rules in a table and resolve them with one join or prefix match, not a loop of lookups. An embedded engine such as DuckDB handles millions of rows in minutes.
  5. Partition. Split by application or path prefix so one failure doesn't rerun everything.

Is a full recompute ever needed?

Not every run. Two kinds of change drive work:

  • A dataset changed → recompute that dataset.
  • A rule or reference list changed (a mapping edited, an application split or retired, a host moved) → recompute only the datasets the old or new rule matches.

A full pass is justified only for:

  • Bootstrap, and after major platform or model changes.
  • Scheduled reconciliation, weekly or monthly, as a backstop for missed diffs, failed runs and direct edits. Compute everything, still write only differences, and report drift.
  • Attestation, when an auditor needs proof the whole population was evaluated at a point in time.

If drift stays near zero for several cycles, stretch the interval. That's your evidence that incremental is safe.

The target shape

harvest → diff → assign changed (set-based) → load differences only
                                            → emit AssignmentChanged events

The one thing to get right: rule changes must emit events too. If someone edits a mapping table and nothing notices, incremental runs go quietly stale. Catching that is what reconciliation is for.

Three scales

  • Business: the same logic applies to any nightly sync. Diff first, then write only what changed, and your 2-hour job becomes 5 minutes.
  • Enterprise: at millions of objects, "full refresh" is a design bug, not a strategy.