---
id: dwd_q3gbjug4
type: pattern
slug: incremental-metadata-feeds
title: "Never recompute everything: incremental metadata feeds"
summary: "A feed that reprocesses millions of datasets every run stops scaling. Recompute only what changed, plus what changed rules touch, and reconcile on a schedule."
status: draft
classification: public
authors:
  - name: Shan Umasankar
    url: https://dealwithdata.com
    role: author
assisted_by: [Claude]
scales: [business, enterprise]
tags: [performance, architecture, harvest]
terms: [dwd_mgglmnwq, dwd_gubdecks, dwd_tthoszn7]
related: [dwd_cb2bgvae]
created: 2026-10-07
updated: 2026-10-07
version: 1
---

**The situation:** a feed assigns every harvested dataset to an owning
application. At a few thousand datasets it runs in minutes. At millions of
datasets and tens of millions of elements, it runs for hours, because it
recomputes and reloads everything, every run.

## Find where the time goes

Split a run into three phases and time each:

1. **Extract:** read datasets and elements out of the catalog.
2. **Compute:** match each dataset to an application (paths, hosts, naming rules).
3. **Load:** write assignments back.

Load is usually the worst. Catalog imports tend to be row by row with
versioning and audit, so rewriting a million unchanged rows costs nearly as
much as writing changed ones.

## Fixes, biggest first

1. **Go incremental.** Drive the feed from a harvest diff: only added or
   changed datasets are reassigned. If 1–2% change per cycle, that's a
   50–100× reduction.
2. **Assign datasets; let elements inherit.** Touching every element
   individually multiplies the work by the elements-per-dataset ratio.
3. **Skip no-op writes.** Compare computed to current, and load only the
   differences. This alone often cuts load time by 90%.
4. **Compute in sets.** Put rules in a table and resolve them with one join
   or prefix match, not a loop of lookups. An embedded engine such as DuckDB
   handles millions of rows in minutes.
5. **Partition.** Split by application or path prefix so one failure
   doesn't rerun everything.

## Is a full recompute ever needed?

Not every run. Two kinds of change drive work:

- **A dataset changed** → recompute that dataset.
- **A rule or reference list changed** (a mapping edited, an application
  split or retired, a host moved) → recompute only the datasets the old
  *or* new rule matches.

A full pass is justified only for:

- **Bootstrap**, and after major platform or model changes.
- **Scheduled reconciliation**, weekly or monthly, as a backstop for missed
  diffs, failed runs and direct edits. Compute everything, still write only
  differences, and report drift.
- **Attestation**, when an auditor needs proof the whole population was
  evaluated at a point in time.

If drift stays near zero for several cycles, stretch the interval. That's
your evidence that incremental is safe.

## The target shape

```
harvest → diff → assign changed (set-based) → load differences only
                                            → emit AssignmentChanged events
```

**The one thing to get right:** rule changes must emit events too. If
someone edits a mapping table and nothing notices, incremental runs go
quietly stale. Catching that is what reconciliation is for.

## Three scales

- **Business:** the same logic applies to any nightly sync. Diff first, then
  write only what changed, and your 2-hour job becomes 5 minutes.
- **Enterprise:** at millions of objects, "full refresh" is a design bug,
  not a strategy.
