{
  "id": "dwd_q3gbjug4",
  "type": "pattern",
  "slug": "incremental-metadata-feeds",
  "title": "Never recompute everything: incremental metadata feeds",
  "summary": "A feed that reprocesses millions of datasets every run stops scaling. Recompute only what changed, plus what changed rules touch, and reconcile on a schedule.",
  "status": "draft",
  "classification": "public",
  "authors": [
    {
      "name": "Shan Umasankar",
      "url": "https://dealwithdata.com",
      "role": "author"
    }
  ],
  "assisted_by": [
    "Claude"
  ],
  "scales": [
    "business",
    "enterprise"
  ],
  "tags": [
    "performance",
    "architecture",
    "harvest"
  ],
  "terms": [
    "dwd_mgglmnwq",
    "dwd_gubdecks",
    "dwd_tthoszn7"
  ],
  "related": [
    "dwd_cb2bgvae"
  ],
  "created": "2026-10-07",
  "updated": "2026-10-07",
  "version": 1,
  "url": "https://dealwithdata.com/c/incremental-metadata-feeds/",
  "markdown_url": "https://dealwithdata.com/c/incremental-metadata-feeds.md",
  "body_markdown": "**The situation:** a feed assigns every harvested dataset to an owning\napplication. At a few thousand datasets it runs in minutes. At millions of\ndatasets and tens of millions of elements, it runs for hours, because it\nrecomputes and reloads everything, every run.\n\n## Find where the time goes\n\nSplit a run into three phases and time each:\n\n1. **Extract:** read datasets and elements out of the catalog.\n2. **Compute:** match each dataset to an application (paths, hosts, naming rules).\n3. **Load:** write assignments back.\n\nLoad is usually the worst. Catalog imports tend to be row by row with\nversioning and audit, so rewriting a million unchanged rows costs nearly as\nmuch as writing changed ones.\n\n## Fixes, biggest first\n\n1. **Go incremental.** Drive the feed from a harvest diff: only added or\n   changed datasets are reassigned. If 1\u20132% change per cycle, that's a\n   50\u2013100\u00d7 reduction.\n2. **Assign datasets; let elements inherit.** Touching every element\n   individually multiplies the work by the elements-per-dataset ratio.\n3. **Skip no-op writes.** Compare computed to current, and load only the\n   differences. This alone often cuts load time by 90%.\n4. **Compute in sets.** Put rules in a table and resolve them with one join\n   or prefix match, not a loop of lookups. An embedded engine such as DuckDB\n   handles millions of rows in minutes.\n5. **Partition.** Split by application or path prefix so one failure\n   doesn't rerun everything.\n\n## Is a full recompute ever needed?\n\nNot every run. Two kinds of change drive work:\n\n- **A dataset changed** \u2192 recompute that dataset.\n- **A rule or reference list changed** (a mapping edited, an application\n  split or retired, a host moved) \u2192 recompute only the datasets the old\n  *or* new rule matches.\n\nA full pass is justified only for:\n\n- **Bootstrap**, and after major platform or model changes.\n- **Scheduled reconciliation**, weekly or monthly, as a backstop for missed\n  diffs, failed runs and direct edits. Compute everything, still write only\n  differences, and report drift.\n- **Attestation**, when an auditor needs proof the whole population was\n  evaluated at a point in time.\n\nIf drift stays near zero for several cycles, stretch the interval. That's\nyour evidence that incremental is safe.\n\n## The target shape\n\n```\nharvest \u2192 diff \u2192 assign changed (set-based) \u2192 load differences only\n                                            \u2192 emit AssignmentChanged events\n```\n\n**The one thing to get right:** rule changes must emit events too. If\nsomeone edits a mapping table and nothing notices, incremental runs go\nquietly stale. Catching that is what reconciliation is for.\n\n## Three scales\n\n- **Business:** the same logic applies to any nightly sync. Diff first, then\n  write only what changed, and your 2-hour job becomes 5 minutes.\n- **Enterprise:** at millions of objects, \"full refresh\" is a design bug,\n  not a strategy.\n"
}