⟨d⟩ dealwithdata

Patterns

Serving metadata to people, systems and agents

One canonical model, one governed write path, many read paths. How to turn catalog changes into subscribable domain events, and which channel serves which consumer.

draft v1 Updated 2026-10-07 enterprise

Once a catalog has good business metadata, everyone asks for it: "Can we get technical metadata with business metadata?", "Can our AI agent use the definitions?" The failure mode is ten teams building ten extracts. The fix is a shape, not a tool.

The shape

One canonical model, one governed write path, many read paths.

The hard, valuable part isn't the streaming. It's the governed link: business term → business element → physical column → lineage. Everything below is delivery.

Getting changes out

Most catalogs sit on a relational database and are poll or harvest based. Changes arrive two ways:

  • UI edits (a steward changes a definition): capture with database change data capture (CDC) such as GoldenGate or Debezium.
  • Harvests (scanners reload technical metadata): diff successive snapshots and emit a change set. This gives cleaner events than the database log.

Three layers, not one

Catalog DB ──CDC──▶ raw.* topics          (private: one team reads these)
Harvests ──diff──▶        │
                          ▼
                    Translator            row changes → domain events
                          │               schema registry, versioned
                          ▼
              public topics               BusinessElementDefinitionChanged
                                          LineageEdgeAdded
                                          CdeStewardReassigned

Never expose raw CDC from a vendor's internal schema. It is undocumented, changes on upgrade, and one business edit fires a dozen row changes. Subscribers break and drown.

Pick the channel by consumer

Consumer Channel Standard
Stewards, analysts Catalog UI n/a
Applications REST or GraphQL over the metadata graph OpenAPI
Change subscribers Domain-event topics CloudEvents envelope, AsyncAPI docs
Bulk analytics Metadata snapshots as tables SQL
AI agents MCP server plus a search index MCP
Pipelines Emit lineage at run time; enforce contracts in CI OpenLineage, ODCS
BI and AI tools Semantic model export OSI (Apache Ossie)

All of these are projections of the same model. When a team asks for "a feed", they get one of these, not a custom extract.

Rules that keep it trustworthy

  1. Stable global IDs for every term, element and physical asset, or the links can't survive across systems.
  2. The link is a governed object. A business-to-technical mapping carries an owner, a confidence score, provenance (human, rule or AI) and a status. AI-proposed mappings enter as proposed; stewards authorize.
  3. The catalog stays the system of record for its domain. Publish outward; don't try to make it the hub for everything.
  4. Agents read through the MCP layer, never the database. That gives one place for entitlements, audit, and "approved definitions only".

When it becomes a nervous system

active metadata only matters if something depends on it. Wire in two or three consumers that break when metadata is wrong, such as change management, a regulatory report pipeline, and one AI agent. Until then, it's documentation.

Using LLMs for long descriptions

For bulk drafting of element descriptions, a fast, cheap model is the right default. Quality comes from grounding, not model size: feed the physical name and type, profile statistics, lineage neighbors, the parent dataset and nearby glossary terms. Route critical elements and low-confidence drafts to a stronger model or a human reviewer.

Sources

  1. OpenLineage accessed 2026-10-07
  2. CloudEvents specification accessed 2026-10-07
  3. AsyncAPI accessed 2026-10-07
  4. Model Context Protocol accessed 2026-10-07
  5. Open Data Contract Standard (Bitol) accessed 2026-10-07