# dealwithdata > Governed data, written as governed data. Practitioner writing on data governance, metadata, lineage and data quality at personal, business and enterprise scale. Every item has a permanent id, a status (draft/reviewed/verified), authors, AI-assistance disclosure and cited sources. Machine channels: https://dealwithdata.com/api/catalog.json (everything), https://dealwithdata.com/api/glossary.skos.jsonld (glossary as SKOS), https://dealwithdata.com/api/lineage.json, https://dealwithdata.com/api/events.jsonl (CloudEvents change log). An MCP server is described at https://dealwithdata.com/consume/. ## Concepts - [Business term vs business element](https://dealwithdata.com/c/business-term-vs-business-element.md): A term is a word and its meaning. An element is governed data that carries that meaning in one context. Governance attaches to the element. [draft, dwd_2p63wxfx] A **business term** is a word and its meaning. A **business element** is a piece of data that carries that meaning in a specific business context. | | Business term | Business element | |---|---|---| | Lives in | [[business-glossary]] | Business data dictionary or domain | | Answers | "What does this word mean?" | "What data do we hold for it, and where?" | | Has | Definition, synonyms, owner | Type, format, allowed values, quality rules, CDE flag | | Links to | Elements (one term → many) | Physical columns (one element → many) | ## The chain ``` Term (meaning) Outstanding Balance └─ Element (data) Mortgage Outstanding Balance ← CDE │ ├─ Column LOAN_DTL.OUTSTANDING_BAL_AMT │ └─ Column MTG_HIST.UPB └─ Element (data) Card Outstanding Balance └─ Column CARD_ACCT.CUR_BAL ``` ## Why the distinction matters - **Governance attaches to the element.** Stewards, [[critical-data-element]] status and quality rules belong to *Mortgage Outstanding Balance*, not to the abstract word. - **Descriptions belong on the element.** "For mortgages, excludes escrow" is true of one element and false of the term in general. When AI drafts long descriptions, it should draft them per element, grounded in that element's columns, profile and lineage. - **Mapping is two hops, not one.** Semantic discovery tools propose column → element → term. Each hop is a proposal with a confidence score until a steward approves it. ## Common mistakes 1. **Making every column a term.** The glossary bloats with `OUTSTANDING_BAL_AMT_2` and nobody trusts it. 2. **Skipping elements.** Terms link straight to columns, so there is nowhere to hang domain-specific rules or criticality. 3. **One element per system.** If three systems hold the same mortgage balance, that is one element with three columns, not three elements. - [The metadata map: every piece, one example](https://dealwithdata.com/c/metadata-map.md): Catalog, glossary, dictionary, lineage, profiling, quality, contracts, semantic layer: what each one is, shown on a single loan-balance column. [draft, dwd_wg5zzqg2] Governance vocabulary is a pile of overlapping words. The fastest way to untangle it is to hold one piece of data still and look at it through each lens in turn. **The example:** a lending table `LOAN_DTL` with a column `OUTSTANDING_BAL_AMT`. ## Every piece, on one column | Piece | What it is | On our column | |---|---|---| | [[technical-metadata]] | Facts about the physical data | `LOAN_DTL.OUTSTANDING_BAL_AMT`, `DECIMAL(18,2)`, Oracle, schema `LND` | | [[data-dictionary]] | One system's column-by-column reference | "Unpaid principal, USD, ≥ 0" | | [[business-glossary]] | The enterprise's words and meanings | *Outstanding Balance*: unpaid principal as of the reporting date | | [[business-element]] | Governed data carrying that meaning in context | *Mortgage Outstanding Balance* | | [[critical-data-element]] | A flag: this one matters most | Yes, because it feeds regulatory reporting | | [[data-catalog]] | Searchable inventory of everything | Search "loan balance" → this column, its owner, its reports | | [[data-lineage]] | Where it comes from and goes | Core banking → batch job → `LOAN_DTL` → exposure report | | [[data-profiling]] | What the data actually looks like | 12M rows, min 0, max 4.2M, 0.3% null | | [[data-quality]] | Rules it must pass | "≥ 0 and not null": 99.7% pass | | [[data-contract]] | Producer's promise to consumers | Stays `DECIMAL`, lands by 06:00, nulls under 1% | | [[reference-data]] | Shared code lists | Status codes `AC` / `CL` / `CO` on the same table | | [[master-data]] | The golden version of an entity | The loan's borrower resolves to Customer #123 | | [[semantic-layer]] | Calculations defined once | `Total Exposure = SUM(OUTSTANDING_BAL_AMT)` | ## How they connect 1. **Harvest** the [[technical-metadata]] and the source [[data-dictionary]]. 2. **Map** columns to [[business-element]]s and [[business-term]]s. This is the business ↔ technical link, and it is the real product. 3. **Govern** the elements: owners, [[critical-data-element]] flags, policies, [[data-quality]] rules. 4. **Trace** them with [[data-lineage]]. 5. **Package** them as data products, contracts or a [[semantic-layer]]. 6. **Serve** them to people, systems and AI agents. ## One line each - **Catalog:** what exists. - **Glossary:** what it means. - **Dictionary:** what it means *here*. - **Lineage:** where it flows. - **Profiling:** what it is. - **Quality:** whether it's what it should be. - **Contract:** what was promised. - **Semantic layer:** how to calculate it. ## Three scales - **Personal:** your spreadsheet of accounts has columns (technical metadata), you know what "balance" means (glossary), and you'd notice if a number went negative (data quality). You're already doing this informally. - **Business:** a 20-person company's CRM and accounting system disagree on what an "active customer" is. A one-page glossary and two quality rules fix more than any tool purchase. - **Enterprise:** millions of columns across hundreds of systems. None of this works by hand; the mapping is proposed by machines and approved by people. ## Patterns - [Serving metadata to people, systems and agents](https://dealwithdata.com/c/active-metadata-consumption.md): One canonical model, one governed write path, many read paths. How to turn catalog changes into subscribable domain events, and which channel serves which consumer. [draft, dwd_cb2bgvae] Once a catalog has good business metadata, everyone asks for it: "Can we get technical metadata with business metadata?", "Can our AI agent use the definitions?" The failure mode is ten teams building ten extracts. The fix is a shape, not a tool. ## The shape **One canonical model, one governed write path, many read paths.** The hard, valuable part isn't the streaming. It's the governed link: [[business-term]] → [[business-element]] → physical column → lineage. Everything below is delivery. ## Getting changes out Most catalogs sit on a relational database and are poll or harvest based. Changes arrive two ways: - **UI edits** (a steward changes a definition): capture with database [[change-data-capture]] such as GoldenGate or Debezium. - **Harvests** (scanners reload technical metadata): diff successive snapshots and emit a change set. This gives cleaner events than the database log. ## Three layers, not one ``` Catalog DB ──CDC──▶ raw.* topics (private: one team reads these) Harvests ──diff──▶ │ ▼ Translator row changes → domain events │ schema registry, versioned ▼ public topics BusinessElementDefinitionChanged LineageEdgeAdded CdeStewardReassigned ``` Never expose raw CDC from a vendor's internal schema. It is undocumented, changes on upgrade, and one business edit fires a dozen row changes. Subscribers break and drown. ## Pick the channel by consumer | Consumer | Channel | Standard | |---|---|---| | Stewards, analysts | Catalog UI | n/a | | Applications | REST or GraphQL over the metadata graph | OpenAPI | | Change subscribers | Domain-event topics | CloudEvents envelope, AsyncAPI docs | | Bulk analytics | Metadata snapshots as tables | SQL | | AI agents | MCP server plus a search index | MCP | | Pipelines | Emit lineage at run time; enforce contracts in CI | OpenLineage, ODCS | | BI and AI tools | Semantic model export | OSI (Apache Ossie) | All of these are projections of the same model. When a team asks for "a feed", they get one of these, not a custom extract. ## Rules that keep it trustworthy 1. **Stable global IDs** for every term, element and physical asset, or the links can't survive across systems. 2. **The link is a governed object.** A business-to-technical mapping carries an owner, a confidence score, provenance (human, rule or AI) and a status. AI-proposed mappings enter as *proposed*; stewards authorize. 3. **The catalog stays the system of record for its domain.** Publish outward; don't try to make it the hub for everything. 4. **Agents read through the MCP layer, never the database.** That gives one place for entitlements, audit, and "approved definitions only". ## When it becomes a nervous system [[active-metadata]] only matters if something depends on it. Wire in two or three consumers that break when metadata is wrong, such as change management, a regulatory report pipeline, and one AI agent. Until then, it's documentation. ## Using LLMs for long descriptions For bulk drafting of element descriptions, a fast, cheap model is the right default. Quality comes from grounding, not model size: feed the physical name and type, profile statistics, lineage neighbors, the parent dataset and nearby glossary terms. Route critical elements and low-confidence drafts to a stronger model or a human reviewer. - [Never recompute everything: incremental metadata feeds](https://dealwithdata.com/c/incremental-metadata-feeds.md): A feed that reprocesses millions of datasets every run stops scaling. Recompute only what changed, plus what changed rules touch, and reconcile on a schedule. [draft, dwd_q3gbjug4] **The situation:** a feed assigns every harvested dataset to an owning application. At a few thousand datasets it runs in minutes. At millions of datasets and tens of millions of elements, it runs for hours, because it recomputes and reloads everything, every run. ## Find where the time goes Split a run into three phases and time each: 1. **Extract:** read datasets and elements out of the catalog. 2. **Compute:** match each dataset to an application (paths, hosts, naming rules). 3. **Load:** write assignments back. Load is usually the worst. Catalog imports tend to be row by row with versioning and audit, so rewriting a million unchanged rows costs nearly as much as writing changed ones. ## Fixes, biggest first 1. **Go incremental.** Drive the feed from a harvest diff: only added or changed datasets are reassigned. If 1–2% change per cycle, that's a 50–100× reduction. 2. **Assign datasets; let elements inherit.** Touching every element individually multiplies the work by the elements-per-dataset ratio. 3. **Skip no-op writes.** Compare computed to current, and load only the differences. This alone often cuts load time by 90%. 4. **Compute in sets.** Put rules in a table and resolve them with one join or prefix match, not a loop of lookups. An embedded engine such as DuckDB handles millions of rows in minutes. 5. **Partition.** Split by application or path prefix so one failure doesn't rerun everything. ## Is a full recompute ever needed? Not every run. Two kinds of change drive work: - **A dataset changed** → recompute that dataset. - **A rule or reference list changed** (a mapping edited, an application split or retired, a host moved) → recompute only the datasets the old *or* new rule matches. A full pass is justified only for: - **Bootstrap**, and after major platform or model changes. - **Scheduled reconciliation**, weekly or monthly, as a backstop for missed diffs, failed runs and direct edits. Compute everything, still write only differences, and report drift. - **Attestation**, when an auditor needs proof the whole population was evaluated at a point in time. If drift stays near zero for several cycles, stretch the interval. That's your evidence that incremental is safe. ## The target shape ``` harvest → diff → assign changed (set-based) → load differences only → emit AssignmentChanged events ``` **The one thing to get right:** rule changes must emit events too. If someone edits a mapping table and nothing notices, incremental runs go quietly stale. Catching that is what reconciliation is for. ## Three scales - **Business:** the same logic applies to any nightly sync. Diff first, then write only what changed, and your 2-hour job becomes 5 minutes. - **Enterprise:** at millions of objects, "full refresh" is a design bug, not a strategy. ## Teardowns - [Open Semantic Interchange (Apache Ossie): what it is, and what it isn't yet](https://dealwithdata.com/c/open-semantic-interchange.md): OSI is a portable format for semantic models: datasets, fields, joins, metrics and AI context. It is not a catalog, glossary, lineage or governance standard. Export to it; don't build your core on it yet. [draft, dwd_5hhpphgh] **In one line:** OSI lets BI and AI tools exchange the *structure* of a [[semantic-layer]] model. Everything else, such as identity, governance, glossary and lineage, is out of scope or still on the roadmap. ## Status - Started by Snowflake with dbt Labs, Databricks, Salesforce and others; the first spec went public in January 2026 under Apache 2.0. - Now **Apache Ossie**, in the Apache Incubator. Oracle, Cloudera, Dataiku, Denodo and Dremio are among the later joiners. - Only formal release: **0.1.1** (December 2025). The in-development **0.2.0.dev0** has already made one breaking change: one model per document, with the old `semantic_model` array removed. The spec itself says not to depend on it in production. ## The whole model | Construct | What it holds | |---|---| | Semantic model | `version`, `name`, `description`, `ai_context`, and the lists below | | Datasets | Logical entities pointing at a physical `source`, with primary and unique keys | | Fields | Row-level attributes defined by an expression, in one or more SQL dialects | | Relationships | Foreign-key joins between datasets, simple or composite | | Metrics | Aggregate expressions at model level; may span datasets | | `ai_context` | Instructions, synonyms and example questions, on almost any object | | `custom_extensions` | Vendor-tagged JSON for anything the core doesn't cover | Supported dialects include ANSI SQL, Snowflake, Databricks, BigQuery, DAX, MDX, Tableau and Ossie's own portable SQL. ## What it does not cover (yet) - **Stable identifiers.** Objects are identified by `name`; stable IDs are an open discussion. - **Governance.** No owner, steward, certification or criticality field. - **Sensitivity.** PII and confidentiality flags are proposals only. - **Meaning.** No [[business-glossary]] or ontology; the project describes its own scope as structural, not conceptual, interoperability. An ontology working group is active. - **Lineage.** Use OpenLineage. - **Storage and versioning.** No registry and no model versioning. The document's `version` is the *spec* version, not yours. Catalog integration and a semantic registry are roadmap items. ## A governed field, the practical way Carry your governance in an extension until the spec grows the slots: ```yaml - name: mortgage_outstanding_balance expression: dialects: - dialect: ANSI_SQL expression: outstanding_bal_amt datatype: Decimal description: Unpaid principal on the mortgage as of the reporting date, excluding escrow. ai_context: synonyms: ["UPB", "outstanding principal"] custom_extensions: - vendor_name: ACME_CATALOG data: '{"global_id":"BE-000123","cde":true,"status":"certified","model_version":"2026.10.1"}' ``` ## Where to store and version OSI files Nothing is standardized, so do what works today: 1. **Generate** OSI from your catalog, the system of record. Never hand-edit. 2. **Keep it in Git**, one file per model. CI runs the project's validator (schema, unique names, references, SQL syntax). 3. **Tag a release** when stewards approve a change set. 4. **Deploy outward** with the community converters (Snowflake Cortex Analyst, GoodData, Salesforce, Apache Polaris). ## Verdict - **Don't build your core model on it.** It is pre-1.0, has broken its schema once, and lacks IDs and governance. - **Do export to it.** It is cheap, and the BI and AI vendors are converging on it. - **Join the working groups if you can.** Governance, identity and catalog integration are exactly where regulated industries have the most to say, and they are under-represented. ## Three scales - **Business:** if you run dbt or a single BI tool, OSI is mostly invisible; it matters when you add a second tool or an AI analyst. - **Enterprise:** OSI is one outbound channel among several, next to lineage, events and an agent interface. See the consumption pattern. ## Glossary - [Active metadata](https://dealwithdata.com/glossary/active-metadata.md): Metadata that triggers action as it changes, through events, rules and agents, instead of waiting to be looked up. [draft, dwd_mgglmnwq] Passive metadata is documentation: someone has to go look. Active metadata emits events when it changes, and other systems subscribe and act: block a deploy that breaks a critical element, alert a steward, update an AI agent's context, reassign ownership. **Test:** if nothing breaks when the metadata is wrong, it isn't active yet. - [Business element](https://dealwithdata.com/glossary/business-element.md): A governed piece of business data that carries a term's meaning in a specific context, linked to physical columns. [draft, dwd_tthoszn7] A business element answers *what data do we hold for this, and where?* It has a type, format, allowed values, quality rules and often a criticality flag, and it links down to one or more physical columns. Governance (stewards, [[critical-data-element]] status, [[data-quality]] rules) usually attaches here, not to the term. **Example:** the term *Outstanding Balance* is realized by the elements *Mortgage Outstanding Balance* (a critical data element) and *Card Outstanding Balance*. The mortgage element maps to `LOAN_DTL.OUTSTANDING_BAL_AMT` and `MTG_HIST.UPB`. **Not to be confused with** a [[business-term]], which is the meaning without the data. - [Business glossary](https://dealwithdata.com/glossary/business-glossary.md): The organization's agreed vocabulary: business terms, their definitions, synonyms and relationships. [draft, dwd_xzf7arim] A business glossary holds the words the business uses and what each one means, independent of any system. It is where "Customer", "Active Account" and "Outstanding Balance" get one agreed definition each, with synonyms and relationships between terms. **Example:** *Outstanding Balance*: unpaid principal owed as of a date. Synonyms: UPB, principal balance. **Not to be confused with** a [[data-dictionary]] (one system's columns) or a [[data-catalog]] (the inventory of everything). - [Business metadata](https://dealwithdata.com/glossary/business-metadata.md): What data means to the business: definitions, owners, rules, and how important it is. [draft, dwd_6azimhas] Business metadata is the human meaning layered on top of [[technical-metadata]]: business terms and their definitions, business elements, owners and stewards, criticality flags, policies and rules. Unlike technical metadata it can rarely be harvested; someone has to decide it, or a tool has to propose it and someone has to approve it. **Example:** "Outstanding Balance: the unpaid principal owed on a loan as of the reporting date. Owner: Lending. Critical data element: yes." **Not to be confused with** [[technical-metadata]], which says where data *is*. - [Business term](https://dealwithdata.com/glossary/business-term.md): A word or phrase the business uses, with one agreed definition. It describes meaning, not data. [draft, dwd_gdefyuzz] A business term lives in the [[business-glossary]] and answers *what does this word mean?* It has a definition, synonyms and an owner, but no data type and no physical location. One term is usually realized by several [[business-element]]s. **Example:** *Outstanding Balance*: the unpaid principal owed as of a date. **Not to be confused with** a [[business-element]], which is a governed piece of data that carries the term's meaning in one context. - [Change data capture (CDC)](https://dealwithdata.com/glossary/change-data-capture.md): Capturing inserts, updates and deletes from a database as a stream of change events, usually by reading its transaction log. [draft, dwd_xbqpgs7h] CDC reads a database's transaction log and publishes each row change as an event, typically to Kafka. Tools include Oracle GoldenGate, Debezium, Qlik Replicate and Striim. **Caution:** raw CDC from a tool's internal database exposes that tool's private schema. Translate row changes into domain events before anyone else subscribes. See the consumption pattern. - [Critical data element (CDE)](https://dealwithdata.com/glossary/critical-data-element.md): A data element important enough to the business to require extra governance: an owner, quality rules, documented lineage and monitoring. [draft, dwd_zx3rfw43] A CDE is a [[business-element]] flagged as critical, usually because it feeds regulatory reports, financial statements or key risk decisions. The flag makes governance mandatory: a named owner, defined [[data-quality]] rules, documented [[data-lineage]] and active monitoring. **Example:** *Mortgage Outstanding Balance* feeds capital and liquidity reporting, so it is a CDE. **Three scales** - *Personal:* the account numbers and balances your tax return depends on. - *Business:* revenue and customer counts in your board report. - *Enterprise:* the few thousand fields that regulators and auditors trace end to end. - [Data catalog](https://dealwithdata.com/glossary/data-catalog.md): A searchable inventory of data assets across systems, pointing to their metadata, owners and lineage. [draft, dwd_lk4rq6xm] A data catalog answers *what data exists, and where?* It indexes datasets and fields across many systems and connects each to its [[business-glossary]] terms, owners, [[data-lineage]] and quality. People use it to find data; increasingly, AI agents use it to ground themselves. **Example:** searching "loan balance" returns `LOAN_DTL.OUTSTANDING_BAL_AMT`, its definition, its owner, and the reports it feeds. **Not to be confused with** a [[data-dictionary]], which documents one system in depth. - [Data contract](https://dealwithdata.com/glossary/data-contract.md): An explicit, versioned agreement between a data producer and its consumers about schema, meaning, quality and freshness. [draft, dwd_qywin2zi] A data contract turns unspoken expectations into something checkable: the schema, the semantics, quality thresholds, freshness and who to call. It is enforced in CI or at load time. The Open Data Contract Standard (ODCS) is the leading open format. **Example:** "`OUTSTANDING_BAL_AMT` stays `DECIMAL(18,2)`, refreshes daily by 06:00, nulls under 1%." *This site applies the idea to itself:* every content file must satisfy `schema/item.schema.json` before it is published. - [Data dictionary](https://dealwithdata.com/glossary/data-dictionary.md): A column-by-column reference for one system: each field's type, allowed values and local definition. [draft, dwd_6prph2yy] A data dictionary documents a single source system from its own point of view: every table and field, its type, its allowed values, and what it means *in that system*. **Example** | Column | Type | Definition | Allowed values | |---|---|---|---| | `OUTSTANDING_BAL_AMT` | DECIMAL(18,2) | Unpaid principal, USD | ≥ 0 | | `LOAN_STAT_CD` | CHAR(2) | Loan status | `AC` active, `CL` closed, `CO` charged off | **Not to be confused with** a [[business-glossary]], which defines terms for the whole enterprise. Linking dictionary entries to glossary terms is the business-to-technical mapping. - [Data lineage](https://dealwithdata.com/glossary/data-lineage.md): The record of where data comes from, how it is transformed, and where it goes. [draft, dwd_2agj6wti] Lineage traces data from its origin through every transformation to every consumer. *Upstream* lineage explains where a number came from; *downstream* lineage shows what breaks if it changes. OpenLineage is the open standard for emitting it from pipelines at run time. **Example:** core banking system → nightly batch job → `LOAN_DTL.OUTSTANDING_BAL_AMT` → exposure report. - [Data profiling](https://dealwithdata.com/glossary/data-profiling.md): Measuring what data actually looks like: counts, ranges, nulls, distinct values and patterns. [draft, dwd_2xjo4xi3] Profiling describes data as it *is*: row counts, minimums and maximums, null rates, distinct values and value patterns. It is the evidence base for [[data-quality]] rules and a strong input for AI that suggests business meaning. **Example:** `OUTSTANDING_BAL_AMT`: 12M rows, min 0, max 4.2M, 0.3% null. **Not to be confused with** [[data-quality]], which checks data against what it *should* be. - [Data quality](https://dealwithdata.com/glossary/data-quality.md): Rules data must satisfy, and the measured results of checking it against them. [draft, dwd_gxekubnp] Data quality is the gap between what data should be and what it is. Rules ("never negative", "not null", "matches the reference list") are checked, scored and monitored. Active data quality blocks or flags bad data *before* it reaches a critical report. **Example:** "Outstanding balance must be ≥ 0 and not null": 99.7% pass this run. **Not to be confused with** [[data-profiling]], which only describes. - [Master data](https://dealwithdata.com/glossary/master-data.md): The single trusted version of a core business entity such as a customer, product or account. [draft, dwd_qtwdi2p5] Master data management produces one golden record for each core entity by matching and merging copies from many systems. It answers *which* customer, while [[reference-data]] answers *which code*. **Example:** five systems hold variations of the same customer; master data resolves them to Customer #123. - [Reference data](https://dealwithdata.com/glossary/reference-data.md): Shared lists of allowed codes and values used across systems. [draft, dwd_p6v7pmhk] Reference data is the set of allowed values that many systems share: country codes, currency codes, status codes, product categories. It changes rarely, matters everywhere, and causes quiet breakage when systems disagree. **Example:** loan status codes `AC` (active), `CL` (closed), `CO` (charged off). - [Semantic layer](https://dealwithdata.com/glossary/semantic-layer.md): A shared definition of business metrics and dimensions, so every tool and AI agent calculates them the same way. [draft, dwd_5k2ibrbv] A semantic layer defines calculations once (metrics, dimensions, joins) over physical data, and every BI tool and AI agent queries through it. The Open Semantic Interchange (OSI), now Apache Ossie, is the emerging format for exchanging these definitions between tools. **Example:** `Total Exposure = SUM(OUTSTANDING_BAL_AMT)`, synonyms "UPB", "principal". **Not to be confused with** *semantic discovery*, a catalog activity that infers what a column *means*. A semantic layer defines how to *calculate*. - [Technical metadata](https://dealwithdata.com/glossary/technical-metadata.md): Facts about how data is physically stored and moved: names, types, locations, schemas and jobs. [draft, dwd_gubdecks] Technical metadata describes the data as the machines see it: database, schema, table and column names, data types, keys, file paths, and the jobs that read and write them. It is usually harvested automatically from the systems themselves. **Example:** `LOAN_DTL.OUTSTANDING_BAL_AMT`, `DECIMAL(18,2)`, in schema `LND` on an Oracle database, loaded nightly by a batch job. **Three scales** - *Personal:* the folder, file name and format of your tax spreadsheet. - *Business:* the columns in your accounting system's invoice table. - *Enterprise:* millions of datasets harvested from hundreds of platforms. **Not to be confused with** [[business-metadata]], which says what the data *means*.