---
id: dwd_5hhpphgh
type: teardown
slug: open-semantic-interchange
title: "Open Semantic Interchange (Apache Ossie): what it is, and what it isn't yet"
summary: "OSI is a portable format for semantic models: datasets, fields, joins, metrics and AI context. It is not a catalog, glossary, lineage or governance standard. Export to it; don't build your core on it yet."
status: draft
classification: public
authors:
  - name: Shan Umasankar
    url: https://dealwithdata.com
    role: author
assisted_by: [Claude]
scales: [business, enterprise]
tags: [standards, semantic-layer, ai]
terms: [dwd_5k2ibrbv, dwd_xzf7arim, dwd_2agj6wti, dwd_lk4rq6xm]
related: [dwd_wg5zzqg2, dwd_cb2bgvae]
as_of: 2026-10-07
sources:
  - title: "Apache Ossie: Core Metadata Specification"
    url: https://github.com/apache/ossie/blob/main/core-spec/spec.md
    accessed: 2026-10-07
  - title: "Apache Ossie Roadmap"
    url: https://github.com/apache/ossie/blob/main/ROADMAP.md
    accessed: 2026-10-07
  - title: "OSI Community Update: What's New and What's Next (April 2026)"
    url: https://ossie.apache.org/updates/osi-april-2026-community-update/
    accessed: 2026-10-07
  - title: "Snowflake: OSI specification now live"
    url: https://www.snowflake.com/en/blog/open-semantic-interchanges-specs-finalized/
    accessed: 2026-10-07
created: 2026-10-07
updated: 2026-10-07
version: 1
---

**In one line:** OSI lets BI and AI tools exchange the *structure* of a
[[semantic-layer]] model. Everything else, such as identity, governance,
glossary and lineage, is out of scope or still on the roadmap.

## Status

- Started by Snowflake with dbt Labs, Databricks, Salesforce and others; the
  first spec went public in January 2026 under Apache 2.0.
- Now **Apache Ossie**, in the Apache Incubator. Oracle, Cloudera, Dataiku,
  Denodo and Dremio are among the later joiners.
- Only formal release: **0.1.1** (December 2025). The in-development
  **0.2.0.dev0** has already made one breaking change: one model per
  document, with the old `semantic_model` array removed. The spec itself
  says not to depend on it in production.

## The whole model

| Construct | What it holds |
|---|---|
| Semantic model | `version`, `name`, `description`, `ai_context`, and the lists below |
| Datasets | Logical entities pointing at a physical `source`, with primary and unique keys |
| Fields | Row-level attributes defined by an expression, in one or more SQL dialects |
| Relationships | Foreign-key joins between datasets, simple or composite |
| Metrics | Aggregate expressions at model level; may span datasets |
| `ai_context` | Instructions, synonyms and example questions, on almost any object |
| `custom_extensions` | Vendor-tagged JSON for anything the core doesn't cover |

Supported dialects include ANSI SQL, Snowflake, Databricks, BigQuery, DAX,
MDX, Tableau and Ossie's own portable SQL.

## What it does not cover (yet)

- **Stable identifiers.** Objects are identified by `name`; stable IDs are an
  open discussion.
- **Governance.** No owner, steward, certification or criticality field.
- **Sensitivity.** PII and confidentiality flags are proposals only.
- **Meaning.** No [[business-glossary]] or ontology; the project describes
  its own scope as structural, not conceptual, interoperability. An ontology
  working group is active.
- **Lineage.** Use OpenLineage.
- **Storage and versioning.** No registry and no model versioning. The
  document's `version` is the *spec* version, not yours. Catalog
  integration and a semantic registry are roadmap items.

## A governed field, the practical way

Carry your governance in an extension until the spec grows the slots:

```yaml
- name: mortgage_outstanding_balance
  expression:
    dialects:
      - dialect: ANSI_SQL
        expression: outstanding_bal_amt
  datatype: Decimal
  description: Unpaid principal on the mortgage as of the reporting date, excluding escrow.
  ai_context:
    synonyms: ["UPB", "outstanding principal"]
  custom_extensions:
    - vendor_name: ACME_CATALOG
      data: '{"global_id":"BE-000123","cde":true,"status":"certified","model_version":"2026.10.1"}'
```

## Where to store and version OSI files

Nothing is standardized, so do what works today:

1. **Generate** OSI from your catalog, the system of record. Never hand-edit.
2. **Keep it in Git**, one file per model. CI runs the project's validator
   (schema, unique names, references, SQL syntax).
3. **Tag a release** when stewards approve a change set.
4. **Deploy outward** with the community converters (Snowflake Cortex
   Analyst, GoodData, Salesforce, Apache Polaris).

## Verdict

- **Don't build your core model on it.** It is pre-1.0, has broken its schema
  once, and lacks IDs and governance.
- **Do export to it.** It is cheap, and the BI and AI vendors are converging
  on it.
- **Join the working groups if you can.** Governance, identity and catalog
  integration are exactly where regulated industries have the most to say,
  and they are under-represented.

## Three scales

- **Business:** if you run dbt or a single BI tool, OSI is mostly invisible;
  it matters when you add a second tool or an AI analyst.
- **Enterprise:** OSI is one outbound channel among several, next to
  lineage, events and an agent interface. See the consumption pattern.
