Data insights

The “open” in Open Data Infrastructure means interoperability

July 29, 2026
The “open” in Open Data Infrastructure means interoperability
Interoperability means the optionality to adapt to rapidly changing future demands.

The Open Data Infrastructure is how Fivetran and dbt Labs describe an architectural approach that allows organizations to store data once in an open format and use it across the full range of tools, compute engines, and AI systems without being locked into a single vendor.

Openness, in particular, is about open standards and interoperability. In an Open Data Infrastructure, data moves freely, and the different elements of the data stack can be tailored as needed for specific use cases. This interoperability, which enables organizations to scale their operations, engineer context as needed, and control costs, is essential for AI readiness and agentic workflows.

The keystone of the Open Data Infrastructure is the modern data lake, because the combination of open table formats and commodity storage decouples compute and storage. This allows teams to maintain their data in a single platform while remaining agnostic about the exact use cases and their corresponding tools that will emerge in the future. In short, openness is about preserving optionality.

A closed data infrastructure, by contrast, is the opposite: limited data portability, restrictions on its usage, and monolithic systems that are costly to exit or replace. A closed data infrastructure is one in which choices are high-commitment and difficult to reverse, and therefore also high-risk and likely to limit future choices in ways that impair your organization’s future agility. Even worse, closed formats, such as traditional data warehouses and data lakes, introduce the risk of building multiple closed data infrastructures for different use cases, leading to escalating engineering burdens, administrative complexity, and costs.

[CTA_MODULE]

“Open standards” doesn’t mean “open source” (and that’s ok)

Open standards and open source are orthogonal concepts and can comfortably coexist. The defining characteristic of the Open Data Infrastructure is the relationship between open standards, which ensure interoperability by allowing disparate systems to talk to each other. The systems in question may be open source or proprietary, so long as the costs of switching are low.

dbt Core, an open source tool, is a good example of an open source system that can be leveraged within an Open Data Infrastructure alongside proprietary systems such as Fivetran. 

A closed data infrastructure may feature open source systems as well, so long as important elements of the data stack are tightly coupled and difficult to switch. An infrastructure in which open source tools are attached to a closed architectural center, such as a traditional data warehouse, is qualitatively different from one in which either open source or proprietary tools are attached to an open architectural center, namely a modern data lake.

Pitfalls that close your data infrastructure

An Open Data Infrastructure must be modular, commercially portable, and context-rich. Organizations should be able to choose and replace components freely, move and use their data without artificial barriers, and preserve the meaning and policies attached to that data across the entire ecosystem.

There are three key pitfalls that stand in the way of an Open Data Infrastructure:

  1. Architectural enclosure
  2. Vendor lock-in
  3. Fragmented semantics and context

Architectural enclosure means that systems are designed so that storage, compute, governance, and other capabilities are tightly coupled, making components difficult to replace or combine independently. Obligately bundled compute and storage is the clearest example, because it turns a technical design choice into a long-term dependency on a particular platform. 

Vendor lock-in means that even when customers technically own their data, vendors may impose egress fees, proprietary interfaces, restrictive contracts, or non-portable metadata and business logic that make moving or using that data elsewhere costly and impractical.

Fragmented semantics and context mean that the information required to understand and responsibly use data is incomplete, inconsistent, or trapped inside individual systems. This includes semantic context, such as definitions, schemas, lineage and relationships, as well as governance context, including ownership, sensitivity, permissions and usage restrictions. Without this context, data may be physically accessible but still difficult to interpret, govern or use safely across tools.

Why the Open Data Infrastructure is essential for the future

AI imposes far greater demands than analytics based on traditional human decision support. Models and agents need data that is persistent, machine-readable, and governed by explicit rules for precedence, conflict resolution, lineage, and reuse that can’t be locked behind a vendor’s application or control plane. 

At the same time, the long-term pattern across technology is clear: standardization and commoditization create modular ecosystems in which components can be improved, replaced, and combined independently. Data storage is following the same path through open table formats and data lakes, which separate durable data from the engines used to process it. AI will follow as well, both at the foundation-model layer and in the agentic harnesses built around models. But AI remains a rapidly developing market: today’s leading model, framework, or platform can be displaced quickly by a new technical breakthrough or business failure. 

Organizations therefore need an architecture in which their data, metadata, and governing context persist even as everything above them changes. The future depends on preserving an organization’s most unique and durable asset — data — while retaining the freedom to swap compute engines, models, orchestration layers, and applications whenever better options emerge, without disruption.

[CTA_MODULE]

Check out our ODI scorecard for a practical look at how well vendors help customers access and use their own data.
Learn more
How "open" is your data infrastructure?
Find out

Related blog posts

Start for free

Join the thousands of companies using Fivetran to centralize and transform their data.

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.