DATA-ANALYTICS
Microsoft and Google join Apache Ossie project
Major tech firms support Apache Ossie to establish a common open specification for sharing semantic data models across analytics and AI platforms.
- Read time
- 5 min read
- Word count
- 1,011 words
- Date
- Oct 1, 2026
- Key Takeaways:
- Microsoft is creating a two-way converter between Power BI and the Ossie semantic format
- The Apache Ossie project currently has the backing of more than 60 major technology companies
- Google is currently in the process of joining the Apache Incubator project for semantic interchange
- The current version of the Ossie specification is a development draft labeled as version 0.2
🌟 Non-members read here
Microsoft and Google are now supporting a new initiative to build an open specification for sharing semantic models across various data, analytics, and artificial intelligence environments. This project currently includes more than 60 major organizations, including industry leaders like Nvidia, Salesforce, Oracle, and Snowflake, aiming to streamline how data is interpreted.
Standardizing Semantic Data Across the Enterprise
The initiative originally launched as the Open Semantic Interchange before moving into the Apache Incubator in June under the name Apache Ossie. It relies on standard formats like JSON and YAML to describe semantic models. These models include essential information such as datasets, fields, internal relationships, business metrics, and contextual data for artificial intelligence applications. By creating a unified language, the project allows these elements to function across different environments like Tableau, Snowflake, and Databricks.
The architectural approach for this project follows a hub-and-spoke configuration. Apache Ossie serves as the central, common format that connects different proprietary systems. Converters act as the spokes, translating data between the individual software and the shared standard. This setup removes the need for businesses to build unique, one-to-one connections every time they want to move data between two specific platforms. It simplifies the infrastructure required to manage complex data ecosystems.
Microsoft is actively contributing to this ecosystem by building a bidirectional converter for Power BI. This tool will allow users to take semantic definitions created in Power BI and represent them in the Ossie format. Conversely, models created in the Ossie standard can be imported back into the Power BI environment. Microsoft is also advocating for the inclusion of DAX, which is the expression language used for Power BI calculations, as a recognized language within the project.
By including DAX, organizations can keep their business logic intact when moving models between different platforms. This ensures that the mathematical formulas and rules that give data its meaning stay consistent. Meanwhile, Google is finalizing its membership in the project. Although Google has not shared all its specific plans, its BigQuery SQL dialect is already part of the supported list. This inclusion reflects the significant presence BigQuery has in modern enterprise data strategies.
Operational Efficiency and Metric Consistency
The participation of major cloud providers could significantly lower the amount of manual work required for data engineering. Teams that move workloads between different analytics platforms often find themselves repeating the same configuration tasks. With a unified standard, these teams can avoid the tedious process of rebuilding their data structures from scratch. This efficiency allows developers to focus on higher-level tasks rather than basic configuration.
One major problem in large organizations is metric drift. This happens when different departments or software tools calculate the same business metric in slightly different ways. When developers have to manually recreate definitions across multiple tools, errors often creep in. By using a shared standard, a metric can be defined once as a code artifact. This artifact can then be versioned and reviewed just like any other software component, ensuring everyone uses the same math.
Consistency is also vital for the future of automated agents and artificial intelligence applications. If an AI agent interprets a business metric differently than a human analyst, it leads to confusion and poor decision-making. A unified semantic layer provides a single source of truth for these autonomous systems. When agents have access to clear, consistent business context, leaders can deploy them with more confidence across the organization.
The improved productivity of developers is another key factor for businesses adopting new technology. When onboarding a new analytics tool or launching a new AI application, the time spent on validation is often a bottleneck. If the semantic definitions are already in a portable format, the setup time drops. This allows companies to be more agile and responsive to market changes by deploying new data tools faster than they could previously.
Challenges in Portability and Vendor Lock-in
While the project offers significant benefits, experts note that portability is not the same as total compatibility. The ability to move a model between platforms depends on how well the new system can understand specific proprietary logic. Even if a metric moves to a new system successfully, the way that system handles filters or time-based calculations might differ. This means the final results could still vary slightly between different platforms.
The technical gap is especially clear with sophisticated languages like DAX used in Microsoft tools. Some of these functions do not have a direct equivalent in standard SQL. Because of this, a complex model might lose some of its unique behavior during the translation process. Senior technical leaders will still need to perform thorough validation to make sure the math remains accurate after a conversion. Automated tools can help, but human oversight remains necessary for production environments.
Security and governance also present hurdles for the current version of the specification. Right now, the standard focuses on data structures and metrics rather than access controls. Details like row-level security and user permissions are not yet part of the core format. This means that even if the data model moves easily, administrators still have to manually set up security rules and certification statuses on each new platform they use.
Furthermore, Apache Ossie is still in its early stages of development. As an incubating project, it has not yet become a permanent fixture of the Apache Software Foundation. The current 0.2 version is a draft, which means the underlying rules and capabilities could change at any time. This lack of a final, stable version creates some risk for companies looking for a long-term solution. Some large-scale features, like references between different models, are also missing from the current draft.
Even if the industry adopts this standard widely, the dream of ending vendor lock-in remains elusive. Software providers will likely continue to develop unique features and high-performance engines that cannot be captured in a general format. While the definitions of the data become easy to move, the actual performance and specialized capabilities will stay tied to specific vendors. Enterprises may find that they have simply traded one type of technical dependency for another.
References
- Attribution: Valentin Podkamennyi, VP Insights
- Citations: Microsoft, Google back Apache Ossie to make enterprise data and AI platforms more interoperable, Info World
- Mentions: Databricks, Nvidia, Oracle, Salesforce, Snowflake
- About: Apache Software Foundation, Microsoft, Google