Why the Cancer AI Alliance selected OMOP as its Common Data Model

Every cancer center documents a wide range of data about patients, including scans, blood work, doctors’ notes, and treatment options. This data is usually tracked via a number of local codes and field names within patients’ electronic health records. These local codes help cancer centers organize and systematize data about their patients. However, when it comes to cross-institutional research, researchers need standardized data so they don’t have to learn each center’s unique data tracking systems. 

Data standardization was one of the first challenges that CAIA set out to solve in the early days of the Alliance. All four participating cancer centers undertook the massive task of standardizing their data to a Common Data Model (CDM), translating their local data into a format that allows federated models to be run across the alliance. As we have previously explained, CAIA selected theObservational Medical Outcomes Partnership (OMOP) Common Data Model version 5.4

In this blog post, we’ll take a deeper look at the reasons why CAIA chose OMOP, how and why data was standardized across seven key domains, and CAIA’s plans for standardizing clinical notes across the Alliance.

What were CAIA’s criteria for a Common Data Model?

CAIA’s criteria for a Common Data Model (CDM)

In CAIA’s federated learning approach, sensitive clinical data remains secure behind each cancer center's firewall. Individual data never leaves its home institution; instead, AI models train locally by traveling to the data. This requires data to be structured identically across all centers with shared quality checks and aligned terminology so that AI models can run consistently. More importantly, aligning to a CDM helps preserve meaning during the federated learning process and allows researchers to interpret the resulting updates. 

CAIA chose OMOP CDM v5.4 because of its mature ecosystem, active open-source community, and widespread adoption. Additionally, it met key technical, legal, and governance criteria:

Data standardization and interoperability

All participating cancer centers must use the same version of the same CDM. The CDM must use standard terminologies — such as SNOMED, LOINC, and RxNorm — to ensure semantic consistency.

Terminology support

The model must provide codes, descriptions, standard mappings, and hierarchies for all key data domains. These mappings must update regularly when source vocabularies change.

Domain coverage

The model must be capable of representing all concepts, facts, and events required by planned CAIA use cases.

Extensibility and scalability

The model must be capable of expanding to new data modalities, including genomics and imaging, as research needs evolve.

Quality assurance

The model must support quality assurance tools and conventions to identify, monitor, and fix data quality issues.

Harmonization and mapping tools

The model must be compatible with existing tools that support data harmonization and vocabulary mapping.

Technology requirements

The model must work with common relational database management systems so all participating cancer centers can adopt it.

Community support

The model must have active communities of practice to provide support and answer questions.

Legal

The model cannot have licensing restrictions that limit its use and extensibility for CAIA's objectives.

Standardizing cancer center data across seven domains

Using OMOP, CAIA’s participating cancer centers organized clinical data into seven distinct domains or categories. These datasets contain de-identified information about each clinical encounter, including: 

  1. Person: Basic patient information, such as age, sex, race, and ethnicity.

  2. Visit: Records of patient encounters, such as clinic visits, hospital stays, or treatment visits.

  3. Condition: Diagnoses and health conditions, including cancer diagnoses.

  4. Procedure: Procedures patients received, such as biopsies, surgeries, or other interventions.

  5. Drug: Medications and treatments patients received, including cancer therapies.

  6. Measurement: Vital signs, lab results, test results, and other recorded clinical measurements.

  7. Death: Mortality information, when available.

De-identified clinical data standardized across 7 domains using OMOP

By standardizing data across these seven domains, researchers can track patients’ experience by mapping how specific diagnoses (Condition) lead to particular therapies (Drug/Procedure) and directly impact patient attributes (Measurement) over time. This dataset has the potential to help researchers understand the effectiveness of particular treatments, find patterns in how different patient groups respond to therapies, and discover which clinical decisions lead to the best long-term outcomes.

Once the data has been de-identified and standardized at a cancer center, data that is relevant to a specific project is moved into a staging area. A local edge node with access to this de-identified data trains AI models locally at each cancer center, sharing only model weights, insights, and summaries with the central CAIA orchestration layer so that sensitive clinical data never leaves the cancer centers’ firewall.

Over the next few months, CAIA will focus on developing the requirements necessary for standardizing clinical notes across the Alliance. This work will include defining scope, de-identification requirements, and mapping how these notes can be integrated with existing OMOP data. Each participating cancer center will assess its current data gaps and determine how to close them.

If you’d like to learn more about CAIA, subscribe to our newsletter for updates and follow us on LinkedIn and X.

Next
Next

Multi-cloud federated learning: The role of edge nodes at CAIA