The promise of federated learning for cancer research
Breakthroughs in cancer research can come from studying the routine clinical data collected during every patient visit.
A single lab result in isolation tells only part of the story. But by mapping routine data across more than a million cancer patients, researchers can track how clinical signals (like standard blood work) may shift over time.
This data can help researchers uncover hidden patterns to
track disease trajectories
anticipate treatment responses
forecast complications
Unlocking clues in this data is one of the great opportunities in modern healthcare.
AI has the potential to help, but only if it can learn from the experiences of millions. While individual cancer centers can build AI models from their own patients’ experiences, truly unique insights and patterns emerge only when millions of these stories are brought together.
While there’s a unique story in every test, scan, or treatment a patient receives, this data rightfully is closely safeguarded by healthcare institutions and is bound by essential privacy and regulatory guidelines. However researchers need the combined insights from millions of cancer patients’ experiences in order to achieve the scale needed to uncover new patterns and shed light on new therapies.
It is not possible moving patient beyond their institution’s firewalls, but with advances in machine learning and AI technologies, researchers can still gather anonymized insights at scale in order to advance cancer research.
Four cancer centers have joined forces with AI technology leaders to explore how researchers can collaborate across institutional silos while protecting patient privacy. CAIA’s goal is to accelerate the pace of cancer with a federated learning platform that allows researchers to securely share dataand build AI models that will ultimately help diagnose, treat, and cure cancer.
With a federated learning approach,AI models learn from diverse data sources without the data ever leaving its home institution. Instead of pooling private patient files, a baseline AI model is sent to each institution. The model learns from the local data, and then instead of sending any patient-level data back, the center sends only summaries and model weights: mathematical updates representing what the model learned.
This updated, more intelligent model is then sent back to the centers, and the cycle begins again. Each iteration makes the model smarter and more accurate, all while the patient records themselves remain safely behind each institution’s firewalls.
Stage 1: Standardize
Stage 2: Federate
Stage 3: Research