AI identifies previously unrecognized health insights in routine sleep studies
Bilal E, Araujo MLD, Beck KL, Heinzinger CM, Ghosn S, Saab CY, Foldvary-Schaefer N, Rogers JL, Mehra R · Nature Communications (2026) 17:7603 · Cleveland Clinic & IBM Research
A single night of clinical polysomnography records millions of physiologic data points — and then gets collapsed into one number. This study shows how much prognostic signal that number throws away.
Beyond the apnea–hypopnea index
Sleep apnea affects nearly a billion people worldwide, yet interpretation of the gold-standard sleep study is still reduced largely to the apnea–hypopnea index (AHI). AHI anchors current diagnostic criteria, but it captures only a narrow slice of sleep physiology and carries limited prognostic value. The Cleveland Clinic–IBM team asked what remains hidden in the rest of the recording.
A transformer foundation model trained on 10,000+ sleep studies
The researchers trained a RoBERTa-style multimodal transformer on raw, high-resolution polysomnography time series — respiratory, cardiac, oxygenation, and neurophysiologic channels — from more than 10,000 clinical recordings linked to electronic medical records. Rather than hand-crafted indices, the model learns high-dimensional embeddings of sleep physiology, which are then clustered into patient phenotypes.
Five latent risk groups with divergent trajectories
Clustering revealed five stable risk groups (RG1–RG5), from low-risk profiles with minimal abnormality to a smaller high-risk group with markedly worse outcomes. RG5 carried more than double the mortality risk of the lowest group (HR = 2.38), with elevated incidence of major adverse cardiovascular events (HR = 1.64), heart failure (HR = 1.65), myocardial infarction (HR = 1.84), atrial fibrillation (HR = 2.23), cognitive impairment (HR = 1.93), and epilepsy (HR = 2.40). Conventional AHI severity categories showed limited predictive value across the same cohort.
Ablation analyses showed that non-respiratory channels contributed critically to correct group assignment — evidence that the risk structure is genuinely multisystem physiology, not repackaged breathing events. The framework also generalized to the independent Sleep Heart Health Study, separating high- from low-risk patients despite lower-resolution legacy data.
Why it matters for the thesis
This is the sleep-medicine analogue of the CT-derived VO₂max hypothesis: a routine diagnostic study already contains latent information about functional and physiologic reserve, and a foundation model can surface it without adding a new test. Both argue that the missing vital sign is not unmeasured — it is unread.
Source
Bilal E, et al. “A foundation model for sleep-based risk stratification and clinical outcomes.” Nature Communications 17, 7603 (2026). DOI: 10.1038/s41467-026-75326-9
Obstructive sleep apnea and cardiorespiratory fitness: a review. Untreated OSA is associated with blunted peak VO₂, impaired chronotropic and oxygen-pulse response, and reduced exercise capacity — linking nocturnal physiology directly to measured cardiorespiratory fitness and, by extension, to the functional-reserve hypothesis behind CT-derived VO₂max.
Cardiovascular effects of continuous positive airway pressure in patients with heart failure and obstructive sleep apnea
Kaneko Y, Floras JS, Usui K, et al. · New England Journal of Medicine (2003) · Stanford Critical Care archive
Treating the airway changed the heart. One month of nightly CPAP measurably improved left ventricular function in patients with heart failure and obstructive sleep apnea.
The trial
Patients with treated heart failure and coexisting obstructive sleep apnea were randomized to receive nightly CPAP or usual care for one month. Outcomes included daytime blood pressure, heart rate, and left ventricular ejection fraction measured by radionuclide angiography, alongside overnight polysomnographic indices.
What CPAP moved
The CPAP group showed reduced daytime systolic blood pressure and heart rate and a significant improvement in left ventricular ejection fraction, with abolition of obstructive events overnight. The control group showed no comparable change. The implication is mechanistic: repetitive nocturnal apnea imposes a measurable, reversible load on the failing heart through negative intrathoracic pressure swings, intermittent hypoxia, and sympathetic surges.
Why it matters for the thesis
This is the interventional bookend to the Cleveland Clinic foundation-model work. If sleep physiology encodes latent cardiovascular risk, this trial shows that the risk is partly modifiable — and that cardiopulmonary reserve is a moving target that responds to therapy. Any imaging-derived surrogate for functional reserve therefore has to be read as a trajectory, not a fixed label.
Source
Kaneko Y, et al. “Cardiovascular Effects of Continuous Positive Airway Pressure in Patients with Heart Failure and Obstructive Sleep Apnea.” N Engl J Med 348:1233–1241 (2003).
Modern medicine can image a perforation in an intestine with millimeter precision. But getting someone to tell you it is there can take hours. Sometimes until morning.
The 2 a.m. availability problem
The article opens with a scenario that every emergency and surgical team recognizes: a CT completes at 2:17 a.m., but the radiologist covering three hospitals from home will not read it until 4 a.m. During that gap, localized infection can spread, obstructed flow can progress to tissue death, and the window for the right intervention narrows. The bottleneck is not the scanner; it is the availability of interpretation.
From narrow detection tools to an intelligence layer
The FDA has authorized nearly 1,000 radiology AI tools, most targeting narrow findings like lung nodules or brain hemorrhages. The authors argue that what clinicians actually need is broader coverage: a regulated intelligence layer that turns pixels into answers for whoever is asking — ED physicians, surgeons, specialists — each with different questions about the same scan.
Speed is the most obvious benefit. Even 30 to 60 minutes of acceleration can change outcomes for time-sensitive conditions. But the deeper opportunity is systematic capability: triage that avoids unnecessary surgery, population screening that flags incidental findings before they become emergencies, and quantitative tracking of tumor response across serial scans.
Semi-autonomy, not replacement
The proposed path is semi-autonomy: AI generates preliminary interpretations that clinicians can act on immediately, while radiologists review, confirm, and focus on complex cases requiring nuanced judgment. This extends radiologist expertise across time and space without requiring full regulatory autonomy overnight.
Why it matters for the thesis
The same “intelligence layer” logic applies to CT-derived VO₂max. A diagnostic CT is already being acquired; the question is whether the data already inside it can be turned into a clinically actionable answer without adding a new test. Building that capability means turning imaging from a static report into an on-demand physiologic query — exactly the shift the Pew authors describe for radiology AI.
Source
“From Medical Scans to AI Answers in Seconds.” The Pew Charitable Trusts, Trend Magazine, November 13, 2025.
A multimodal sleep foundation model for disease prediction
Published in Nature Medicine (2025). Stanford-led collaboration on SleepFM, a foundation model for overnight physiology.
Sleep is the longest continuous physiologic recording most humans ever produce. SleepFM asks what happens when a foundation model is trained directly on that signal — across brain, heart, and breath — and then used as a general substrate for predicting disease.
What the paper shows
The authors train a multimodal foundation model on large-scale polysomnography — EEG, ECG, respiratory effort, oximetry, EMG, and related channels — from hundreds of thousands of overnight studies. Rather than optimising a single classifier for a single condition, they learn a shared representation of overnight physiology that can be adapted, with minimal supervision, to downstream tasks.
Across those tasks, SleepFM predicts disease onset and progression spanning cardiometabolic (hypertension, atrial fibrillation, type-2 diabetes), neurologic (Parkinson's, cognitive decline), and psychiatric (depression) domains — often outperforming models trained only on the target label, and doing so from a single night of passively collected signal.
Why it matters for the thesis
SleepFM is the sleep analogue of what CT-derived VO₂max estimation aims to do for imaging: extract a high-value physiologic readout from data that is already being captured, without adding a new test to the clinical pathway. The transferable idea is that foundation models make the marginal cost of a new prediction close to zero once the substrate exists.
The measurement bottleneck in medicine is rarely the sensor. It is the workflow that surrounds the sensor.
Overnight PSGs, contrast-enhanced CTs, ambulatory ECGs, and echoes all contain far more physiologic information than the report ever extracts. A foundation model turns those recordings into a general-purpose feature space that specialists can query for questions the original study was never ordered to answer.
Open questions
Generalisation across sleep-lab hardware, demographic shift, and home-based recordings from consumer devices remains the honest limit of any PSG-trained model. The same is true of imaging foundation models trained on tertiary-centre scanners. Prospective validation across sites, and pre-specified analysis of subgroup performance, will decide whether SleepFM becomes clinical infrastructure or a benchmark trophy.
Source
“A multimodal sleep foundation model for disease prediction.” Nature Medicine, 2025.
The most important number your doctor has never measured
By Joel Selanikio, MD. Originally published at Future Health.
There's a number that predicts whether you'll be alive in ten years better than your blood pressure, your cholesterol, your BMI, or whether you smoke. It's called VO2 max. And you've probably never heard your doctor mention it.
The number
VO2 max measures the maximum amount of oxygen your body can use during intense exercise. It is a single figure — milliliters of oxygen per kilogram of body weight per minute — that captures how well your heart, lungs, blood vessels, and muscles work together. Researchers call the broader concept cardiorespiratory fitness, or CRF. VO2 max is how you measure it.
And it may be the single most powerful predictor of whether you live or die that medicine has ever identified.
In 2022, a study of more than 750,000 U.S. veterans found that even modest improvements in VO2 max were associated with a 13 to 15 percent reduction in mortality risk — regardless of age, BMI, sex, or existing conditions. A 46-year follow-up of over 5,000 men in Copenhagen found that each 1 ml/kg/min increase was associated with 45 additional days of life.
A 2018 Cleveland Clinic analysis of 122,007 adults found that those with the lowest VO2 max had four times the mortality risk of those with the highest. The mortality difference between the least fit and the most fit people was larger than the difference between non-smokers and smokers, or between people with and without diabetes.
The recommendation no one followed
In 2016, the American Heart Association published a landmark scientific statement arguing that cardiorespiratory fitness should be treated as a clinical vital sign — measured routinely, alongside blood pressure, heart rate, and temperature. At a minimum, all adults should have it estimated every year.
That was nearly a decade ago. Has your doctor measured it? Mentioned it?
The single most powerful predictor of mortality we have doesn't fit inside a 15-minute office visit. So it doesn't get done.
The migration
While the clinical system was busy not measuring VO2 max, someone else was — on millions of people's wrists. Apple Watch has been estimating VO2 max since 2020, passively, using heart-rate sensors and GPS during outdoor walks, runs, and hikes. No mask. No technician. No appointment.
A wrist estimate is not as good as a lab test. But it is better than a gold-standard test that never gets done. And because wearables measure it every time you take a brisk walk, they capture trajectory — response to training, illness, and seasonal variation.
Source
Joel Selanikio, MD. The original essay expands on the clinical evidence, the AHA statement, and the implications of consumer-wearable measurement for preventive care.
Why personalized customer experiences are the future of healthcare
By Eangelica Aton. Originally published in The AI Journal.
The typical healthcare customer service experience probably sounds familiar: a customer calls their insurer, navigates a labyrinth of menus, and is promptly placed on hold. It is a toss-up whether they actually get the help they need.
Conversation intelligence AI can present relevant plan information to representatives in real time, reduce call times, and free them to focus on being engaged and empathetic listeners. Omnichannel operations meet patients on the channel they already prefer — phone, portal, or telehealth.
AI should remove the work around the relationship, not replace the relationship itself.
The best healthcare AI will not feel like a product
By Roberto Cruz, Co-Founder & CEO, TietAI.
Every healthcare AI demonstration follows the same script. The interface is polished. The patient is summarised in seconds. Then it reaches a hospital on a Monday morning, and nobody opens it.
The question is how the value of the AI becomes part of the work clinicians are already doing.
Embedded, not absent
The most valuable systems will increasingly operate behind the scenes — preparing information before it is requested, reconciling records before discrepancies reach the clinician, and routing work to the correct team. They act through the systems hospitals already run.
Capacity is the real economic case
Healthcare is under capacity pressure, not just financial pressure. The companies that win will help existing institutions do materially more with the people and systems they already have. That demands proof in operational terms: time returned, delays removed, errors prevented, capacity created.
The best healthcare AI will not be the system clinicians remember using. It will be the reason the work was easier to do.
Attribution
Roberto Cruz, Co-Founder & CEO, TietAI — makers of Hydra, the AI infrastructure layer for healthcare.
Imaging & AI · Case Study 08
Automated detection of traumatic white matter injury using voxel-based morphometry of DTI
Le TH, Mukherjee P, Manley GT, et al. Proceedings of the International Society for Magnetic Resonance in Medicine, Miami, 2005.
Traumatic brain injury often leaves white matter damage that conventional scans miss. This early work asked whether diffusion tensor imaging, paired with voxel-based morphometry, could make that injury visible and quantifiable.
The clinical problem
Closed head trauma produces edema, hemorrhage, and contusions, but diffuse axonal injury — shearing of axons from rotational acceleration — is a major source of long-term disability. Conventional CT and MR often underestimate its extent. T2* gradient echo improves on CT by detecting blood products, yet most DAI lesions are non-hemorrhagic. FLAIR detects some of these foci, but still misses a large fraction of the true white matter burden.
The result is a gap between what imaging reports and what patients experience. A scan can read as nearly normal while cognitive, psychiatric, and functional deficits persist.
Why DTI changes the picture
Diffusion tensor imaging measures the directional motion of water in brain tissue. In healthy white matter, water diffusion is anisotropic — it moves preferentially along axonal bundles. Fractional anisotropy captures that directionality. When axons are sheared or demyelinated, FA drops.
Unlike conventional sequences that look for focal lesions, DTI probes microstructural integrity across the entire white matter skeleton. That makes it inherently suited to diffuse injury patterns that evade lesion-counting approaches.
Voxel-based morphometry against a normal database
The authors built a normal FA database from 15 adult volunteers, spatially normalizing each subject's DTI data to a standardized FA template. Individual trauma patients — 11 in this pilot — were then registered to the same template and compared voxel-by-voxel against the normal distribution.
This is the imaging analogue of a reference-range laboratory test: instead of relying on radiologist impression, the analysis flags voxels where a patient's FA deviates from expected normal variation. The automation removes inter-reader variability and scales to whole-brain screening.
Parallel imaging at 3T
High-field DTI near the skull base is traditionally degraded by susceptibility artifacts. By using parallel imaging, the team reduced geometric distortion enough to assess inferior frontal and temporal regions — areas particularly vulnerable to trauma and previously difficult to evaluate with DTI.
Technical advances in acquisition are what turn a promising biomarker into a clinically usable one.
Relevance today
This 2005 study prefigures the current wave of quantitative neuroimaging: normative atlases, voxel-wise deviation mapping, and AI-assisted detection of subtle injury. The same logic — compare a patient's imaging against a learned distribution of normal — now powers large-scale efforts in brain age estimation, lesion detection, and radiomic phenotyping.
For the thesis, the paper is a reminder that the hardest clinical problems are often measurement problems first. Whether the target is white matter integrity or cardiorespiratory fitness, the path is similar: acquire a quantitative signal, establish a normative reference, and automate the comparison so clinicians can act on it.
Source
Le TH, Mukherjee P, Manley GT, et al. “Automated detection of traumatic white matter injury using voxel-based morphometry of diffusion tensor images: a 3T study with parallel imaging.” Proceedings of the 13th Annual Meeting of the International Society for Magnetic Resonance in Medicine, Miami, 2005.