Flagship: Computing for Health
Foundation Models for Biosignal Analysis [Luca Benini]
This project develops a new generation of artificial intelligence models for analyzing biological signals such as brain activity (EEG), heart signals (ECG), and blood flow signals (PPG), muscle signals (EMG). These signals are widely used in healthcare, for example in monitoring neurological disorders, detecting heart conditions, or enabling brain-computer interfaces.
Current AI systems are typically designed for a single type of signal and struggle to generalize across different sensors, hospitals, or patient populations. Our project addresses this limitation by building foundation models that can learn from different types of biosignals. This model can adapt to different sensor configurations and operate efficiently even on long recordings.
To achieve this, we combine three key innovations: (1) a method to standardize heterogeneous sensor layouts, (2) efficient architectures that scale to large datasets, and (3) adaptive computation mechanisms that adjust processing effort based on signal complexity.
The outcome is a robust and efficient AI system that improves diagnostic accuracy, reduces the need for labeled data, and enables deployment on resource-constrained devices such as wearable monitors. This contributes to more scalable and accessible healthcare technologies.
Key achievements:
- Developed a multi-modal biosignal foundation model architecture integrating topology-agnostic input handling, efficient temporal modeling, and scalable attention mechanisms.
- Established a scalable training pipeline on HPC infrastructure, demonstrating near-linear scaling and high parallel efficiency for large-scale biosignal pre-training
- Built the BioFoundation framework, enabling reproducible, modular, and distributed training across heterogeneous biosignal datasets
Contact
Fall Injury Classification [Torsten Hoefler]
Our objective is to decrease the rate of hospital admissions due to fall-related injuries while also minimizing the complications that arise when such injuries do occur.
The injury risk to the hip when falling can be assessed by using medical imaging techniques to build a three-dimensional model of a person’s hip, which is then used as input for finite element simulations of different fall scenarios. Such simulations are computationally expensive, and its inputs involve confidential patient data, and thus cannot be performed in the cloud.
The key to this project is to reduce the computation cost of fall-related hip injury prediction using AI methods. The most effective way to do this is not known, we will investigate multiple approaches, such as training a GNN on voxel data (which completely replaces the existing simulation pipeline) and using automatic differentiation to accelerate the existing simulation.
We will train our models using federated learning, enabling decentralized training while preserving data privacy. This approach allows collaboration without sharing sensitive data, improving model accuracy with a diverse dataset. We aim to train a large model on the server and then scale it down for use on client devices like those in hospitals, making it efficient for hardware with limited computational resources.
To summarize, our model will enable healthcare providers and individuals to take proactive steps to reduce fall-related injury risks, improving patient outcomes and lowering healthcare costs. With a privacy-focused, data-driven approach, we aim to advance fall prevention and intervention strategies.
Contact
Metagenomic Analysis on Near-Data-Processing Platforms [Onur Mutlu]
Genome sequence analysis, which examines the DNA sequences of organisms, drives advances in many critical medical fields. Given its importance and the exponentially growing volumes of genomic data, extensive efforts target acceleration of genome analysis. In this work, we demonstrate a major bottleneck that greatly limits the benefits of state-of-the-art genome sequence analysis accelerators: the data preparation bottleneck, where genomic sequence data is stored in compressed form and needs to be first decompressed and formatted before an accelerator can operate on it.
To mitigate this bottleneck, we propose SAGe, an algorithm-architecture co-design for highly-compressed storage and high-performance access of large-scale genomic sequence data. The key challenge is to improve data preparation performance while maintaining high compression ratios at low hardware cost. We address this by leveraging key properties of genomic datasets to co-design (i) a lossless (de)compression algorithm, (ii) hardware that decompresses data with lightweight operations and efficient streaming accesses and (iii) storage data layout.
Due to its lightweight design, SAGe can be seamlessly integrated with a broad range of hardware accelerators, such as an in-storage NDP genome analysis system on the SSD controller. Our results demonstrate that SAGe improves the average end-to-end performance and energy efficiency of two state-of-the-art genome sequence analysis accelerators by 3.0×–32.1× and 13.0×–34.0×, respectively, compared to when the accelerators rely on state-of-the-art software and hardware decompression tools.
Key achievements
Our initial research in the domain has resulted in the following contributions:
- Systematic identification of a critical bottleneck in genome analysis pipelines. Through detailed profiling and end-to-end analysis of state-of-the-art genome sequence analysis workflows, we identify the data preparation stage as a major and previously underexplored bottleneck that limits the achievable performance and energy efficiency of modern hardware accelerators.
- Hardware/software co-design for mitigating data movement and preparation overheads. Moving towards near-data processing (NDP) systems, we introduce a principled algorithm-architecture co-design approach that rethinks how genomic data is stored, accessed, and prepared for computation, minimizing costly data movement and enabling efficient streaming-based processing.
- Design of SAGe: a lightweight, high-performance data access framework. We develop SAGe, a versatile framework that combines compression, data layout, hardware support, and interface design to enable efficient handling of large-scale genomic datasets. SAGe maintains high compression ratios while enabling low-cost, high-throughput data preparation compatible with accelerator-friendly and NDP-style execution.