| Home > Publications database > Ultrafast and reference-free sequence discovery in single-cell data. |
| Journal Article | DKFZ-2026-02123 |
; ;
2026
Nature Publ. Group
London [u.a.]
Abstract: Knowledge of RNA sequences, expression, splicing, isoforms, structure and modifications is central for understanding and targeting cellular processes. Revolutionary single-cell and spatial transcriptomics technologies-for example, as deployed by consortia such as the Human Cell Atlas-partially capture this diversity and generate cellular profiles that expand at petabyte scale each year1-5. Yet researchers cannot search sequences across these datasets: standard pipelines do not scale or rely on references, retaining only gene or isoform counts, whereas accessing raw sequences requires collecting, downloading and processing millions of large files. Here we present Malva, a computational platform that enables ultrafast, species-agnostic and reference-free interrogation of the raw sequence space, enabling searching for any sequence, mutation, splice junction or pathogen, or spatial location of arbitrary transcripts. The continuously expanding Malva Index currently comprises around 74 million cells from thousands of experiments in health and disease. Malva enables reference-free discovery-researchers can, for example, identify cell types and predict cell-cell similarity directly from sequence composition. Building on Malva's speed and accuracy, we demonstrate how Malva can be flexibly connected to state-of-the-art neural networks and how to execute complex searches and enable automated analyses. Malva transforms single-cell atlases from static gene count tables into dynamic, sequence-resolved resources that may help to bridge human-machine reasoning about biology.
|
The record appears in these collections: |