A Google-like search engine for single-cell RNA
Imagine doctors had the potential to understand exactly what cells caused a patient’s cancer, or whether pathogens contributed to the disease. They could then use the information to tailor a treatment plan to the patient’s specific cancer. But answering such questions would mean wading through data from thousands of experiments locked in massive databases around the globe. Moreover, the search would take several days at the very least.
Now, researchers at the Berlin Institute of Medical Systems Biology of the Max Delbrück Center (MDC-BIMSB) present a search engine that radically simplifies such tasks: “Malva.” It is the first platform that can quickly sort through massive single-cell data using sequence information only, explains Daniel León-Periñán, first author of the study in “Nature.” León-Periñán is a doctoral student in the Systems Biology of Gene Regulatory Elements lab of Dr. Nikolaus Rajewsky, Director of MDC-BIMSB.
“Like Google did for the internet 30 years ago, Malva allows scientists and AI tools to search across millions of cells in seconds — without downloading huge files or needing a reference genome, and without deep computational expertise,” adds Rajewsky, senior author of the paper. “Malva transforms static transcriptomic atlases into dynamic resources, which will further our understanding of RNA biology. It will also be potentially transformative in helping researchers understand how health slides into disease, or how and which cells respond to specific medical treatments.”
The need for a tool to search RNA data
Single-cell RNA sequencing gives researchers a remarkably detailed view of what is happening inside individual cells at any given point in time. Over the past decade, researchers around the world have amassed terabytes of data. But anyone wishing to mine it to learn more about a DNA or RNA sequence of interest faces multiple hurdles. They would need to download and reprocess petabytes of raw files, which no single lab has the capacity to do, and figure out how to standardize data from different sources.
What’s more, because of the way the data is indexed, information about RNA isoforms — multiple RNA variants coded by the same gene — is extremely limited.
Daniel León-Periñán, Nikos Karaiskos, Nikolaus Rajewsky, creators of Malva
To simplify the task and to expand the types of questions the data can answer, León-Periñán and co-first author Dr. Nikos Karaiskos, also from the Rajewsky lab, have reprocessed data from public repositories and made it searchable by nucleotide sequence. Additionally, Malva indexes spatial data, so researchers can also find where in a tissue section a particular RNA is located.
Malva is designed to expand continuously. Every time new single-cell RNA data becomes available in the literature, it gets downloaded to a server located at the Max Delbrück Center, processed and added to existing data.
Broad application
There are innumerable use-case scenarios, says Karaiskos. “They can range from the very simple, like: ‘In what cell type is this particular gene expressed?’ to much more complex.”
Other platforms can also be used to answer simple questions, Karaiskos adds. But they can’t, for example, answer questions about RNA biology. This is because RNA isoforms are not each indexed to a reference gene individually, but rather treated as a single entity. This makes it impossible to distinguish whether any particular RNA isoform is expressed in any specific cell type, he says. “Having the flexibility to search by RNA sequence in Malva gives us the ability to answer questions from this data that were previously not answerable.”
Moreover, because these platforms map their data to a reference genome, which is a composite of a few individual human DNA samples, anyone looking for information about RNA produced from rare or unique gene variants may not find much. Malva, on the other hand, enables access to a much more diverse RNA dataset.
The Malva platform also includes sequence information from multiple species, including bacteria, viruses and fungi. Thus, researchers can study how these microorganisms affect human cells and cause disease.
The platform is currently freely available to scientists. The three researchers are in the early phases of launching a start-up to make Malva commercially available. A patent on the technology is pending.
Text: Gunjan Sinha
Further information
Literature
Daniel León-Periñán, Nikos Karaiskos, Nikolaus Rajewsky (2026): “Ultrafast and reference-free sequence discovery in single-cell data.” Nature. DOI: 10.1038/s41586-026 – 10975‑w
Contacts
Dr. Nikolaus Rajewsky
Director, Berlin Institute of Medical Systems Biology of the Max Delbrück Center (MDC-BIMSB)
Group Leader, Gene Regulatory Elements lab
+49 30 9406 2999 und +49 30 9406 3068 (Office)
rajewsky@mdc-berlin.de
Dr. Nikos Karaiskos
Scientist, Gene Regulatory Elements lab
Berlin Institute of Medical Systems Biology of the Max Delbrück Center
Nikolaos.Karaiskos@mdc-berlin.de
Dr. Daniel León-Periñán
Scientist, Gene Regulatory Elements lab
Berlin Institute of Medical Systems Biology of the Max Delbrück Center
Daniel.LeonPerinan@mdc-berlin.de
Gunjan Sinha
Editor, Communications
Max Delbrück Center
+49 30 9406 – 2118
presse@mdc-berlin.de
- Max Delbrück Center
The Max Delbrück Center for Molecular Medicine in the Helmholtz Association lays the foundation for the medicine of tomorrow through today’s discoveries. At locations in Berlin-Buch, Berlin-Mitte, Heidelberg, and Mannheim, interdisciplinary teams investigate the complexity of disease at the systems level – from molecules and cells to organs and entire organisms. Together with academic, clinical, and industry partners, and as part of global networks, we turn biological insights into innovations for early detection, personalized therapies, and disease prevention. Founded in 1992, the Max Delbrück Center is home to a vibrant, international research community of around 1,800 people from over 70 countries. We are 90 percent funded by the German federal government and 10 percent by the state of Berlin.