2026年8月26日,德国柏林亥姆霍兹协会Max Delbrück分子医学中心柏林医疗系统生物学研究所调控元件系统生物学实验室Nikolaus Rajewsky等科学家在《自然》发表研究,推出了Malva平台,使单细胞数据中超快且无需参考的序列发现成为可能。
RNA序列、表达、剪接、异构体、结构和修饰的知识对于理解和靶向细胞过程至关重要。革命性的单细胞和空间转录组学技术——例如人类细胞图谱等联盟所部署的技术——部分捕获了这种多样性,并生成了每年以拍字节规模扩展的细胞图谱。然而,研究人员无法在这些数据集中搜索序列:标准流程无法扩展或依赖于参考,仅保留基因或异构体计数,而访问原始序列则需要收集、下载和处理数百万个大型文件。在此,研究人员介绍了Malva,一个能够在原始序列空间中实现超快、物种非依赖且无需参考的检索的计算平台,可搜索任何序列、突变、剪接连接点或病原体,或任意转录本的空间位置。不断扩展的Malva索引目前包含来自健康和疾病中数千个实验的约7400万个细胞。Malva实现了无需参考的发现——例如,研究人员可以直接从序列组成中识别细胞类型并预测细胞-细胞相似性。基于Malva的速度和准确性,研究人员展示了如何将Malva灵活地连接到最先进的神经网络,以及如何执行复杂搜索和实现自动化分析。Malva将单细胞图谱从静态基因计数表转变为动态的、序列解析的资源,可能有助于弥合关于生物学的人机推理。
附:英文原文
Title: Ultrafast and reference-free sequence discovery in single-cell data
Author: Len-Perin, Daniel, Karaiskos, Nikos, Rajewsky, Nikolaus
Issue&Volume: 2026-08-26
Abstract: Knowledge of RNA sequences, expression, splicing, isoforms, structure and modifications is central for understanding and targeting cellular processes. Revolutionary single-cell and spatial transcriptomics technologies—for example, as deployed by consortia such as the Human Cell Atlas—partially capture this diversity and generate cellular profiles that expand at petabyte scale each year. Yet researchers cannot search sequences across these datasets: standard pipelines do not scale or rely on references, retaining only gene or isoform counts, whereas accessing raw sequences requires collecting, downloading and processing millions of large files. Here we present Malva, a computational platform that enables ultrafast, species-agnostic and reference-free interrogation of the raw sequence space, enabling searching for any sequence, mutation, splice junction or pathogen, or spatial location of arbitrary transcripts. The continuously expanding Malva Index currently comprises around 74 million cells from thousands of experiments in health and disease. Malva enables reference-free discovery—researchers can, for example, identify cell types and predict cell–cell similarity directly from sequence composition. Building on Malva’s speed and accuracy, we demonstrate how Malva can be flexibly connected to state-of-the-art neural networks and how to execute complex searches and enable automated analyses. Malva transforms single-cell atlases from static gene count tables into dynamic, sequence-resolved resources that may help to bridge human–machine reasoning about biology.
DOI: 10.1038/s41586-026-10975-w
Source: https://www.nature.com/articles/s41586-026-10975-w
Nature:《自然》,创刊于1869年。隶属于施普林格·自然出版集团,最新IF:69.504
官方网址:http://www.nature.com/
投稿链接:http://www.nature.com/authors/submit_manuscript.html
