Three decisions inside the index
Start with what was built. Malva Index was compiled from more than 140 TB of public single-cell data in about 70 h of wall time, around 9,700 h of CPU time, and in the human index the count of unique 24-mers sits above 100 billion, drawn from around 51 million cells, 592 studies and 7,966 samples, with about 10 million mouse cells alongside. On disk the whole thing takes less than 10 per cent of the storage its source files occupy. That compression ratio is the first clue about where the speed comes from: the index throws away everything except which cell carried which short subsequence.[1]
Three choices do the work. Non-overlapping k-mer sampling removes the de Bruijn graph construction that other k-mer indices need. Per-k-mer sparse inverted lists store only the identifiers of the cells that carry a given k-mer, so the cost of a query stops tracking the total size of the index. Independent per-sample indexing with incremental merging keeps the layout on disk, which is what lets the corpus keep growing without a rebuild. The payoff is a latency figure bounded by disk access: 70 ms for a single k-mer, 0.9 s for a 1 kb transcript and about 1 min for 1,000 transcripts on a single CPU core.[1]
What each decision costs
Every one of those choices is paid for somewhere, and the paper is unusually direct about where. The counts Malva returns come straight out of k-mer matches with no alignment behind them and no correction for unique molecular identifiers, so they track read counts rather than molecule counts; the authors call them semi-quantitative and mean it. A library-specific normalization based on sequencing saturation pulls out the mean inflation that amplification adds, but it does not touch gene- or sequence-specific amplification biases. Anyone reaching for these numbers as expression measurements is using the wrong instrument.[1]
Two more prices sit further down. Barcode error correction and doublet removal are omitted to keep processing throughput up at atlas scale, which pushes that cleanup onto whoever interprets the result. And because k-mers are sampled without overlap, sensitivity for a short probe of length k depends on whether it lands on an indexed position; the sliding window query algorithm guarantees detection only for probes longer than k. Exact matching also excludes the Hamming distance 1 mismatches a tool like Flexiplex permits, which is a deliberate trade of tolerance for speed rather than an oversight.[1]
The headline indexing figure reads the same way once the trade is visible. Malva took 24 CPU hours and an 8 GB peak to index a Stereo-seq mouse liver atlas of about 61 billion reads across 20 sections, between 4 and 25 times faster than the other preprocessing algorithms. That margin looks like the consequence of moving annotation off memory and onto disk more than of a faster matching step, since the memory ceiling is exactly what the sparse inverted lists remove. One competing reading deserves saying out loud: the comparison methods were benchmarked without memory constraints, so part of the gap may reflect how each tool was configured rather than the data structure itself.[1]
Where the index stops
Coverage is the honest boundary. As of May 2025 the index held around 80 per cent of the Human Cell Atlas but around 20 per cent of the human single-cell runs in the SRA, and the team says plainly that Malva does not replace reference-based quantification pipelines, which remain suitable for gene-level analysis. Read together, those two sentences describe a tool that answers a question no existing pipeline could answer at this scale, over a slice of the public data that is large in one archive and thin in another.[1]
Two numbers are the ones to check again as the corpus grows. The first is that SRA share, because a crawler that keeps pace with new submissions is a different resource from one that falls behind them. The second is how far the pseudocount correlations hold as protocols the benchmark did not include enter the index; the paper's own robustness test keeps sensitivity, specificity and those correlations stable up to a 2 per cent simulated error rate and lets them degrade gradually toward 10 per cent. Both are measurable, both are published, and both are the sort of thing an index has to keep proving rather than claim once.[1]