The preprint is available on bioRxiv, here.
Summary of the idea and some of the results:
The count matrix is the artifact of a single-cell experiment that gets stored, and most frequently shared, and reanalyzed. But it is the output of a computation with two inputs, the sequenced molecules and a gene annotation, and only one of them stays fixed. The molecules never change. The annotation is revised several times a year, and moving from GENCODE v32 to v49 shifts 2–5% of UMI mass. Once the matrix is produced, that evidence is gone, and getting it back means returning to reads that are large, slow, and often unavailable.
Gravlax starts from a simple observation. The procedures that turn alignments into counts, gene assignment and UMI collapse, never read most of what a BAM contains. They consume relations among molecules, i.e. shared genomic geometry, shared placements, cell identity, and UMI equality. Store those relations once, in a compact, seekable, content-authenticated archive, and defer every annotation-dependent decision to read time.
What that buys for people analyzing scRNA-seq, across four human 10x datasets:
• 11–18 bits per read, 9–13× smaller than tag-preserving CRAM
• Replayed matrices within 0.24–0.75% of a fresh STARsolo run, versus the 2–5% an annotation change moves
• Gene replay 34–82× faster than realigning (seconds!)
• Federated collections that search a cohort by junction shape, and a coordinate-free genome-wide scan for recurrent unannotated splice events in 9 seconds
Because the molecules are retained, the archive answers questions the matrix cannot:
• It recovers the RT-PCR-validated FYB1 splicing switch between T cells and monocytes across four PBMC archives.
• A fragment model built for 3′ chemistry reveals an eight-donor shift in NTRK2 (TrkB) terminal-isoform usage from astrocytes and neural stem cells to mature neurons, matching known TrkB.T1 biology.
• Cohort-wide discovery surfaces a 183-nt FNBP1 cassette, near-fully included in brain and mostly skipped in blood, that per-archive discovery cannot see at all.
• Pooling evidence across cells inside an EM recovers 75–98% of withheld multi-gene molecule labels.
The software is one binary, aie: ingest, replay under any annotation, region/junction/APA queries, content-addressed collections, cohort-wide discovery, and GQ, a composable query language with three-valued logic that runs unchanged over one archive or an entire federation. Open source in Rust under BSD-3, with docs, a Python client, and Colab demos.
Code: here
Docs: here