Atlas versions
The model card, data availability, and the protocol for releasing a new atlas version.
Current release
Each answer is given at the level the evidence supports: one of the reference's
blood cell types, its group, or its lineage. Where
the evidence runs out, it abstains and says why. Full detail and sources:
service/model/MODEL_CARD.md.
NB2's evaluation describes its rule before the two conservative flags the service adds (a non-empty set always contains the best guess; a restricted request is renormalised), which are not yet evaluated on RNA. The development datasets are measured as served.
Releases
The reference moved from CrossModalNet, a single model jointly
trained on RNA and proteomics together, to a frozen RNA-only reference
with a separately trained query encoder. This was a structural change,
not a metric update: the joint model could implicitly see SCoPE2 during
training, which is part of why its zero-shot numbers below read higher than the
honestly separated architecture's do. Both sets of numbers are real; they answer
different questions, and the lower number is the trustworthy one, not a
regression.
The pca_centroid_cosine column in the diagnostics
table has dropped twice, for one reason. The first model was trained on both
modalities, so its RNA and protein centroids coincided. The RNA-only references never
saw protein, so theirs don't. The column now measures v3.1's space. Read the current
values in that table; this page does not copy manifest numbers into static text.
The architecture was decided by a five-seed comparison over the widened
-gene feature space (module-pooling encoder, uniform masking; see the v3
card below for the settled details), and the reference is trained with it:
reference_model.pt, reference_embedding.npy, and
reference_centroids.npy are real, live artifacts. The projection
pipeline runs against them end to end. It is not yet hosted:
the projection page sends an upload to a service you run locally.
v3, the previous release
The first model, kept rather than erased
v3 model card
The v3 reference encoder is trained, supervised, on RNA only:
cells across
immune and blood cell classes. It is applied
zero-shot to proteomics uploads: no protein labels and no
paired RNA-protein cells were used at any point in training. Full detail and
sources: service/model/MODEL_CARD.md.
Each row is scored under one decision rule. Nearest centroid is what the service runs; shared kNN is the rule every benchmark method is scored with. Compare a shipped-checkpoint number with a 5-seed mean only under the same rule.
A previously published RNA to RNA figure of is superseded by the test-cells-only above: the published number's test split included cells the model had trained on.
The restricted is
v3_seed0's own number under the service's rule, and is exact for
SCoPE2 by construction, on two separate counts. First: SCoPE2's protein
cells are only ever macrophage or monocyte, so restricting to exactly those two
classes matches its true label set on this one dataset. On other data with other
cell types present, the same restriction is wrong (see the first failure mode).
Second: v3_seed0, the checkpoint v3 serves when selected,
is the best of the independently trained seeds of the same architecture by a
wide margin: against a
5-seed mean of under the same
rule (range ).
Against scArches/scANVI, the strongest available cross-modal baseline, on SCoPE2 under the shared kNN rule:
- Restricted to macrophage and monocyte, scANVI leads v3. v3's
five-seed mean balanced accuracy is
(
research/notebook-outputs/nb1d/ours_scope2_5seed_family_summary.csv), and only v3-seed and scANVI-seed pairings favour v3. - Across all classes, v3 leads. pairings favour v3, by a mean of .
- The shipped checkpoint,
v3_seed0, is ahead of each of the three scANVI seeds in both regimes.
The pairings come from
research/notebook-outputs/nb1d/paired_bootstrap_ours_vs_scanvi.csv, with the
narrative in research/benchmark/results.md. One methodological asymmetry
favours scANVI throughout: scANVI trains on the query cells (transductive), while this
encoder is fixed and zero-shot on the query.
Known failure modes:
- Two-class label space, now opt-in: when a request restricts labels to the supported classes, a non-myeloid protein upload (T cell, NK cell, B cell, and so on) can only ever be labelled macrophage, monocyte, or abstain, never its true label. Under v3.1, the default, lymphoid uploads get lymphoid answers (see the model card).
- Macrophage placement: cross-modal alignment is weak for macrophage specifically (latent centroid cosine against for monocyte), and across all classes most macrophage protein cells are assigned to another class (see the model card).
- Only of reference classes (macrophage, monocyte) have ever been validated against real protein data.
- No donor-level holdout in training: every donor contributed a large share of its own cells to the training set (see the model card).
- High seed-to-seed variance on real cross-modal transfer, not yet explained:
retraining the same architecture moves SCoPE2 restricted balanced accuracy
by up to
( across
seeds, shared kNN); the shipped seed happens
to be the best of them, not a principled choice. A separate development-only
check on real PBMC240 data (lineage level) shows the same pattern; see
service/model/MODEL_CARD.mdfor the full detail.
Data availability
Release protocol
- Freeze inputs: record accession numbers, download dates, and checksums for every dataset.
- Train and export, then regenerate the manifest with
python scripts/build_manifest.py. - Bump
ATLAS_VERSIONinscripts/build_manifest.pyand commit the regenerated manifest. - Tag the release and archive it for a persistent identifier.
- Record label stability against the previous version and add it to this page.
This is the first release with a predecessor, and its label stability against is not yet measured. v3.1 gives its answers at class, group or lineage level, and v3 at class level only, so the measure needs a definition first.