Atlas versions

The model card, data availability, and the protocol for releasing a new atlas version.

Current release

Each answer is given at the level the evidence supports: one of the reference's   blood cell types, its group, or its lineage. Where the evidence runs out, it abstains and says why. Full detail and sources: service/model/MODEL_CARD.md.

NB2's evaluation describes its rule before the two conservative flags the service adds (a non-empty set always contains the best guess; a restricted request is renormalised), which are not yet evaluated on RNA. The development datasets are measured as served.

Releases

The reference moved from CrossModalNet, a single model jointly trained on RNA and proteomics together, to a frozen RNA-only reference with a separately trained query encoder. This was a structural change, not a metric update: the joint model could implicitly see SCoPE2 during training, which is part of why its zero-shot numbers below read higher than the honestly separated architecture's do. Both sets of numbers are real; they answer different questions, and the lower number is the trustworthy one, not a regression.

The pca_centroid_cosine column in the diagnostics table has dropped twice, for one reason. The first model was trained on both modalities, so its RNA and protein centroids coincided. The RNA-only references never saw protein, so theirs don't. The column now measures v3.1's space. Read the current values in that table; this page does not copy manifest numbers into static text.

The architecture was decided by a five-seed comparison over the widened  -gene feature space (module-pooling encoder, uniform masking; see the v3 card below for the settled details), and the reference is trained with it: reference_model.pt, reference_embedding.npy, and reference_centroids.npy are real, live artifacts. The projection pipeline runs against them end to end. It is not yet hosted: the projection page sends an upload to a service you run locally.

v3, the previous release

The first model, kept rather than erased

v3 model card

The v3 reference encoder is trained, supervised, on RNA only:   cells across   immune and blood cell classes. It is applied zero-shot to proteomics uploads: no protein labels and no paired RNA-protein cells were used at any point in training. Full detail and sources: service/model/MODEL_CARD.md.

Each row is scored under one decision rule. Nearest centroid is what the service runs; shared kNN is the rule every benchmark method is scored with. Compare a shipped-checkpoint number with a 5-seed mean only under the same rule.

A previously published RNA to RNA figure of   is superseded by the test-cells-only   above: the published number's test split included cells the model had trained on.

The restricted   is v3_seed0's own number under the service's rule, and is exact for SCoPE2 by construction, on two separate counts. First: SCoPE2's protein cells are only ever macrophage or monocyte, so restricting to exactly those two classes matches its true label set on this one dataset. On other data with other cell types present, the same restriction is wrong (see the first failure mode). Second: v3_seed0, the checkpoint v3 serves when selected, is the best of the independently trained seeds of the same architecture by a wide margin:   against a 5-seed mean of   under the same rule (range  ).

Against scArches/scANVI, the strongest available cross-modal baseline, on SCoPE2 under the shared kNN rule:

The pairings come from research/notebook-outputs/nb1d/paired_bootstrap_ours_vs_scanvi.csv, with the narrative in research/benchmark/results.md. One methodological asymmetry favours scANVI throughout: scANVI trains on the query cells (transductive), while this encoder is fixed and zero-shot on the query.

Known failure modes:

  1. Two-class label space, now opt-in: when a request restricts labels to the supported classes, a non-myeloid protein upload (T cell, NK cell, B cell, and so on) can only ever be labelled macrophage, monocyte, or abstain, never its true label. Under v3.1, the default, lymphoid uploads get lymphoid answers (see the model card).
  2. Macrophage placement: cross-modal alignment is weak for macrophage specifically (latent centroid cosine   against   for monocyte), and across all classes most macrophage protein cells are assigned to another class (see the model card).
  3. Only   of   reference classes (macrophage, monocyte) have ever been validated against real protein data.
  4. No donor-level holdout in training: every donor contributed a large share of its own cells to the training set (see the model card).
  5. High seed-to-seed variance on real cross-modal transfer, not yet explained: retraining the same architecture moves SCoPE2 restricted balanced accuracy by up to   (  across   seeds, shared kNN); the shipped seed happens to be the best of them, not a principled choice. A separate development-only check on real PBMC240 data (lineage level) shows the same pattern; see service/model/MODEL_CARD.md for the full detail.

Data availability

Release protocol

  1. Freeze inputs: record accession numbers, download dates, and checksums for every dataset.
  2. Train and export, then regenerate the manifest with python scripts/build_manifest.py.
  3. Bump ATLAS_VERSION in scripts/build_manifest.py and commit the regenerated manifest.
  4. Tag the release and archive it for a persistent identifier.
  5. Record label stability against the previous version and add it to this page.

This is the first release with a predecessor, and its label stability against   is not yet measured. v3.1 gives its answers at class, group or lineage level, and v3 at class level only, so the measure needs a definition first.