Skip to content
VIRDX

Histo domain representations

Find, load and read the per-slide domain representations the histo pipeline publishes, and run the pipeline yourself.

A domain representation is what the histo pipeline produces for one H&E whole-slide image: five label maps that describe the slide down to single cells. This page is for teams that want to use them.

There are three ways in:

  1. Read the published outputs from the data platform. This is the normal route. You need vxdata-sdk and zarr, but not the histo package.
  2. Production runs on the cluster, which publish new outputs to the platform.
  3. Local runs, which write outputs to shared storage for your own experiments.

1. Read published outputs

What is published

Each processed scan has one HistoDomainRep bundle in the histo_domain_reps namespace. The bundle points to five HistoMaps.

Bundle field HistoMap task Map
tissue_segmentation TISSUE_SEGMENTATION tissue_seg.zarr
cell_segmentation CELL_SEGMENTATION cell_seg.zarr
nuclei_segmentation NUCLEI_SEGMENTATION nuclei_seg.zarr
gleason_segmentation GLEASON_SEGMENTATION gleason_seg.zarr
structure_tensor STRUCTURE_TENSOR structure_tensor.zarr

Identifiers carry the pipeline version, so several versions can exist side by side:

histodomainrep/<scan path>/v2.0
histomap/<scan path>/tissue_seg-v2.0

<scan path> is the source HistoScan identifier without the histoscan/ prefix. The bundle’s parent_identifier is that HistoScan, so you can find the H&E the maps were made from.

As of 2026-09-25 there are 8,474 v2.0 bundles: histai 4,230, essen01 1,865, frankfurt 893, essen02 818, aggc2022 387, chimera 281.

Load one slide

from pathlib import Path

import zarr
from vxdata.sdk import Client, F

client = Client()

# Pick bundles: by dataset, by scan, or by version.
reps = (
    client.histo_domain_reps.query()
    .filter(F.datasource_id == "datasource/essen01", F.histo_version == "v2.0")
    .collect()
)
rep = reps.row(0, named=True)

# Resolve the five maps to storage urls.
fields = ["tissue_segmentation", "cell_segmentation", "nuclei_segmentation",
          "gleason_segmentation", "structure_tensor"]
maps = client.histo_maps.query().filter(F.identifier.is_in([rep[f] for f in fields])).collect()
urls = dict(zip(maps["task"], maps["url"]))

# Download the maps you need (each url is a zarr directory) and open them.
dest = Path("domain_reps")
tissue_path = Path(client.storage.download(urls["TISSUE_SEGMENTATION"], dest))
cell_path = Path(client.storage.download(urls["CELL_SEGMENTATION"], dest))

tissue_group = zarr.open_group(tissue_path, mode="r")
cell_group = zarr.open_group(cell_path, mode="r")
tissue_seg = tissue_group["scale_1"]          # (H, W) uint8
cell_seg = cell_group["scale_1"]              # (H, W) uint32
cell_lut = cell_group["cell_lut"][:]          # (n_cells, 3) uint32
classes = tissue_group.attrs["virdx_attrs"]["misc"]  # class names, ids, colours

The H&E comes from the parent scan: client.histo_scans.query() .filter(F.identifier == rep["parent_identifier"]). Its url is the preprocessed pyramidal TIFF.

The maps

Every map is an OME-Zarr (zarr v3) group. The image is the array scale_1; there is no pyramid. cell_seg.zarr also holds the cell_lut array. The four label maps share one pixel grid: the H&E pyramid level closest to 10x (about 0.97 µm/px). The structure tensor is on a 4x coarser grid. The exact µm/px of each map is the scale in its OME metadata and the affine in virdx_attrs.

Map Shape dtype Pixel value
tissue_seg (H, W) uint8 tissue class, with Gleason grade on epithelium
cell_seg (H, W) uint32 cell id on cell pixels, tissue class elsewhere
nuclei_seg (H, W) uint32 nucleus id (= its cell id) on nucleus pixels, tissue class elsewhere
gleason_seg (H, W) uint8 0 background, 1 benign, 3/4/5 Gleason grade
structure_tensor (⌈H/4⌉, ⌈W/4⌉, 2) uint8 fibre orientation and coherence, 4x subsampled

One id space

tissue_seg, cell_seg and nuclei_seg use the same id space:

Ids Meaning
0–27 base tissue classes (the fine28 taxonomy; 0 = background)
128–139 Gleason-graded epithelium, e.g. 132 = luminal_epithelium_g4
255 unknown
≥ 256 cell instance ids (only in cell_seg and nuclei_seg)

Do not hard-code the class table. tissue_seg, cell_seg and nuclei_seg store it in attrs["virdx_attrs"]["misc"]: region_ids, region_classes, region_palette_rgb and id_space. The same table is in their HistoMap’s labels column. gleason_seg stores its own five-class table in the same place, without id_space. The source config is config/domain_representation/taxonomy/fine28.yaml.

Cells

cell_seg.zarr/cell_lut has one row per cell: [cell_id, base_class, gleason_grade]. base_class is a 0–27 id. gleason_grade is 3, 4 or 5 for graded epithelium cells and 0 for all others.

The maps agree by construction:

  • On cell pixels, tissue_seg is the cell’s class, graded if the cell has a grade.
  • Off cell pixels, cell_seg == tissue_seg.
  • Every nucleus id in nuclei_seg is also a cell id in cell_seg.
  • Every cell id in the map has a LUT row. A few LUT rows (about 0.3 % on a test slide) have no pixels, so join from the map to the LUT, not the reverse.

Only these classes hold cells: epithelia (1, 2, 22, 23), stroma (3, 4), inflammation, nerve, vessels and arteries. Lumen, fat, blood and artefact classes have none.

Structure tensor

Channel 0 is the dominant gradient orientation in radians, in [-π/2, π/2]. The fibres run perpendicular to it. Channel 1 is coherence in [0, 1]. Decode with value * scale + offset. The scale and offset are in attrs["virdx_attrs"]["misc"]["structure_tensor_encoding"]. Pixel (r, c) samples full-grid pixel (4r, 4c). With histo installed, histo.domain_representation.representation.decode_structure_tensor(array) does the decode.

Provenance and versions

Each bundle records:

  • histo_version: the pipeline version (v2.0). It changes whenever the outputs change meaning.
  • tissue_model_id, gleason_model_id, cell_model_id: the ClearML model ids.
  • generation_parameters: the full pipeline config, including the Docker image tag and the taxonomy.
  • generation_outputs: µm/px (mpp), cell count, timings, and urls of the run log (run.json) and event log (events.jsonl).

A bundle is written only after its five maps are uploaded. If a bundle exists, the scan is complete.

Size and practical notes

  • Biopsy slides are small (for example 27k × 6k px). Frankfurt whole-mounts are about 47k × 47k px. There, tissue_seg is about 100 MB and cell_seg several times more. Read windows (tissue_seg[y0:y1, x0:x1]) instead of whole arrays when you can.
  • Label map chunks are 1024 × 1024 px; structure tensor chunks are 256 × 256. All are compressed with zstd.
  • storage.download(url, dest) keeps the key path under dest, without the bucket name.

2. Production runs: publish to the data platform

A production run processes HistoScans on the cluster and registers the results on the platform, as described in section 1. The cluster already has everything the run needs: the image holds the code, the environment and the config, the job mounts the ClearML, S3 and vxData credentials from cluster secrets, and the backbone weights come from shared storage. You need argo access to the argo-workflows namespace and a histo checkout for the job file.

Test on staging first. run_environment=staging (the default) registers the bundles on the staging platform and uploads to the domrep_v2_testing group. run_environment=production writes the official v2.0 bundles that other teams read (upload group domainrep_v2.0). Agree production runs with the histo team.

argo submit -n argo-workflows jobs/run-domain-representation-v2.yaml \
  -p image=histo -p image_tag=main-088def2 -p username=<you> \
  -p run_environment=staging -p datasource=essen01

main-088def2 is the image that made most v2.0 bundles. It uses the same models as main.

Parameter What it sets
datasource Process every preprocessed scan of this datasource, e.g. essen01.
scans_file Instead of datasource: a JSON list of HistoScan identifiers, under /mnt/artifacts so the pods can read it.
run_environment staging (default) or production.
gpu_type Pin the GPU: NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition or NVIDIA-RTX-6000-Ada-Generation. Empty takes either.
worklist_mpx Slide area per GPU pod, in megapixels at 10x. The default 40000 takes about 2 h on Blackwell.
overrides Extra Hydra overrides for the pipeline, e.g. output.worklist_log_dir=/mnt/artifacts/<you>/logs to keep the per-pod logs.

What happens:

  • A list step finds the scans that have no bundle with this pipeline version and these model ids, and packs them into worklists. Resubmitting the same job only processes what is left.
  • One GPU pod (128 Gi) per worklist loads the models once and processes its scans one by one. A failed scan does not stop the others; the pod then exits with code 13.
  • Inputs are preprocessed HistoScans. A scan’s FOREGROUND_MASK map is used when it has one; without it every window is embedded, which is slower.

The job runs in your own Kueue queue (user-yourname). To also use idle GPUs elsewhere on the cluster, add -l kueue.x-k8s.io/queue-name=preemptible,virdx.dev/preemptible=true to argo submit. Those pods can be stopped and rerun, which is safe because done scans are skipped.

Watch the run with argo list -n argo-workflows, argo get -n argo-workflows <workflow> and argo logs -n argo-workflows <workflow>. Finished scans appear as bundles on the platform.

3. Local runs: outputs for your own experiments

A local run writes the same five maps to shared storage and never touches the platform. Use it to try other settings, or slides that are not on the platform.

You need:

  • A GPU. The coder pod’s 24 GB GPU slice works with the local profiles.
  • A histo checkout with its environment: pixi install --manifest-path packages/histo/pyproject.toml -e cuda-dev.
  • ClearML credentials in ~/clearml.conf. The three heads are downloaded from clearml.fra.virdx.dev on the first run and cached in ~/.clearml/cache.
  • /mnt/artifacts mounted. The backbone weights are read from there, so no Hugging Face token is needed.
  • A vxData token only if your inputs are platform scans.

Run from the repo root:

SSL_CERT_FILE=/etc/ssl/certs/ca-certificates.crt \
REQUESTS_CA_BUNDLE=/etc/ssl/certs/ca-certificates.crt \
pixi run --manifest-path packages/histo/pyproject.toml -e cuda-dev \
  python scripts/domain_representation/generate_domain_representation.py \
  compute=local 'inputs.paths=[/path/to/slide.tiff]'

The two certificate variables let Python trust the internal CA of clearml.fra.virdx.dev. On the coder pod the system bundle already contains it.

Inputs, one of:

  • inputs.paths=[...]: preprocessed pyramidal TIFFs. If <name>_tissue_mask.ome.zarr sits next to <name>.tiff, only tissue is embedded.
  • inputs.list_file=slides.json: a JSON list of {"path": ..., "identifier": ...}. identifier is optional and names the output folder.
  • inputs.scans=[histoscan/...]: platform scans, downloaded first.

inputs.limit=N processes only the first N slides.

Profiles: compute=local (fits the coder pod) and compute=local_small (smaller windows, for quick checks).

Each slide gets a folder /mnt/artifacts/$USER/histo/domain_representation_v2/<slide>/ with the five maps (tissue_seg.zarr, cell_seg.zarr, …), run.json (settings, models, timings) and events.jsonl. Set output.scratch_dir=... to write elsewhere. Open the maps with zarr as in section 1; their layout is the same.

A slide whose folder already has a run.json is skipped. Delete the folder, or use a new output.scratch_dir, to compute it again.

Any setting in config/domain_representation/generate.yaml can be changed on the command line, for example watershed.enabled=false. Settings are checked on load, and a bad combination stops the run with a message.

Python API

The same run from Python. The public API is what histo.domain_representation exports:

Name What it does
load_settings(cfg) -> PipelineSettings Types and checks the composed Hydra config.
load_models(settings, device=None) -> Models Loads H-optimus-0 and the tissue, Gleason and cell heads.
process_scans(scans, settings, models, *, platform) -> ScansReport Processes a list of slides. Skips slides already done, and one failure does not stop the rest. report.done, report.failed.
process_scan(scan, settings, models, *, output_dir, ...) Processes one slide into output_dir.
local_scans(settings.inputs) -> list[ScanInput] Slides from local paths or a JSON list file.
ScanInput, PlatformScan One slide on disk / one HistoScan on the platform.
Platform Resolves scans and publishes outputs. Pass platform=None to write locally only.
from hydra import compose, initialize_config_dir
from histo.domain_representation import load_models, load_settings, local_scans, process_scans

with initialize_config_dir(version_base=None, config_dir="/abs/path/to/histo/config/domain_representation"):
    cfg = compose("generate", overrides=["compute=local", "inputs.paths=[/data/slide.tiff]"])
settings = load_settings(cfg)
report = process_scans(local_scans(settings.inputs), settings, load_models(settings), platform=None)

The config is in the repo, not in the histo package, so this also needs a histo checkout. The other modules of histo are internal; their interfaces may change without notice.

Contact

Ask the histo team before you depend on anything not listed on this page.

Navigation

Type to search…

↑↓ navigate↵ selectEsc close