A domain representation is what the histo pipeline produces for one H&E whole-slide image: five label maps that describe the slide down to single cells. This page is for teams that want to use them.
There are three ways in:
- Read the published outputs from the data platform. This is the normal
route. You need
vxdata-sdkandzarr, but not thehistopackage. - Production runs on the cluster, which publish new outputs to the platform.
- Local runs, which write outputs to shared storage for your own experiments.
1. Read published outputs
What is published
Each processed scan has one HistoDomainRep bundle in the histo_domain_reps
namespace. The bundle points to five HistoMaps.
| Bundle field | HistoMap task | Map |
|---|---|---|
tissue_segmentation |
TISSUE_SEGMENTATION |
tissue_seg.zarr |
cell_segmentation |
CELL_SEGMENTATION |
cell_seg.zarr |
nuclei_segmentation |
NUCLEI_SEGMENTATION |
nuclei_seg.zarr |
gleason_segmentation |
GLEASON_SEGMENTATION |
gleason_seg.zarr |
structure_tensor |
STRUCTURE_TENSOR |
structure_tensor.zarr |
Identifiers carry the pipeline version, so several versions can exist side by side:
histodomainrep/<scan path>/v2.0
histomap/<scan path>/tissue_seg-v2.0<scan path> is the source HistoScan identifier without the histoscan/
prefix. The bundle’s parent_identifier is that HistoScan, so you can find the
H&E the maps were made from.
As of 2026-09-25 there are 8,474 v2.0 bundles: histai 4,230, essen01 1,865,
frankfurt 893, essen02 818, aggc2022 387, chimera 281.
Load one slide
from pathlib import Path
import zarr
from vxdata.sdk import Client, F
client = Client()
# Pick bundles: by dataset, by scan, or by version.
reps = (
client.histo_domain_reps.query()
.filter(F.datasource_id == "datasource/essen01", F.histo_version == "v2.0")
.collect()
)
rep = reps.row(0, named=True)
# Resolve the five maps to storage urls.
fields = ["tissue_segmentation", "cell_segmentation", "nuclei_segmentation",
"gleason_segmentation", "structure_tensor"]
maps = client.histo_maps.query().filter(F.identifier.is_in([rep[f] for f in fields])).collect()
urls = dict(zip(maps["task"], maps["url"]))
# Download the maps you need (each url is a zarr directory) and open them.
dest = Path("domain_reps")
tissue_path = Path(client.storage.download(urls["TISSUE_SEGMENTATION"], dest))
cell_path = Path(client.storage.download(urls["CELL_SEGMENTATION"], dest))
tissue_group = zarr.open_group(tissue_path, mode="r")
cell_group = zarr.open_group(cell_path, mode="r")
tissue_seg = tissue_group["scale_1"] # (H, W) uint8
cell_seg = cell_group["scale_1"] # (H, W) uint32
cell_lut = cell_group["cell_lut"][:] # (n_cells, 3) uint32
classes = tissue_group.attrs["virdx_attrs"]["misc"] # class names, ids, coloursThe H&E comes from the parent scan: client.histo_scans.query() .filter(F.identifier == rep["parent_identifier"]). Its url is the
preprocessed pyramidal TIFF.
The maps
Every map is an OME-Zarr (zarr v3) group. The image is the array scale_1;
there is no pyramid. cell_seg.zarr also holds the cell_lut array. The four
label maps share one pixel grid: the H&E pyramid level closest to 10x (about
0.97 µm/px). The structure tensor is on a 4x coarser grid. The exact µm/px of
each map is the scale in its OME metadata and the affine in virdx_attrs.
| Map | Shape | dtype | Pixel value |
|---|---|---|---|
tissue_seg |
(H, W) | uint8 | tissue class, with Gleason grade on epithelium |
cell_seg |
(H, W) | uint32 | cell id on cell pixels, tissue class elsewhere |
nuclei_seg |
(H, W) | uint32 | nucleus id (= its cell id) on nucleus pixels, tissue class elsewhere |
gleason_seg |
(H, W) | uint8 | 0 background, 1 benign, 3/4/5 Gleason grade |
structure_tensor |
(⌈H/4⌉, ⌈W/4⌉, 2) | uint8 | fibre orientation and coherence, 4x subsampled |
One id space
tissue_seg, cell_seg and nuclei_seg use the same id space:
| Ids | Meaning |
|---|---|
| 0–27 | base tissue classes (the fine28 taxonomy; 0 = background) |
| 128–139 | Gleason-graded epithelium, e.g. 132 = luminal_epithelium_g4 |
| 255 | unknown |
| ≥ 256 | cell instance ids (only in cell_seg and nuclei_seg) |
Do not hard-code the class table. tissue_seg, cell_seg and nuclei_seg
store it in attrs["virdx_attrs"]["misc"]: region_ids, region_classes,
region_palette_rgb and id_space. The same table is in their HistoMap’s
labels column. gleason_seg stores its own five-class table in the same
place, without id_space. The source config is config/domain_representation/taxonomy/fine28.yaml.
Cells
cell_seg.zarr/cell_lut has one row per cell: [cell_id, base_class, gleason_grade]. base_class is a 0–27 id. gleason_grade is 3, 4 or 5 for
graded epithelium cells and 0 for all others.
The maps agree by construction:
- On cell pixels,
tissue_segis the cell’s class, graded if the cell has a grade. - Off cell pixels,
cell_seg == tissue_seg. - Every nucleus id in
nuclei_segis also a cell id incell_seg. - Every cell id in the map has a LUT row. A few LUT rows (about 0.3 % on a test slide) have no pixels, so join from the map to the LUT, not the reverse.
Only these classes hold cells: epithelia (1, 2, 22, 23), stroma (3, 4), inflammation, nerve, vessels and arteries. Lumen, fat, blood and artefact classes have none.
Structure tensor
Channel 0 is the dominant gradient orientation in radians, in [-π/2, π/2]. The
fibres run perpendicular to it. Channel 1 is coherence in [0, 1]. Decode with
value * scale + offset. The scale and offset are in
attrs["virdx_attrs"]["misc"]["structure_tensor_encoding"]. Pixel
(r, c) samples full-grid pixel (4r, 4c). With histo installed,
histo.domain_representation.representation.decode_structure_tensor(array)
does the decode.
Provenance and versions
Each bundle records:
histo_version: the pipeline version (v2.0). It changes whenever the outputs change meaning.tissue_model_id,gleason_model_id,cell_model_id: the ClearML model ids.generation_parameters: the full pipeline config, including the Docker image tag and the taxonomy.generation_outputs: µm/px (mpp), cell count, timings, and urls of the run log (run.json) and event log (events.jsonl).
A bundle is written only after its five maps are uploaded. If a bundle exists, the scan is complete.
Size and practical notes
- Biopsy slides are small (for example 27k × 6k px). Frankfurt whole-mounts are
about 47k × 47k px. There,
tissue_segis about 100 MB andcell_segseveral times more. Read windows (tissue_seg[y0:y1, x0:x1]) instead of whole arrays when you can. - Label map chunks are 1024 × 1024 px; structure tensor chunks are 256 × 256. All are compressed with zstd.
storage.download(url, dest)keeps the key path underdest, without the bucket name.
2. Production runs: publish to the data platform
A production run processes HistoScans on the cluster and registers the results
on the platform, as described in section 1. The cluster already has everything
the run needs: the image holds the code, the environment and the config, the
job mounts the ClearML, S3 and vxData credentials from cluster secrets, and the
backbone weights come from shared storage. You need argo access to the
argo-workflows namespace and a histo checkout for the job file.
Test on staging first. run_environment=staging (the default) registers
the bundles on the staging platform and uploads to the domrep_v2_testing
group. run_environment=production writes the official v2.0 bundles that
other teams read (upload group domainrep_v2.0). Agree production runs with
the histo team.
argo submit -n argo-workflows jobs/run-domain-representation-v2.yaml \
-p image=histo -p image_tag=main-088def2 -p username=<you> \
-p run_environment=staging -p datasource=essen01main-088def2 is the image that made most v2.0 bundles. It uses the same
models as main.
| Parameter | What it sets |
|---|---|
datasource |
Process every preprocessed scan of this datasource, e.g. essen01. |
scans_file |
Instead of datasource: a JSON list of HistoScan identifiers, under /mnt/artifacts so the pods can read it. |
run_environment |
staging (default) or production. |
gpu_type |
Pin the GPU: NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition or NVIDIA-RTX-6000-Ada-Generation. Empty takes either. |
worklist_mpx |
Slide area per GPU pod, in megapixels at 10x. The default 40000 takes about 2 h on Blackwell. |
overrides |
Extra Hydra overrides for the pipeline, e.g. output.worklist_log_dir=/mnt/artifacts/<you>/logs to keep the per-pod logs. |
What happens:
- A list step finds the scans that have no bundle with this pipeline version and these model ids, and packs them into worklists. Resubmitting the same job only processes what is left.
- One GPU pod (128 Gi) per worklist loads the models once and processes its scans one by one. A failed scan does not stop the others; the pod then exits with code 13.
- Inputs are preprocessed HistoScans. A scan’s
FOREGROUND_MASKmap is used when it has one; without it every window is embedded, which is slower.
The job runs in your own Kueue queue (user-yourname). To also use idle GPUs
elsewhere on the cluster, add
-l kueue.x-k8s.io/queue-name=preemptible,virdx.dev/preemptible=true to
argo submit. Those pods can be stopped and rerun, which is safe because done
scans are skipped.
Watch the run with argo list -n argo-workflows, argo get -n argo-workflows <workflow> and argo logs -n argo-workflows <workflow>. Finished scans appear
as bundles on the platform.
3. Local runs: outputs for your own experiments
A local run writes the same five maps to shared storage and never touches the platform. Use it to try other settings, or slides that are not on the platform.
You need:
- A GPU. The coder pod’s 24 GB GPU slice works with the
localprofiles. - A histo checkout with its environment:
pixi install --manifest-path packages/histo/pyproject.toml -e cuda-dev. - ClearML credentials in
~/clearml.conf. The three heads are downloaded fromclearml.fra.virdx.devon the first run and cached in~/.clearml/cache. /mnt/artifactsmounted. The backbone weights are read from there, so no Hugging Face token is needed.- A vxData token only if your inputs are platform scans.
Run from the repo root:
SSL_CERT_FILE=/etc/ssl/certs/ca-certificates.crt \
REQUESTS_CA_BUNDLE=/etc/ssl/certs/ca-certificates.crt \
pixi run --manifest-path packages/histo/pyproject.toml -e cuda-dev \
python scripts/domain_representation/generate_domain_representation.py \
compute=local 'inputs.paths=[/path/to/slide.tiff]'The two certificate variables let Python trust the internal CA of
clearml.fra.virdx.dev. On the coder pod the system bundle already contains
it.
Inputs, one of:
inputs.paths=[...]: preprocessed pyramidal TIFFs. If<name>_tissue_mask.ome.zarrsits next to<name>.tiff, only tissue is embedded.inputs.list_file=slides.json: a JSON list of{"path": ..., "identifier": ...}.identifieris optional and names the output folder.inputs.scans=[histoscan/...]: platform scans, downloaded first.
inputs.limit=N processes only the first N slides.
Profiles: compute=local (fits the coder pod) and compute=local_small
(smaller windows, for quick checks).
Each slide gets a folder
/mnt/artifacts/$USER/histo/domain_representation_v2/<slide>/ with the five
maps (tissue_seg.zarr, cell_seg.zarr, …), run.json (settings, models,
timings) and events.jsonl. Set output.scratch_dir=... to write elsewhere.
Open the maps with zarr as in section 1; their layout is the same.
A slide whose folder already has a run.json is skipped. Delete the folder, or
use a new output.scratch_dir, to compute it again.
Any setting in config/domain_representation/generate.yaml can be changed on
the command line, for example watershed.enabled=false. Settings are checked
on load, and a bad combination stops the run with a message.
Python API
The same run from Python. The public API is what histo.domain_representation
exports:
| Name | What it does |
|---|---|
load_settings(cfg) -> PipelineSettings |
Types and checks the composed Hydra config. |
load_models(settings, device=None) -> Models |
Loads H-optimus-0 and the tissue, Gleason and cell heads. |
process_scans(scans, settings, models, *, platform) -> ScansReport |
Processes a list of slides. Skips slides already done, and one failure does not stop the rest. report.done, report.failed. |
process_scan(scan, settings, models, *, output_dir, ...) |
Processes one slide into output_dir. |
local_scans(settings.inputs) -> list[ScanInput] |
Slides from local paths or a JSON list file. |
ScanInput, PlatformScan |
One slide on disk / one HistoScan on the platform. |
Platform |
Resolves scans and publishes outputs. Pass platform=None to write locally only. |
from hydra import compose, initialize_config_dir
from histo.domain_representation import load_models, load_settings, local_scans, process_scans
with initialize_config_dir(version_base=None, config_dir="/abs/path/to/histo/config/domain_representation"):
cfg = compose("generate", overrides=["compute=local", "inputs.paths=[/data/slide.tiff]"])
settings = load_settings(cfg)
report = process_scans(local_scans(settings.inputs), settings, load_models(settings), platform=None)The config is in the repo, not in the histo package, so this also needs a
histo checkout. The other modules of histo are internal; their interfaces may
change without notice.
Contact
Ask the histo team before you depend on anything not listed on this page.