---
title: "Histo domain representations"
description: "Find, load and read the per-slide domain representations the histo pipeline publishes, and run the pipeline yourself."
image: "https://docs.virdx.dev/img/virdx-social-card.png"
---

> Documentation Index
> Fetch the complete documentation index at: https://docs.virdx.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# Histo domain representations

A **domain representation** is what the histo pipeline produces for one H&E
whole-slide image: five label maps that describe the slide down to single cells.
This page is for teams that want to use them.

There are three ways in:

1. **Read the published outputs** from the data platform. This is the normal
   route. You need `vxdata-sdk` and `zarr`, but not the `histo` package.
2. **Production runs** on the cluster, which publish new outputs to the
   platform.
3. **Local runs**, which write outputs to shared storage for your own
   experiments.

## 1. Read published outputs

### What is published

Each processed scan has one `HistoDomainRep` bundle in the `histo_domain_reps`
namespace. The bundle points to five `HistoMap`s.

| Bundle field           | HistoMap task          | Map                |
| ---------------------- | ---------------------- | ------------------ |
| `tissue_segmentation`  | `TISSUE_SEGMENTATION`  | `tissue_seg.zarr`  |
| `cell_segmentation`    | `CELL_SEGMENTATION`    | `cell_seg.zarr`    |
| `nuclei_segmentation`  | `NUCLEI_SEGMENTATION`  | `nuclei_seg.zarr`  |
| `gleason_segmentation` | `GLEASON_SEGMENTATION` | `gleason_seg.zarr` |
| `structure_tensor`     | `STRUCTURE_TENSOR`     | `structure_tensor.zarr` |

Identifiers carry the pipeline version, so several versions can exist side by side:

```
histodomainrep/<scan path>/v2.0
histomap/<scan path>/tissue_seg-v2.0
```

`<scan path>` is the source `HistoScan` identifier without the `histoscan/`
prefix. The bundle's `parent_identifier` is that `HistoScan`, so you can find the
H&E the maps were made from.

As of 2026-09-25 there are 8,474 `v2.0` bundles: histai 4,230, essen01 1,865,
frankfurt 893, essen02 818, aggc2022 387, chimera 281.

### Load one slide

```python
from pathlib import Path

import zarr
from vxdata.sdk import Client, F

client = Client()

# Pick bundles: by dataset, by scan, or by version.
reps = (
client.histo_domain_reps.query()
.filter(F.datasource_id == "datasource/essen01", F.histo_version == "v2.0")
.collect()
)
rep = reps.row(0, named=True)

# Resolve the five maps to storage urls.
fields = ["tissue_segmentation", "cell_segmentation", "nuclei_segmentation",
      "gleason_segmentation", "structure_tensor"]
maps = client.histo_maps.query().filter(F.identifier.is_in([rep[f] for f in fields])).collect()
urls = dict(zip(maps["task"], maps["url"]))

# Download the maps you need (each url is a zarr directory) and open them.
dest = Path("domain_reps")
tissue_path = Path(client.storage.download(urls["TISSUE_SEGMENTATION"], dest))
cell_path = Path(client.storage.download(urls["CELL_SEGMENTATION"], dest))

tissue_group = zarr.open_group(tissue_path, mode="r")
cell_group = zarr.open_group(cell_path, mode="r")
tissue_seg = tissue_group["scale_1"]          # (H, W) uint8
cell_seg = cell_group["scale_1"]              # (H, W) uint32
cell_lut = cell_group["cell_lut"][:]          # (n_cells, 3) uint32
classes = tissue_group.attrs["virdx_attrs"]["misc"]  # class names, ids, colours
```

The H&E comes from the parent scan: `client.histo_scans.query()
.filter(F.identifier == rep["parent_identifier"])`. Its `url` is the
preprocessed pyramidal TIFF.

### The maps

Every map is an OME-Zarr (zarr v3) group. The image is the array `scale_1`;
there is no pyramid. `cell_seg.zarr` also holds the `cell_lut` array. The four
label maps share one pixel grid: the H&E pyramid level closest to 10x (about
0.97 µm/px). The structure tensor is on a 4x coarser grid. The exact µm/px of
each map is the `scale` in its OME metadata and the `affine` in `virdx_attrs`.

| Map                | Shape         | dtype  | Pixel value |
| ------------------ | ------------- | ------ | ----------- |
| `tissue_seg`       | (H, W)        | uint8  | tissue class, with Gleason grade on epithelium |
| `cell_seg`         | (H, W)        | uint32 | cell id on cell pixels, tissue class elsewhere |
| `nuclei_seg`       | (H, W)        | uint32 | nucleus id (= its cell id) on nucleus pixels, tissue class elsewhere |
| `gleason_seg`      | (H, W)        | uint8  | 0 background, 1 benign, 3/4/5 Gleason grade |
| `structure_tensor` | (⌈H/4⌉, ⌈W/4⌉, 2) | uint8  | fibre orientation and coherence, 4x subsampled |

#### One id space

`tissue_seg`, `cell_seg` and `nuclei_seg` use the same id space:

| Ids       | Meaning |
| --------- | ------- |
| 0–27      | base tissue classes (the `fine28` taxonomy; 0 = background) |
| 128–139   | Gleason-graded epithelium, e.g. 132 = `luminal_epithelium_g4` |
| 255       | unknown |
| ≥ 256     | cell instance ids (only in `cell_seg` and `nuclei_seg`) |

Do not hard-code the class table. `tissue_seg`, `cell_seg` and `nuclei_seg`
store it in `attrs["virdx_attrs"]["misc"]`: `region_ids`, `region_classes`,
`region_palette_rgb` and `id_space`. The same table is in their HistoMap's
`labels` column. `gleason_seg` stores its own five-class table in the same
place, without `id_space`. The source config is `config/domain_representation/taxonomy/fine28.yaml`.

#### Cells

`cell_seg.zarr/cell_lut` has one row per cell: `[cell_id, base_class,
gleason_grade]`. `base_class` is a 0–27 id. `gleason_grade` is 3, 4 or 5 for
graded epithelium cells and 0 for all others.

The maps agree by construction:

- On cell pixels, `tissue_seg` is the cell's class, graded if the cell has a grade.
- Off cell pixels, `cell_seg == tissue_seg`.
- Every nucleus id in `nuclei_seg` is also a cell id in `cell_seg`.
- Every cell id in the map has a LUT row. A few LUT rows (about 0.3 % on a
  test slide) have no pixels, so join from the map to the LUT, not the reverse.

Only these classes hold cells: epithelia (1, 2, 22, 23), stroma (3, 4),
inflammation, nerve, vessels and arteries. Lumen, fat, blood and artefact
classes have none.

#### Structure tensor

Channel 0 is the dominant gradient orientation in radians, in [-π/2, π/2]. The
fibres run perpendicular to it. Channel 1 is coherence in [0, 1]. Decode with
`value * scale + offset`. The scale and offset are in
`attrs["virdx_attrs"]["misc"]["structure_tensor_encoding"]`. Pixel
`(r, c)` samples full-grid pixel `(4r, 4c)`. With `histo` installed,
`histo.domain_representation.representation.decode_structure_tensor(array)`
does the decode.

### Provenance and versions

Each bundle records:

- `histo_version`: the pipeline version (`v2.0`). It changes whenever the
  outputs change meaning.
- `tissue_model_id`, `gleason_model_id`, `cell_model_id`: the ClearML model ids.
- `generation_parameters`: the full pipeline config, including the Docker
  image tag and the taxonomy.
- `generation_outputs`: µm/px (`mpp`), cell count, timings, and urls of the
  run log (`run.json`) and event log (`events.jsonl`).

A bundle is written only after its five maps are uploaded. If a bundle exists,
the scan is complete.

### Size and practical notes

- Biopsy slides are small (for example 27k × 6k px). Frankfurt whole-mounts are
  about 47k × 47k px. There, `tissue_seg` is about 100 MB and `cell_seg` several
  times more. Read windows (`tissue_seg[y0:y1, x0:x1]`) instead of whole arrays
  when you can.
- Label map chunks are 1024 × 1024 px; structure tensor chunks are 256 × 256.
  All are compressed with zstd.
- `storage.download(url, dest)` keeps the key path under `dest`, without the
  bucket name.

## 2. Production runs: publish to the data platform

A production run processes HistoScans on the cluster and registers the results
on the platform, as described in section 1. The cluster already has everything
the run needs: the image holds the code, the environment and the config, the
job mounts the ClearML, S3 and vxData credentials from cluster secrets, and the
backbone weights come from shared storage. You need `argo` access to the
`argo-workflows` namespace and a histo checkout for the job file.

**Test on staging first.** `run_environment=staging` (the default) registers
the bundles on the staging platform and uploads to the `domrep_v2_testing`
group. `run_environment=production` writes the official `v2.0` bundles that
other teams read (upload group `domainrep_v2.0`). Agree production runs with
the histo team.

```bash
argo submit -n argo-workflows jobs/run-domain-representation-v2.yaml \
  -p image=histo -p image_tag=main-088def2 -p username=<you> \
  -p run_environment=staging -p datasource=essen01
```

`main-088def2` is the image that made most `v2.0` bundles. It uses the same
models as `main`.

| Parameter | What it sets |
| --------- | ------------ |
| `datasource` | Process every preprocessed scan of this datasource, e.g. `essen01`. |
| `scans_file` | Instead of `datasource`: a JSON list of HistoScan identifiers, under `/mnt/artifacts` so the pods can read it. |
| `run_environment` | `staging` (default) or `production`. |
| `gpu_type` | Pin the GPU: `NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition` or `NVIDIA-RTX-6000-Ada-Generation`. Empty takes either. |
| `worklist_mpx` | Slide area per GPU pod, in megapixels at 10x. The default 40000 takes about 2 h on Blackwell. |
| `overrides` | Extra Hydra overrides for the pipeline, e.g. `output.worklist_log_dir=/mnt/artifacts/<you>/logs` to keep the per-pod logs. |

What happens:

- A list step finds the scans that have no bundle with this pipeline version
  and these model ids, and packs them into worklists. Resubmitting the same
  job only processes what is left.
- One GPU pod (128 Gi) per worklist loads the models once and processes its
  scans one by one. A failed scan does not stop the others; the pod then exits
  with code 13.
- Inputs are preprocessed HistoScans. A scan's `FOREGROUND_MASK` map is used
  when it has one; without it every window is embedded, which is slower.

The job runs in your own Kueue queue (`user-yourname`). To also use idle GPUs
elsewhere on the cluster, add
`-l kueue.x-k8s.io/queue-name=preemptible,virdx.dev/preemptible=true` to
`argo submit`. Those pods can be stopped and rerun, which is safe because done
scans are skipped.

Watch the run with `argo list -n argo-workflows`, `argo get -n argo-workflows
<workflow>` and `argo logs -n argo-workflows <workflow>`. Finished scans appear
as bundles on the platform.

## 3. Local runs: outputs for your own experiments

A local run writes the same five maps to shared storage and never touches the
platform. Use it to try other settings, or slides that are not on the platform.

You need:

- A GPU. The coder pod's 24 GB GPU slice works with the `local` profiles.
- A histo checkout with its environment:
  `pixi install --manifest-path packages/histo/pyproject.toml -e cuda-dev`.
- ClearML credentials in `~/clearml.conf`. The three heads are downloaded from
  `clearml.fra.virdx.dev` on the first run and cached in `~/.clearml/cache`.
- `/mnt/artifacts` mounted. The backbone weights are read from there, so no
  Hugging Face token is needed.
- A vxData token only if your inputs are platform scans.

Run from the repo root:

```bash
SSL_CERT_FILE=/etc/ssl/certs/ca-certificates.crt \
REQUESTS_CA_BUNDLE=/etc/ssl/certs/ca-certificates.crt \
pixi run --manifest-path packages/histo/pyproject.toml -e cuda-dev \
  python scripts/domain_representation/generate_domain_representation.py \
  compute=local 'inputs.paths=[/path/to/slide.tiff]'
```

The two certificate variables let Python trust the internal CA of
`clearml.fra.virdx.dev`. On the coder pod the system bundle already contains
it.

Inputs, one of:

- `inputs.paths=[...]`: preprocessed pyramidal TIFFs. If
  `<name>_tissue_mask.ome.zarr` sits next to `<name>.tiff`, only tissue is
  embedded.
- `inputs.list_file=slides.json`: a JSON list of `{"path": ..., "identifier": ...}`.
  `identifier` is optional and names the output folder.
- `inputs.scans=[histoscan/...]`: platform scans, downloaded first.

`inputs.limit=N` processes only the first N slides.

Profiles: `compute=local` (fits the coder pod) and `compute=local_small`
(smaller windows, for quick checks).

Each slide gets a folder
`/mnt/artifacts/$USER/histo/domain_representation_v2/<slide>/` with the five
maps (`tissue_seg.zarr`, `cell_seg.zarr`, ...), `run.json` (settings, models,
timings) and `events.jsonl`. Set `output.scratch_dir=...` to write elsewhere.
Open the maps with `zarr` as in section 1; their layout is the same.

A slide whose folder already has a `run.json` is skipped. Delete the folder, or
use a new `output.scratch_dir`, to compute it again.

Any setting in `config/domain_representation/generate.yaml` can be changed on
the command line, for example `watershed.enabled=false`. Settings are checked
on load, and a bad combination stops the run with a message.

### Python API

The same run from Python. The public API is what `histo.domain_representation`
exports:

| Name | What it does |
| ---- | ------------ |
| `load_settings(cfg) -> PipelineSettings` | Types and checks the composed Hydra config. |
| `load_models(settings, device=None) -> Models` | Loads H-optimus-0 and the tissue, Gleason and cell heads. |
| `process_scans(scans, settings, models, *, platform) -> ScansReport` | Processes a list of slides. Skips slides already done, and one failure does not stop the rest. `report.done`, `report.failed`. |
| `process_scan(scan, settings, models, *, output_dir, ...)` | Processes one slide into `output_dir`. |
| `local_scans(settings.inputs) -> list[ScanInput]` | Slides from local paths or a JSON list file. |
| `ScanInput`, `PlatformScan` | One slide on disk / one `HistoScan` on the platform. |
| `Platform` | Resolves scans and publishes outputs. Pass `platform=None` to write locally only. |

```python
from hydra import compose, initialize_config_dir
from histo.domain_representation import load_models, load_settings, local_scans, process_scans

with initialize_config_dir(version_base=None, config_dir="/abs/path/to/histo/config/domain_representation"):
cfg = compose("generate", overrides=["compute=local", "inputs.paths=[/data/slide.tiff]"])
settings = load_settings(cfg)
report = process_scans(local_scans(settings.inputs), settings, load_models(settings), platform=None)
```

The config is in the repo, not in the `histo` package, so this also needs a
histo checkout. The other modules of `histo` are internal; their interfaces may
change without notice.

## Contact

Ask the histo team before you depend on anything not listed on this page.

Source: https://docs.virdx.dev/histo/index.mdx
