Skip to content

BoxToMaskD and MaskToBoxD Transpose Results #8998

Description

@oseymour

Describe the bug
The outputs from BoxToMaskD and MaskToBoxD are transposed in the spatial dims relative to the inputs.

To Reproduce

from monai.apps.detection.transforms.dictionary import BoxToMaskD, MaskToBoxD
import numpy as np
import matplotlib.pyplot as plt

ins = {
    "image": np.zeros([1,10,10]),
    "boxes": np.array([
        [1, 1, 3, 6],  # xmin, ymin, xmax, ymax
    ]),
    "labels": np.array([1,]),
}

box2mask_out = BoxToMaskD(
    box_keys="boxes", label_keys="labels", box_ref_image_keys="image",
    box_mask_keys="box_masks", min_fg_label=1, ellipse_mask=False
)(ins)

mask2box_out = MaskToBoxD(
    box_keys="boxes", box_mask_keys="box_masks", label_keys="label",
    min_fg_label=1,
)(box2mask_out)

plt.figure(figsize=[8,12])

plt.subplot(3,2,1)
plt.title("Input Box")
plt.imshow(ins["image"].squeeze())
plt.hlines(y=[ins["boxes"][0,1], ins["boxes"][0,3]], xmin=ins["boxes"][0,0], xmax=ins["boxes"][0,2])
plt.vlines(x=[ins["boxes"][0,0], ins["boxes"][0,2]], ymin=ins["boxes"][0,1], ymax=ins["boxes"][0,3])

plt.subplot(3,2,3)
plt.title("Box from BoxToMaskD")
plt.imshow(box2mask_out["image"].squeeze())
plt.hlines(y=[box2mask_out["boxes"][0,1], box2mask_out["boxes"][0,3]], xmin=box2mask_out["boxes"][0,0], xmax=box2mask_out["boxes"][0,2])
plt.vlines(x=[box2mask_out["boxes"][0,0], box2mask_out["boxes"][0,2]], ymin=box2mask_out["boxes"][0,1], ymax=box2mask_out["boxes"][0,3])

plt.subplot(3,2,4)
plt.title("Mask from BoxToMaskD")
plt.imshow(box2mask_out["image"].squeeze())
plt.imshow(box2mask_out["box_masks"].squeeze())

plt.subplot(3,2,5)
plt.title("Box from MaskToBoxD")
plt.imshow(mask2box_out["image"].squeeze())
plt.hlines(y=[mask2box_out["boxes"][0,1], mask2box_out["boxes"][0,3]], xmin=mask2box_out["boxes"][0,0], xmax=mask2box_out["boxes"][0,2])
plt.vlines(x=[mask2box_out["boxes"][0,0], mask2box_out["boxes"][0,2]], ymin=mask2box_out["boxes"][0,1], ymax=mask2box_out["boxes"][0,3])

plt.subplot(3,2,6)
plt.title("Mask from MaskToBoxD")
plt.imshow(mask2box_out["image"].squeeze())
plt.imshow(mask2box_out["box_masks"].squeeze())

plt.show()

Expected behavior
I would expect the masks and the boxes to be oriented in the same direction.

Screenshots
Plot from example code

Image

Environment

Ensuring you use the relevant python executable, please paste the output of:

================================
Printing MONAI config...
================================
MONAI version: 1.5.2
Numpy version: 2.4.4
Pytorch version: 2.5.1+cu121
MONAI flags: HAS_EXT = False, USE_COMPILED = False, USE_META_DICT = False
MONAI rev id: d18565fb3e4fd8c556707f91ac280a2dc3f681c1
MONAI __file__: C:\Users\<username>.HCAD\Documents\GitHub\eval-iu-dino\.venv\Lib\site-packages\monai\__init__.py

Optional dependencies:
Pytorch Ignite version: NOT INSTALLED or UNKNOWN VERSION.
ITK version: NOT INSTALLED or UNKNOWN VERSION.
Nibabel version: NOT INSTALLED or UNKNOWN VERSION.
scikit-image version: NOT INSTALLED or UNKNOWN VERSION.
scipy version: 1.18.0
Pillow version: 12.2.0
Tensorboard version: 2.20.0
gdown version: NOT INSTALLED or UNKNOWN VERSION.
TorchVision version: 0.20.1+cu121
tqdm version: 4.66.5
lmdb version: NOT INSTALLED or UNKNOWN VERSION.
psutil version: 7.2.2
pandas version: NOT INSTALLED or UNKNOWN VERSION.
einops version: NOT INSTALLED or UNKNOWN VERSION.
transformers version: NOT INSTALLED or UNKNOWN VERSION.
mlflow version: NOT INSTALLED or UNKNOWN VERSION.
pynrrd version: NOT INSTALLED or UNKNOWN VERSION.
clearml version: NOT INSTALLED or UNKNOWN VERSION.

For details about installing the optional dependencies, please visit:
    https://docs.monai.io/en/latest/installation.html#installing-the-recommended-dependencies


================================
Printing system config...
================================
System: Windows
Win32 version: ('11', '10.0.22631', 'SP0', 'Multiprocessor Free')
Win32 edition: Enterprise
Platform: Windows-11-10.0.22631-SP0
Processor: Intel64 Family 6 Model 186 Stepping 3, GenuineIntel
Machine: AMD64
Python version: 3.12.9
Process name: python.exe
Command: ['C:\\Users\\223113289.HCAD\\AppData\\Roaming\\uv\\python\\cpython-3.12.9-windows-x86_64-none\\python.exe', '-c', 'import monai; monai.config.print_debug_info()']
Open files: [popenfile(path='C:\\Windows\\System32\\en-US\\KernelBase.dll.mui', fd=-1)]
Num physical CPUs: 10
Num logical CPUs: 12
Num usable CPUs: 12
CPU usage (%): [0.6, 0.0, 1.3, 0.0, 49.1, 30.2, 5.4, 0.3, 24.9, 4.7, 0.0, 0.9]
CPU freq. (MHz): 1472
Load avg. in last 1, 5, 15 mins (%): [0.0, 0.0, 0.0]
Disk usage (%): 88.7
Avg. sensor temp. (Celsius): UNKNOWN for given OS
Total physical memory (GB): 31.4
Available memory (GB): 10.3
Used memory (GB): 21.1

================================
Printing GPU config...
================================
Num GPUs: 0
Has CUDA: False
cuDNN enabled: True
NVIDIA_TF32_OVERRIDE: None
TORCH_ALLOW_TF32_CUBLAS_OVERRIDE: None
cuDNN version: 90100

Additional context
Add any other context about the problem here.

Activity

aymuos15 commented on Jul 17, 2026

@aymuos15
Contributor

Thanks for the repro!

The transforms aren't transposing tho actually. The box round-trips exactly: [1, 1, 3, 6] goes through BoxToMaskD then MaskToBoxD and comes back as [1., 1., 3., 6.]. The mask sits exactly where the box says it should:

Box coordinate Value Array axis it indexes Mask foreground
xmin, xmax 1, 3 axis 0, which imshow draws vertically 1:3
ymin, ymax 1, 6 axis 1, which imshow draws horizontally 1:6

So x is the first spatial array axis, not the horizontal one, and a box [xmin, ymin, xmax, ymax] covers image[..., xmin:xmax, ymin:ymax]. Your hlines/vlines calls treat coordinate 0 as horizontal, so the outline and the mask get drawn perpendicular to each other even though the data agrees:

box/mask orientation

Swapping the arguments lines them up:

b = box2mask_out["boxes"][0] - 0.5   # imshow centres pixel i at i
plt.imshow(box2mask_out["box_masks"].squeeze())
plt.hlines(y=[b[0], b[2]], xmin=b[1], xmax=b[3])   # x -> vertical
plt.vlines(x=[b[1], b[3]], ymin=b[0], ymax=b[2])   # y -> horizontal

That said, the docs never say which axis xmin indexes, and "xyxy" means the opposite in torchvision.

@ericspod does this warrant a doc change?

oseymour commented on Jul 17, 2026

@oseymour
Author

Is the near-universal standard for axes not that X is the horizontal axis (spatial axis 1) and Y is vertical (spatial axis 0)? So if MONAI's standard box format is xyxy and you're saying y=b[0] in your plt.hlines() call then you're saying y is equal to an x coordinate.

Also, if my [xmin, ymin, xmax, ymax] is [1, 1, 3, 6] then the box should be 2 pixels wide and 5 pixels tall. So it should be taller than it is wide.

aymuos15 commented on Jul 17, 2026

@aymuos15
Contributor

I do realise that's the general comp vision norm but medical imaging follows the opposite as seen here (NIfTI maps voxel (i,j,k) to world (x,y,z) in order, so the first array axis is x). MONAI does too, e.g. MetaTensor defaults to RAS space and flip_boxes on axis 0 matches Flipd(spatial_axis=0).

So the transpose look is cos imshow draws axis 0 vertically.

oseymour commented on Jul 22, 2026

@oseymour
Author

I would strongly push back against you saying medical imaging diverges from computer vision here. For example, I could point to C8.5.5.1.14 in the DICOM standard which says "the lower right corner is x=image width - 1, and y=image length - 1" implying that x is the horizontal axis in the DICOM standard.

I agree that imshow draws axis 0 vertically, but I bet the vast majority of MONAI users would call the vertical axis the y axis.

I don't follow your comment about Flipd(spatial_axis=0).

I'd be really curious to see what other maintainers' opinions are on this is.

ericspod commented on Aug 25, 2026

@ericspod
Member

Hi @oseymour sorry for the late catchup. The two box transforms boil down to using convert_box_to_mask and convert_mask_to_box. What both of these understand for a box with value np.array([[1, 1, 3, 6]]) is that the masked area in the first spatial dimension starts from 1 and goes to 3, and in the second dimension starts from 1 and goes to 6. This is independent of what we choose to label these dimensions. We can see this by looking at the actual array data itself rather than through Matplotlib:

import numpy as np
from monai.apps.detection.transforms.box_ops import convert_box_to_mask, convert_mask_to_box

img = np.zeros([1, 10, 10])
boxes = np.array([[1, 1, 3, 6]])  # xmin, ymin, xmax, ymax
labels = np.array([1])

# construct a mask with the helper function
m1 = convert_box_to_mask(boxes, labels, img.shape[1:], bg_label=0)

# construct a mask manually using the expected behaviour
m2 = np.zeros_like(m1)
m2[0, 1:3, 1:6] = 1

print(np.allclose(m1, m2))  # True

We would then expect to get the original box parameters back from convert_mask_to_box:

print(convert_mask_to_box(m1, bg_label=0))  # (array([[1., 1., 3., 6.]], dtype=float32), array([1]))
print(convert_mask_to_box(m2, bg_label=0))  # (array([[1., 1., 3., 6.]], dtype=float32), array([1]))

These outputs conform with what is stored in your transform outputs:

print(np.allclose(box2mask_out["box_masks"], m1))  # True
print(np.allclose(mask2box_out["boxes"], boxes))  # True

I believe the transforms are correct and the issue is with the interpretation of dimensions along with what Matplotlib expects. It's much more helpful to look at the array data directly for this reason, here it's feasible to just print m1 rather than rely on visualisation:

array([[[0, 0, 0, 0, 0, 0, 0, 0, 0, 0],
        [0, 1, 1, 1, 1, 1, 0, 0, 0, 0],
        [0, 1, 1, 1, 1, 1, 0, 0, 0, 0],
        [0, 0, 0, 0, 0, 0, 0, 0, 0, 0],
        [0, 0, 0, 0, 0, 0, 0, 0, 0, 0],
        [0, 0, 0, 0, 0, 0, 0, 0, 0, 0],
        [0, 0, 0, 0, 0, 0, 0, 0, 0, 0],
        [0, 0, 0, 0, 0, 0, 0, 0, 0, 0],
        [0, 0, 0, 0, 0, 0, 0, 0, 0, 0],
        [0, 0, 0, 0, 0, 0, 0, 0, 0, 0]]], dtype=int16)

We can definitely discuss how the dimensions are labelled as this seems be a common point of confusion, we've consistently stated that array/tensor dimensions are NCHW or NCHWD which aligns with what Matplotlib expects and common image orientations like RAS. PyTorch labels dimensions NCDHW so we're not consistent with them, updating docs to state a more sensible interpretation has been proposed (#7365) but this requires feedback.

modelpath-dev commented on Sep 10, 2026

@modelpath-dev

I will take this issue. Please assign it to me.

The problem seems to be that the spatial dimensions are transposed between the outputs of BoxToMaskD and MaskToBoxD. I will first check the transformation logic in both classes, focusing on how spatial dimensions are handled. A likely fix is to ensure consistent dimension ordering during these transformations. I will reproduce the issue using the provided code and verify the fix by checking the orientation of the boxes and masks.

modelpath-dev commented on Sep 13, 2026

@modelpath-dev

I will take this issue. Please assign it to me.

The problem seems to be with the interpretation of spatial dimensions between the outputs of BoxToMaskD and MaskToBoxD. I will look into the transformation logic in both classes to ensure consistent handling of dimensions. I will use the provided code to reproduce the issue and verify the fix by checking the orientation of the boxes and masks.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions