Repository navigation
BoxToMaskD and MaskToBoxD Transpose Results #8998
Description
Activity
Thanks for the repro!
The transforms aren't transposing tho actually. The box round-trips exactly: [1, 1, 3, 6] goes through BoxToMaskD then MaskToBoxD and comes back as [1., 1., 3., 6.]. The mask sits exactly where the box says it should:
| Box coordinate | Value | Array axis it indexes | Mask foreground |
|---|---|---|---|
xmin, xmax |
1, 3 | axis 0, which imshow draws vertically |
1:3 |
ymin, ymax |
1, 6 | axis 1, which imshow draws horizontally |
1:6 |
So x is the first spatial array axis, not the horizontal one, and a box [xmin, ymin, xmax, ymax] covers image[..., xmin:xmax, ymin:ymax]. Your hlines/vlines calls treat coordinate 0 as horizontal, so the outline and the mask get drawn perpendicular to each other even though the data agrees:
Swapping the arguments lines them up:
b = box2mask_out["boxes"][0] - 0.5 # imshow centres pixel i at i
plt.imshow(box2mask_out["box_masks"].squeeze())
plt.hlines(y=[b[0], b[2]], xmin=b[1], xmax=b[3]) # x -> vertical
plt.vlines(x=[b[1], b[3]], ymin=b[0], ymax=b[2]) # y -> horizontalThat said, the docs never say which axis xmin indexes, and "xyxy" means the opposite in torchvision.
@ericspod does this warrant a doc change?
Is the near-universal standard for axes not that X is the horizontal axis (spatial axis 1) and Y is vertical (spatial axis 0)? So if MONAI's standard box format is xyxy and you're saying y=b[0] in your plt.hlines() call then you're saying y is equal to an x coordinate.
Also, if my [xmin, ymin, xmax, ymax] is [1, 1, 3, 6] then the box should be 2 pixels wide and 5 pixels tall. So it should be taller than it is wide.
I do realise that's the general comp vision norm but medical imaging follows the opposite as seen here (NIfTI maps voxel (i,j,k) to world (x,y,z) in order, so the first array axis is x). MONAI does too, e.g. MetaTensor defaults to RAS space and flip_boxes on axis 0 matches Flipd(spatial_axis=0).
So the transpose look is cos imshow draws axis 0 vertically.
I would strongly push back against you saying medical imaging diverges from computer vision here. For example, I could point to C8.5.5.1.14 in the DICOM standard which says "the lower right corner is x=image width - 1, and y=image length - 1" implying that x is the horizontal axis in the DICOM standard.
I agree that imshow draws axis 0 vertically, but I bet the vast majority of MONAI users would call the vertical axis the y axis.
I don't follow your comment about Flipd(spatial_axis=0).
I'd be really curious to see what other maintainers' opinions are on this is.
Hi @oseymour sorry for the late catchup. The two box transforms boil down to using convert_box_to_mask and convert_mask_to_box. What both of these understand for a box with value np.array([[1, 1, 3, 6]]) is that the masked area in the first spatial dimension starts from 1 and goes to 3, and in the second dimension starts from 1 and goes to 6. This is independent of what we choose to label these dimensions. We can see this by looking at the actual array data itself rather than through Matplotlib:
import numpy as np
from monai.apps.detection.transforms.box_ops import convert_box_to_mask, convert_mask_to_box
img = np.zeros([1, 10, 10])
boxes = np.array([[1, 1, 3, 6]]) # xmin, ymin, xmax, ymax
labels = np.array([1])
# construct a mask with the helper function
m1 = convert_box_to_mask(boxes, labels, img.shape[1:], bg_label=0)
# construct a mask manually using the expected behaviour
m2 = np.zeros_like(m1)
m2[0, 1:3, 1:6] = 1
print(np.allclose(m1, m2)) # TrueWe would then expect to get the original box parameters back from convert_mask_to_box:
print(convert_mask_to_box(m1, bg_label=0)) # (array([[1., 1., 3., 6.]], dtype=float32), array([1]))
print(convert_mask_to_box(m2, bg_label=0)) # (array([[1., 1., 3., 6.]], dtype=float32), array([1]))These outputs conform with what is stored in your transform outputs:
print(np.allclose(box2mask_out["box_masks"], m1)) # True
print(np.allclose(mask2box_out["boxes"], boxes)) # TrueI believe the transforms are correct and the issue is with the interpretation of dimensions along with what Matplotlib expects. It's much more helpful to look at the array data directly for this reason, here it's feasible to just print m1 rather than rely on visualisation:
array([[[0, 0, 0, 0, 0, 0, 0, 0, 0, 0],
[0, 1, 1, 1, 1, 1, 0, 0, 0, 0],
[0, 1, 1, 1, 1, 1, 0, 0, 0, 0],
[0, 0, 0, 0, 0, 0, 0, 0, 0, 0],
[0, 0, 0, 0, 0, 0, 0, 0, 0, 0],
[0, 0, 0, 0, 0, 0, 0, 0, 0, 0],
[0, 0, 0, 0, 0, 0, 0, 0, 0, 0],
[0, 0, 0, 0, 0, 0, 0, 0, 0, 0],
[0, 0, 0, 0, 0, 0, 0, 0, 0, 0],
[0, 0, 0, 0, 0, 0, 0, 0, 0, 0]]], dtype=int16)
We can definitely discuss how the dimensions are labelled as this seems be a common point of confusion, we've consistently stated that array/tensor dimensions are NCHW or NCHWD which aligns with what Matplotlib expects and common image orientations like RAS. PyTorch labels dimensions NCDHW so we're not consistent with them, updating docs to state a more sensible interpretation has been proposed (#7365) but this requires feedback.
I will take this issue. Please assign it to me.
The problem seems to be that the spatial dimensions are transposed between the outputs of BoxToMaskD and MaskToBoxD. I will first check the transformation logic in both classes, focusing on how spatial dimensions are handled. A likely fix is to ensure consistent dimension ordering during these transformations. I will reproduce the issue using the provided code and verify the fix by checking the orientation of the boxes and masks.
I will take this issue. Please assign it to me.
The problem seems to be with the interpretation of spatial dimensions between the outputs of BoxToMaskD and MaskToBoxD. I will look into the transformation logic in both classes to ensure consistent handling of dimensions. I will use the provided code to reproduce the issue and verify the fix by checking the orientation of the boxes and masks.

Describe the bug
The outputs from BoxToMaskD and MaskToBoxD are transposed in the spatial dims relative to the inputs.
To Reproduce
Expected behavior
I would expect the masks and the boxes to be oriented in the same direction.
Screenshots
Plot from example code
Environment
Ensuring you use the relevant python executable, please paste the output of:
Additional context
Add any other context about the problem here.