Mecha Research · Research / Medical Imaging

Locating Any Pathology Anywhere in Any Medical Image

A modality-agnostic approach to grounding free-text findings in XR, CT, and MRI.

An interactive snake moves through an adaptive quadtree field. On desktop, move the pointer to invert the field and click to relocate its target.

Executive summary

We are delighted to announce that we have invented a new way to localise any textual finding in any CT, XR or MRI study. We can localise each text finding in any level of detail, to a specific region of a specific slice within a specific series of any study. What is perhaps more remarkable is that this new approach requires no human annotations or key images. We can localise findings in any level of textual detail. For example we can map the precise text “exophytic focus at the lower pole of the left kidney” to the exact region of the CT where the model believes that finding is present.

Introduction

Solving the interpretive moment in radiology requires three key elements:

  1. The ability to accurately map medical images to complete text reports,
  2. The ability to take account of relevant patient context when producing text reports,
  3. The ability to localize any finding described in the text reports to the images.

General localization of findings in medical images is therefore something of a holy grail in AI for radiology.

In this post, we present high-level findings of a novel approach we've developed to locate any pathology anywhere in any medical image.

Preliminaries

Pathology localization itself is not novel. Many classical algorithms use segmentation models to locate specific findings in images.

Outside of segmentation models, approaches which use model outputs to approximate what's being looked at also exist. In this section, we will briefly discuss these alternative techniques and outline their main limitations.

Segmentation models

Training a segmentation model usually requires the following procedure:

  1. A large number of radiologists must read a large number of studies.
  2. As they read the studies, they look for pre-agreed pathological findings (e.g. pulmonary nodules).
  3. Whenever they find an instance of the pre-agreed findings, they must label it. This can be done by manually filling in the region of interest, or by using semi-automated tools which nevertheless require that they click in the right spot.
  4. Accrue a large dataset of medical images + labels for every instance of the pre-agreed findings.
  5. Optional: Perform quality assurance (QA) to ensure the labels make sense and are correct.
  6. Optional: Train a QA model which predicts low-quality cases and then either (i) fix them manually in a second pass or (ii) remove them from training.
  7. Train a segmentation model, usually an nnU-Net [1], to then locate the pre-agreed findings.

The above can work well if you:

  • Have many radiologists willing to label a large number of medical studies,
  • Are looking for visually unambiguous features like brain bleeds,
  • Don't mind restricting the number of things you're able to look for.

However, one big issue with classical segmentation models is that they are incredibly rigid; that is, they're only able to detect the narrow set of pathologies they're explicitly trained to look for. For instance, a model trained to locate pulmonary nodules cannot also be used to look for subdural brain bleeds.

There are many other issues too. Human labels are very noisy: some people systematically 'under-segment' objects, some 'over-segment' them, and some mislabel or miscount objects altogether (e.g. confusing vertebral level L2 with L3).

Whilst architectures such as nnU-Net [1] can be data efficient, accuracy requires data scaling, and it is often not practical to expect that experts can label million-scale datasets for a small number of pathologies at a time. This means that these systems are much less able to make good use of data scaling laws, and therefore almost always generalize relatively poorly out of distribution.

Finally, consider a Vision-Language Model (VLM) which outputs the following sentence: "Subtle hyperintensity in the middle cerebral artery (MCA)". Unless you have trained a segmentation model to detect this pathology, you'll be unable to predict what the VLM was looking at when making this claim. Supposing you did in fact have a segmentation model trained to look for hyperintensity in the MCA, then there is still no guarantee that its output matches what the VLM itself was looking for. Therefore, agentic workflows that combine VLMs with segmentation models (or, more generally, classifiers of any type) naively can lead to non-faithful explanations of the VLM’s reasoning, a setup we believe to be sub-optimal.

Post-hoc visual attribution methods

There are a number of techniques which can analyze model gradients or latents, or more broadly estimate model behaviour in controlled ways so as to allow us to approximate which features the model is using to make predictions.

In particular, we care about local explanation techniques, which attempt to explain how individual predictions are made.

These include feature importance techniques like LIME [2] and SHAP [3], rule-based approaches such as Anchors [4], and saliency map techniques like SmoothGrad [5], Integrated Gradients [6], Gradient × Input [7], and Grad-CAM [8]. There are also approaches which modify backpropagation, which is the mechanism that is used to update neural network weights. These include Guided Backpropagation [9] and Layer-wise Relevance Propagation (LRP) [10]. Meanwhile, prototype approaches attempt to explain a model with synthetic or natural input examples [11], and counterfactual approaches try to figure out which input features need to change to meaningfully change the prediction of a model [12].

Post-hoc attribution has several limitations. An explanation may look plausible without reflecting the computation that produced the prediction; some modified-backpropagation methods even remain similar after model parameters are randomized. Explanations can also be manipulated without changing the model’s output, shift after small input changes, and depend heavily on settings such as random seed or perturbation count. Depending on the approach used, faithfulness, i.e. the extent to which the explanation reflects the model’s actual computation, is not guaranteed.

Saliency map techniques are quite common and deserve an additional word. Grad-CAM, for instance, is spatially coarse by construction: it weights a late convolutional feature map and then upsamples it, which can obscure small findings and precise anatomical boundaries [8]. Its output can also vary with training randomness, with different model initializations producing qualitatively different maps [13].

mecha Wayfinder

In order to tackle the above issues, we present mecha Wayfinder, a new technique for general semantic localization.

Wayfinder's results directly reflect a VLM's actual computations and are therefore a more faithful explanation of what the model is looking at. Furthermore, Wayfinder is not limited to looking at a dozen, or even a hundred different pathologies, but can instead locate any finding of arbitrary descriptive length or complexity.

Wayfinder does not require any labelled data or segmentation masks, which means that the approach benefits directly from simple data scaling of medical images and their associated radiology reports.

General semantic grounding in 2D

In the 2D case, we use an illustrative lumbar spine study to illustrate the localization ability of Wayfinder. The VLM's report was as follows:

Mild right convexity thoracolumbar curvature. Marked loss of disc space height at L1-L2. Marked loss of disc space height at L2-L3. Moderate to marked degenerative changes at L3-L4. Moderate loss of disc space height at L4-L5. Mild changes at L5-S1. Mild lower lumbar facet arthropathy. Right total hip arthroplasty. Vascular calcifications.

We can then iterate through each individual finding to see what the VLM was focussing on when it was generating that finding.

Lumbar spine
Lumbar spine radiograph, AP viewLoading heatmap…
AP
Lumbar spine radiograph, Lateral viewLoading heatmap…
Lateral

Mild right convexity thoracolumbar curvature.

Start with the overall spinal alignment on the AP view.

General semantic grounding in 3D

Our approach extends straightforwardly to 3D imaging. Extending our approach to 3D not only allows us to highlight a given pathology but also allows us to automatically load the precise series (and window) where the VLM believes the finding is most clearly demonstrated. This is useful because a CT might contain many series and many different view-based reconstructions.

The viewer below brings together six CT studies and 30 findings. Choose a study, then move through its findings to see the regions the VLM associates with each description. Each finding opens at its selected series, window and location; the preferred view appears first, alongside the other two planes.

Use the slice controls to explore the volume, change the series or CT window, or turn off the heatmap to inspect the underlying images.

8 series · 7 findings
CT images load as you reach this section.

multinodular goiter of the thyroid

Case studies for pulmonary nodules

The following selected examples compare Wayfinder with human nodule annotations in JSRT chest radiographs and LIDC-IDRI CT scans. Select Human, AI, or Both to inspect the annotations separately or together.

Chest X-ray

These 12 radiographs show the human-annotated nodule alongside the part of the image the VLM most strongly associates with a lung nodule. Use the case selector to move through the examples.

Loading image…
Human noduleAI region

The cyan circle is the human nodule annotation. The amber square marks the part of the image the VLM most strongly associates with a lung nodule.

Chest CT

Here we illustrate two CT cases that demonstrate nodule detection in 3D. This works in the same way as the 2D case above. Move through the slices to compare the regions highlighted by the VLM with the human annotations.

Axial · Lung windowLoading CT study…
Loading complete CT study · 0% · 0.0 / 13.1 MB
Preparing all slices and annotations

Drag the slider, or click the image and scroll. Arrow keys also move through slices.

Human maskAI heatmap

Cyan shows the human nodule annotations. The heatmap shows the regions the VLM associates with a lung nodule, with brighter areas indicating a stronger association.

Limitations

Wayfinder is only as good at localization as the VLM is at finding pathology. If the VLM has no conception of a pathology, perhaps because it's vanishingly rare, Wayfinder would be unable to go beyond the VLM's perceptive ability to find it. This can also mean that the localizations themselves are noisy if the VLM only has a partial understanding of the pathology.

However, Wayfinder benefits from simple data scaling with no human labelling of the images, and this means that as the VLM becomes stronger with more data, localization ability improves in tandem.

Conclusion

In this post, we present a novel approach for general semantic localization in medical images. In particular, our approach can look for any finding, of any complexity, in any medical image. The approach works for XR, CT, and MRI, as well as any other imaging modality where a foundation model can be trained.

Traceability, where the outputs of a VLM can be attributed to the specific parts of the images from which they were derived, is likely to be extremely important for future generations of AI models in radiology. Localization allows radiologists to check findings more quickly than mapping from raw text back to images; this can dramatically improve future workflows. We also see this line of work as important for enhancing the future safety of VLMs, as it allows for substantially easier detection of hallucinations.

References

  1. Isensee, F., Jaeger, P. F., Kohl, S. A. A., Petersen, J., & Maier-Hein, K. H. “nnU-Net: a self-configuring method for deep learning-based biomedical image segmentation.” Nature Methods, 18, 203–211 (2021).
  2. Ribeiro, M. T., Singh, S., & Guestrin, C. “Why Should I Trust You?”: Explaining the Predictions of Any Classifier. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 1135–1144 (2016).
  3. Lundberg, S. M., & Lee, S.-I. “A Unified Approach to Interpreting Model Predictions.” Advances in Neural Information Processing Systems, 30, 4765–4774 (2017).
  4. Ribeiro, M. T., Singh, S., & Guestrin, C. “Anchors: High-Precision Model-Agnostic Explanations.” Proceedings of the AAAI Conference on Artificial Intelligence, 32(1), 1527–1535 (2018).
  5. Smilkov, D., Thorat, N., Kim, B., Viégas, F., & Wattenberg, M. “SmoothGrad: removing noise by adding noise.” arXiv:1706.03825 (2017).
  6. Sundararajan, M., Taly, A., & Yan, Q. “Axiomatic Attribution for Deep Networks.” Proceedings of the 34th International Conference on Machine Learning, 70, 3319–3328 (2017).
  7. Ancona, M., Ceolini, E., Öztireli, C., & Gross, M. “Towards better understanding of gradient-based attribution methods for Deep Neural Networks.” 6th International Conference on Learning Representations (2018).
  8. Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., & Batra, D. “Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization.” Proceedings of the IEEE International Conference on Computer Vision, 618–626 (2017).
  9. Springenberg, J. T., Dosovitskiy, A., Brox, T., & Riedmiller, M. A. “Striving for Simplicity: The All Convolutional Net.” ICLR Workshop (2015).
  10. Bach, S., Binder, A., Montavon, G., Klauschen, F., Müller, K.-R., & Samek, W. “On Pixel-Wise Explanations for Non-Linear Classifier Decisions by Layer-Wise Relevance Propagation.” PLOS ONE, 10(7), e0130140 (2015).
  11. Chen, C., Li, O., Tao, C., Barnett, A. J., Rudin, C., & Su, J. K. “This Looks Like That: Deep Learning for Interpretable Image Recognition.” Advances in Neural Information Processing Systems, 32, 8928–8939 (2019).
  12. Goyal, Y., Wu, Z., Ernst, J., Batra, D., Parikh, D., & Lee, S. “Counterfactual Visual Explanations.” Proceedings of the 36th International Conference on Machine Learning, 97, 2376–2384 (2019).
  13. Woerl, A.-C., Disselhoff, J., & Wand, M. “Initialization Noise in Image Gradients and Saliency Maps.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 1766–1775 (2023).

Talk to us

Interested in next-generation AI for radiology?

Talk to a founder

From the journal

Continue reading

Research10 min read

We Taught a Fly to Read X-Rays

Training a connectome-constrained network to classify pleural effusion on chest X-rays using sparse retinal inputs.

Read article
Research20 min read

Qualitative Case Studies for Assessing Foundation Models

Most metrics used to evaluate machine learning models are quantitative, but qualitative assessment is crucial.

Read article