Assessment of AI segmentation models in histopathology whole slide images: the effect of the unit of analysis
- Author(s)
- Arab, A; Garcia, V; Kahaki, S; van Rijthoven, M; Salgado, R; Gallas, BD; Ciompi, F; Petrick, N; Chen, W;
- Journal Title
- SPIE Medical Imaging
- Publication Type
- Meeting report
- Abstract
- Performance assessment of AI models for histopathology image segmentation often involves aggregating results from only regions of interest (ROIs) with expert-annotated reference standards, as whole slide images are cost-prohibitive to annotate entirely. However, standardized aggregation methods for performance assessment are lacking. We introduce three aggregation approaches based on distinct analytical units: pixel-level, ROI-level, and slide-level analysis. While uncertainty quantification is feasible for each analytical unit, slide-level uncertainty estimates may provide greater clinical relevance. Therefore, we developed a bootstrap method to estimate performance uncertainty at the slide level. We demonstrate the impact of different aggregation methods using synthetic data, then apply these approaches to evaluate a U-Net-based segmentation model designed to identify tumoral and tumor-associated stromal regions—essential preprocessing steps for automated quantification of tumor-infiltrating lymphocyte (TIL) density in breast cancer tissue slides. The segmentation model was trained and validated on independent datasets using the Dice coefficient as the primary performance metric. Results demonstrated substantial variation in segmentation accuracy across aggregation methods, with Dice coefficients ranging from 0.634 (95% CI [0.428, 0.806]) to 0.717 (95% CI [0.5, 0.835]) for tumor segmentation, and from 0.686 (95% CI [0.621, 0.765]) to 0.874 (95% CI [0.839, 0.91]) for stroma segmentation. These findings demonstrate that AI segmentation performance results are not comparable when different aggregation methods are employed, even with the same nominal metric. Explicit specification of the analytical unit is essential for meaningful performance assessment and comparison of AI models in histopathology image segmentation applications. Our code for performance assessment is publicly available at: https://github.com/DIDSR/SegVal-WSI.
- Publisher
- SPIE
- Department(s)
- Laboratory Research
- Publisher's Version
- https://doi.org/10.1117/12.3085917
- Terms of Use/Rights Notice
- Refer to copyright notice on published article.
Creation Date: 2026-06-16 12:05:20
Last Modified: 2026-06-16 12:05:36