SLICEChat: Progressive In-Encoder Token Pruning for Whole-Slide Pathology Language Models

Ali Kerem Bozkurt1,2, Barış Cem Bakay1,2, İbrahim Kulaç1, Çiğdem Gündüz-Demir1,2, Erkut Erdem2,3 Aykut Erdem1,2
1Koç University, Istanbul, Turkey 2KUIS AI Center, Istanbul, Turkey 3Hacettepe University, Ankara, Turkey
SLICEChat architecture: whole-slide patches are encoded with CONCH v1.5, progressively pruned inside a hybrid Mamba–Transformer slide encoder, and projected into a multimodal language model.
SLICEChat learns which regions to keep during slide encoding, building compact visual representations for accurate and efficient whole-slide reasoning.

Abstract

Whole-slide pathology images (WSIs) contain gigapixel-scale visual content, creating a major scalability challenge for slide-level multimodal large language models (MLLMs). Existing approaches process thousands of patch tokens and typically apply compression only after slide encoding, leaving multimodal attention computationally expensive. We introduce SLICEChat, a slide-level MLLM that integrates progressive token pruning within a hybrid Mamba–Transformer slide encoder. Mamba layers enable efficient long-range propagation, while Transformer layers preserve global interactions as the sequence is progressively shortened. Between stages, language-supervised, region-aware pruning removes spatially coherent low-utility regions under a controlled keep-rate schedule, producing compact slide representations before multimodal fusion. On SlideBench VQA, SLICEChat achieves 79.84% accuracy on TCGA and 59.09% on BCNB cohorts, outperforming the evaluated prior slide-level pathology MLLMs in overall accuracy. It also achieves the highest overall WSI-Precision and WSI-Relevance on WSI-Bench, with competitive memory usage and the second-lowest inference latency among the evaluated models. These results demonstrate accurate and computationally efficient multimodal reasoning over gigapixel WSIs.

Progressive In-Encoder Token Pruning

A compact slide representation, built progressively from thousands of tissue patches.

Hybrid slide encoding

We use CONCH v1.5 the patch encoder. In the slide encoder, each hybrid stage combines three Vision Mamba blocks with a Transformer block, pairing efficient long-range propagation with global attention.

Learned region selection

Language-supervised routers score contextualized tokens between encoder stages. Five pruning modules progressively remove low-scoring, spatially coherent regions to reach a controlled token budget.

Spatially aware reasoning

A two-layer MLP projects retained tokens into Qwen2.5-VL-7B. Original patch coordinates are preserved through M-RoPE, maintaining spatial relationships during multimodal fusion.

A whole-slide image is progressively reduced from 4,785 patch tokens to 957 across five pruning stages. Orange marks newly removed tokens; dark blue marks tokens removed earlier.
From 4,785 to 957 tokens. Each stage removes a coherent tissue region. Newly removed tokens appear in orange; previously removed tokens appear in dark blue. Retained tokens preserve their original row-major order.

Preserving tissue neighborhoods

Irregular tissue boundaries can make contiguous raster windows span disconnected strips. Hilbert indexing keeps nearby patches close in a temporary pruning sequence, producing compact candidate regions. The router removes the lowest-scoring window.

Comparison of raster and Hilbert indexing on an irregular tissue mask: Hilbert windows form more compact spatial regions.
Spatial support of pruning windows under raster and Hilbert indexing.

Training SLICEChat

Slide encoder training: Stage 1 uses masked autoencoding without pruning; Stage 2 learns pruning through contrastive slide–text alignment.
Slide encoder training. Blue arrows show self-supervised pretraining; red arrows show language-supervised pruning.
  1. Self-supervised pretraining

    Learn slide representations from unlabeled TCGA slides using masked autoencoding. Pruning is disabled while the hybrid backbone learns long-range context.

  2. Language-supervised pruning

    Train the routing modules with slide–caption pairs and CLIP-style objectives. The slide backbone and PubMedBERT text encoder remain frozen.

  3. Multimodal instruction tuning

    Jointly train the projection layer and language model on slide captions and visual question answering examples, keeping the slide encoder frozen.

Results

Strong whole-slide understanding with a compact visual representation.

SlideBench VQA

SLICEChat achieves the highest overall accuracy on both TCGA and BCNB among the evaluated models. Slide-level models in this comparison are trained on SlideInstruct.

SlideBench VQA accuracy (%). Higher is better; bold indicates the best score in each column.
ModelTCGABCNB
ClinicalDiagnosisMicroscopyOverall
SlideChat69.3972.6182.4674.9554.04
HistoSelect80.6174.1281.6876.5153.24
SLICEChat (ours)75.5179.2582.4679.8459.09

WSI-Bench

SLICEChat leads in overall WSI-Precision and WSI-Relevance. These models are trained on the WSI-Bench training set.

WSI-Bench results. Higher is better; bold indicates the best score for each metric and task.
MetricModelDiagnosisMorphology
analysis
Treatment
planning
ReportOverall
WSI-PrecisionWSI-LLaVA0.5970.5300.7940.3370.538
HistoSelect0.5600.5160.7180.3260.518
SLICEChat (ours)0.6720.5940.7900.4700.607
WSI-RelevanceWSI-LLaVA0.8220.7450.8660.6780.760
HistoSelect0.8160.7430.8230.6790.756
SLICEChat (ours)0.8450.7570.8890.7170.776

Computational Efficiency

On a typical whole-slide input with 11,948 patches, SLICEChat has the second-lowest latency and peak memory consumption among the compared models.

Measured on an NVIDIA A100 GPU. Lower is better; bold indicates the best score.
ModelPeak memory (GiB)Latency (ms)
SlideChat16.5602,405
WSI-LLaVA15.0481,775
HistoSelect18.6043,048
SLICEChat (ours)15.0542,001

The comparison uses inputs with equal patch counts. WSI-LLaVA and HistoSelect preprocessing components were integrated into their respective end-to-end evaluation pipelines. Running the original repositories without this integration would result in significantly higher latency.

BibTeX

@misc{bozkurt2026slicechatprogressiveinencodertoken,
      title={SLICEChat: Progressive In-Encoder Token Pruning for Whole-Slide Pathology Language Models}, 
      author={Ali Kerem Bozkurt and Baris Cem Bakay and Ibrahim Kulac and Cigdem Gunduz-Demir and Erkut Erdem and Aykut Erdem},
      year={2026},
      eprint={2609.24894},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2609.24894}, 
}