SurgTPGS++: Text-Promptable Gaussian Splatting with Dense Features Aggregation for Semantic 4D Surgical Scene Understanding

A unified Gaussian Splatting framework for dynamic 4D reconstruction, real-time language-assisted 3D querying, and semantic segmentation in surgical scenes.

Medical Image Analysis 2026
Extended work of SurgTPGS, MICCAI 2025
1The Chinese University of Hong Kong   2Technical University of Munich   3University of Strasbourg & IHU Strasbourg   4The University of Manchester
*Equal contribution, †Corresponding author
SurgTPGS++ teaser showing 4D reconstruction, vision-language query, and text-promptable 3D segmentation.
SurgTPGS++ complements the limitations of standalone 4D Gaussian reconstruction, vision-language models, and segmentation networks by unifying real-time novel-view rendering, language-based querying, and semantic 3D segmentation.

Abstract

With the growing demand for surgical embodied intelligence, accurate semantic understanding of 3D surgical scenes with language-based interaction has become increasingly important. Such capability can assist surgeons in identifying and interacting with surgical instruments and anatomical structures during pre-operative planning and real-time intra-operative guidance. However, existing methods typically address surgical vision-language modeling, 3D reconstruction, and semantic segmentation as separate tasks, leaving real-time language-assisted 3D querying in dynamic surgical scenes largely unexplored.

We present SurgTPGS++, a Gaussian Splatting pipeline for text-promptable 3D surgical scene understanding. SurgTPGS++ introduces semantic feature aggregation (SFA) to extract and integrate rich vision-language features from multi-domain VLMs, embeds the aggregated features into semantic-aware 3D Gaussians, and applies semantic-aware deformation tracking (SADT) to model temporal deformation of both texture and semantic features. A codebook-based query (CQ) module then matches rendered semantic features with text-codebook embeddings, enabling real-time language-assisted 3D segmentation from arbitrary viewpoints and timestamps.

Highlights

Unified Surgical 4D Understanding

SurgTPGS++ jointly represents appearance, geometry, dynamics, and language-aligned semantics in one deformable Gaussian scene.

Semantic Feature Aggregation

SFA fuses patch-level CLIP features with intermediate ViT features to produce dense semantic supervision without explicit mask proposals.

Semantic-Aware Deformation

SADT deforms semantic features alongside texture and geometry, preserving temporal consistency under tissue and tool motion.

Codebook-Based Query

CQ converts multi-text prompting into a stable in-codebook assignment, enabling efficient real-time 3D segmentation.

Dynamic Text-Promptable Segmentation

SurgTPGS++ showcase of consistent 3D text-promptable segmentation in a dynamic surgical scene.
SurgTPGS++ provides consistent 3D text-promptable segmentation over time while maintaining real-time novel-view rendering above 100 FPS in dynamic surgical scenes.

Pipeline

Overview of the SurgTPGS++ pipeline with SFA, SADT, and CQ modules.
The pipeline extracts dense semantic features with SFA, optimizes a deformable semantic Gaussian representation with SADT, and performs text-promptable 3D segmentation by matching rendered semantic features against a text-codebook through CQ.

Extended From SurgTPGS

Dense feature aggregation. The extended work reduces reliance on SAM-based region proposals and simple VLM finetuning by aggregating multi-level VLM features into dense language-aligned maps.

Dynamic semantic tracking. Semantic features are embedded into the Gaussian representation and tracked through deformation, improving 4D consistency for surgical tools and deformable anatomy.

Efficient multi-text querying. CQ performs a single codebook matching operation instead of repeated prompt-wise querying, increasing query speed from 67.5 FPS to 246.64 FPS.

Results

89.23
mIoU on CholecSeg8K
60.29
mIoU on EndoVis18
70.57
mIoU on CaDISv2
190
FPS rendering speed
246.64
FPS query speed
3 min
Training time on EndoVis18

Quantitative Comparison

Box plot of text-query segmentation performance across CholecSeg8K, EndoVis18, and CaDISv2.
Across CholecSeg8K, EndoVis18, and CaDISv2, SurgTPGS++ achieves higher and more compact mIoU distributions than LangSplat, OpenGaussian, DGD, and SurgTPGS, indicating stronger robustness across surgical scenes and query categories.

Qualitative Comparison

Qualitative comparison of text-promptable segmentation results across surgical datasets.
Qualitative results on CholecSeg8K, EndoVis18, and CaDISv2 show that SurgTPGS++ produces cleaner semantic responses, smoother object boundaries, and more complete queried regions in novel views.

BibTeX

@article{huang2026surgtpgs++,
        title={SurgTPGS++: Text-promptable Gaussian Splatting with dense features aggregation for semantic 4D surgical scene understanding},
        author={Huang, Yiming and Bai, Long and Cui, Beilei and Yuan, Kun and Wang, Guankun and Hoque, Mobarak I and Padoy, Nicolas and Navab, Nassir and Ren, Hongliang},
        journal={Medical Image Analysis},
        pages={104318},
        year={2026},
        publisher={Elsevier}
      }