SCOPE-4D: Endoscopic 4D Geometry Foundation Models

Structured Camera and Organ Motion Prediction for Endoscopy

Camera, dense geometry and 3D tissue trajectories from monocular video, in one forward pass

Chaoyi Zhou1,2,*, Zhongpai Gao1,†, Anwesa Choudhuri1, Meng Zheng1,
Benjamin Planche1, Run Wang2, Terrence Chen1, Siyu Huang2,†, Ziyan Wu1

1 United Imaging Intelligence, Boston, MA, USA
2 Clemson University, Clemson, SC, USA

* The work conducted by the first author was carried out during an internship at United Imaging Intelligence.
† Corresponding authors.

1SCOPE-5K · Data 4,910 clips · gastrointestinal endoscopy and laparoscopy

In-domain benchmark training + held-out evaluation · 4,855 clips · 7 sources
SimCol3D33 clips
GISynthetic
video
label geometry
C3VD112 clips
GIReal
video
label geometry
EndoMapper-Sim5 clips
GISynthetic
video
label geometry
EndoMapper-Real2,379 clips
GIRealour labels
video
label geometry
StereoMIS354 clips
LaparoscopyRealour labels
video
label geometry
Sano1,938 clips
LaparoscopySynthetic
video
label geometry
SCARED34 clips
LaparoscopyReal
video
label geometry
Out-of-domain benchmark evaluation only · new
SCOPE-colon-phantom14 clips
GIRealnew
video
tracked camera path
SCOPE-colon-real10 clips
GIRealnew
video
second clip

2SCOPE-4D · Model one model, one forward pass

Depth and geometry

Depth
Geometry + cameras
Input video
Ground truth
VGGT-Omega
SCOPE-4D (Ours)
VGGT-Omega
SCOPE-4D (Ours)
Input video
Ground truth
Endo3R
SCOPE-4D (Ours)
Endo3R
SCOPE-4D (Ours)

Laparoscopy (top) and colonoscopy (bottom).

3D tissue trajectories

Instrument pulling tissue
Liver breathing
D4RT*
SCOPE-4D (Ours)
4RC
SCOPE-4D (Ours)

Long sequences

Abstract

Geometric understanding supports endoscopic navigation and robotic assistance, but learning reliable endoscopic geometry faces two challenges: scarce geometric annotations and ambiguity between camera motion and tissue deformation. We present SCOPE-4D, an endoscopic 4D geometry foundation model that jointly predicts camera parameters, dense geometry, and 3D tissue trajectories from monocular RGB video in a single forward pass. Our curation and annotation pipeline constructs SCOPE-5K, a collection of approximately 5,000 clips spanning real and synthetic gastrointestinal endoscopy and laparoscopy. The collection provides rich geometric supervision and includes newly collected phantom and real-colonoscopy evaluation sets. Geometric supervised fine-tuning on SCOPE-5K learns endoscopic priors that improve camera and depth estimation. Common–Residual Motion (CRM) further constrains local deformation relative to common tissue movement. Together with geometric supervision, CRM and trajectory supervision further improve camera and depth estimation over geometric fine-tuning alone while enabling dense 3D tissue tracking. Evaluations on public and newly collected benchmarks demonstrate strong in-domain and out-of-domain geometry, superior 3D tracking, and more stable long-sequence colon reconstruction. A blinded user study further supports the perceived reconstruction quality on real clinical video. Together, these results demonstrate the value of large-scale endoscopic supervision and motion constraints for joint geometry estimation and tissue tracking.

Overview

SCOPE-4D overview

Method

SCOPE-4D builds on a feed-forward multi-view transformer. Geometric fine-tuning on SCOPE-5K supervises camera and dense depth. CRM represents each 3D point's motion as a shared common motion plus a local residual. Together with geometric supervision, CRM and trajectory supervision further improve camera and depth estimation over geometric fine-tuning alone while enabling dense 3D tissue tracking.

SCOPE-4D method

SCOPE-5K data

Real endoscopic video rarely comes with geometry. Our pipeline (panel a of the method figure) turns raw videos into clips with camera and depth labels: screening and clipping, per-clip structure-from-motion with tracking-based correspondences, dense labels from a camera-conditioned geometry model, and an automatic quality-control stage that accepts, flags or rejects each clip.

Data composition

Data sourceGILaparoSyntheticRealDepthCameraClips
EndoMapper-Real✓✓✓†✓†2,379
C3VD✓✓✓✓112
SimCol3D✓✓✓✓33
EndoMapper-Sim✓✓✓✓5
Sano✓✓✓✓1,938
StereoMIS✓✓✓†✓354
SCARED✓✓✓✓34
Hamlyn✓✓✓22
RealSynCol✓✓✓✓9
SCOPE-colon-phantom (new)✓✓✓14
SCOPE-colon-real (new)✓✓10
SCOPE-5K✓✓✓✓✓✓4,910

† annotations generated by our pipeline.

Quality-control examples: one clip per verdict

Accept

Borderline

Reject

Results

Geometry and camera stability Hamlyn, out-of-domain

More examples: EndoMapper-Real; StereoMIS; SCOPE-colon-phantom

EndoMapper-Real

EndoMapper-Real

StereoMIS

SCOPE-colon-phantom

3D tissue tracking StereoMIS

Input video
SCOPE-4D (Ours)
SM4RT
D4RT*
4RC

Liver breathing

Input video
SCOPE-4D (Ours)
SM4RT
D4RT*
4RC

Instrument pulling tissue

Long-sequence colon reconstruction RealSynCol, 701 views

Interactive demo

Comparison

Colon 2 point clouds
Colon 2 trajectories

Predicted trajectory GT trajectory

Quantitative results

MethodIn-domain averageHamlyn (OOD)SCOPE-colon-phantom (OOD)
PoseDepthPoseDepthPose
AUC@5°↑AUC@30°↑ATE↓AbsRel↓δ1.25↑ATE↓ARE↓AbsRel↓δ1.25↑ATE↓RTE↓ARE*↓
General-domain
VGGT2.9429.110.1280.16074.614.6101.9640.19069.1916.5316.7812.14
VGGT-Ω7.8148.070.0490.09391.074.4961.7460.13580.828.5410.798.22
MapAnything2.6326.220.1220.14185.004.9841.9600.16674.1515.9817.7026.98
CUT3R0.2715.550.1310.30054.064.5711.6930.18570.2713.6416.5622.48
SpatialTrackerV26.0144.850.0580.11386.985.4992.2410.14277.4210.9213.8011.23
4RC6.1241.850.0680.13879.713.6391.4380.12084.2613.9815.339.13
Any4D0.2911.180.1600.21565.254.7551.1140.16772.2316.9417.5712.15
SM4RT4.8638.720.0870.14777.783.5171.3500.12083.9414.4615.509.17
Endoscopy-specific
Endo3R0.2611.140.1790.18570.0620.9829.4920.22362.7716.4217.0516.13
AF-SfMLearner1.4520.760.1340.25356.9415.5625.9790.19068.2010.3513.8810.31
EndoDAC1.4119.080.1300.20166.2218.0578.9680.19168.8711.5115.0310.54
Endo-FASt3R4.1335.020.0620.23060.399.1523.7940.21164.758.7011.399.22
SCOPE-4DSFT (Ours)29.4467.330.0160.04298.522.9271.0140.12983.778.4110.717.37
SCOPE-4D (Ours)29.8667.410.0170.04098.682.9761.0130.12485.238.2910.827.21

Camera and depth. Hamlyn depth excludes instruments. Bold best, underline second.

MethodSanoStereoMIS
TrackingPoseDepthTrackingPoseDepth
EPE↓APD↑ATE↓RTE↓RRE↓AbsRel↓δ1.25↑EPE↓APD↑ATE↓RTE↓RRE↓AbsRel↓δ1.25↑
D4RT*2.03572.940.0140.2390.2150.19365.767.54034.120.0430.9400.3440.14787.20
SpatialTrackerV20.93390.580.0170.2230.1250.14578.766.87038.980.0490.6770.1950.10795.24
4RC0.52394.720.0060.1000.0620.18666.156.26644.280.0390.4810.1460.11490.42
Any4D0.53594.410.0100.3760.3090.18866.017.27742.450.0612.2030.7650.14482.72
SM4RT0.41195.640.0070.1040.0810.18267.316.27643.950.0250.4270.1580.11692.57
SCOPE-4D (Ours)0.44296.110.0010.0540.0510.03399.755.02253.080.0100.2110.1140.05098.80

3D tracking (mm) and geometry. D4RT* is an open-source reproduction.

Ablation: training data

Training dataIn-domain averageHamlyn (OOD)SCOPE-colon-phantom (OOD)
PoseDepthDepthPose
AUC@5°↑AUC@30°↑AbsRel↓δ1.25↑AbsRel↓δ1.25↑RTE↓ARE*↓
No finetune (VGGT-Ω weights)9.5553.030.07794.030.13580.8210.798.22
SC+SIM+C3VD+EM-S+SA35.9974.080.03199.500.14380.9311.287.12
+ StereoMIS33.1070.020.03299.530.14281.4010.516.83
+ EndoMapper-Real36.6274.070.03299.570.12883.0410.377.37
Full dataset36.0273.680.03099.590.12983.7710.717.37

Data composition, geometric fine-tuning only. Hamlyn depth excludes instruments. SC: SCARED, SIM: SimCol3D, EM-S: EndoMapper-Sim, SA: Sano.

Ablation: motion representation

VariantTrackingPoseDepth
EPE↓APD↑AUC@5°↑AUC@30°↑AbsRel↓δ1.25↑
SoM (baseline)5.9446.585.6726.860.05098.82
w/ zero-motion anchor5.3050.116.0927.170.04898.82
Full CRM5.2752.636.9627.660.04998.85

Long-sequence reconstruction

MethodATE↓RTE↓RRE↓
VGGT93.2123.6515.11
VGGT-Ω85.9221.6731.21
SM4RT88.9521.2916.98
Endo3R77.8714.7522.92
SCOPE-4DSFT (Ours)73.5210.3813.38
SCOPE-4D (Ours)71.1710.9413.61

User study

MethodCamera trajectory↓Point cloud↓Depth map↓
VGGT-Ω1.93 [1.65, 2.25]1.95 [1.65, 2.28]1.98 [1.67, 2.32]
SM4RT2.90 [2.60, 3.08]3.22 [2.80, 3.57]3.33 [2.87, 3.73]
Endo3R3.83 [3.53, 4.00]3.52 [3.23, 3.78]3.37 [3.00, 3.70]
SCOPE-4D (Ours)1.33 [1.12, 1.62]1.32 [1.13, 1.55]1.32 [1.13, 1.53]

SCOPE-colon-real: mean rank from six clinical and biomedical experts on ten clips (1 best, 4 worst), 95% bootstrap intervals.

Study interface
User study comparison pageUser study ranking criteria