DE²TR: Dual Evidence Detection Transformer
for Video Temporal Grounding
1School of Information and Communication Engineering  ·  2School of Software Engineering
Xi’an Jiaotong University, Xi’an, China
  ECCV 2026  ·  September 8–12  ·  Malmö, Sweden

Ground truth Predictions Top-5 frames

How to use: click a GT, a predicted moment or a top-5 saliency frame to play it — playback stops automatically at the segment end.

Overview of the DE²TR framework

Overview of DE²TR. (a) Modal fusion and interaction generate saliency and boundary scores as structural priors; guided by these priors, the DEFormer decouples the probing of semantic and boundary evidence under dual evidence alignment losses. (b) The DEFormer layer: a semantic evidence branch aggregates salient visual context, while a parallel boundary evidence branch captures fine-grained cues and iteratively refines the temporal reference anchors.

Abstract

Video temporal grounding (VTG) aims to localize relevant moments and predict saliency scores in untrimmed videos based on natural language queries. Recently, DETR-inspired methods have achieved instance-level moment predictions through query-based semantic evidence aggregation. However, unlike spatial objects with clear physical edges, video moments often exhibit ambiguous temporal boundaries due to the semantic gradient. Consequently, relying solely on global semantics can lead to severe boundary misalignment, yielding sub-optimal grounding results.

To address this issue, we propose a novel Dual Evidence Detection Transformer (DE²TR) framework for better boundary localization. In particular, we introduce a dual-branch decoder that simultaneously captures global semantic evidence and fine-grained boundary cues. To prevent feature conflict between these distinct objectives, we channel-wise decouple the moment queries, facilitating independent evidence learning under the supervision of evidence alignment losses. Building upon this dual-branch design, we introduce a prior-guided refinement mechanism that employs explicit prior maps to steer queries toward potential regions of interest. Furthermore, a boundary augmentation strategy synthesizes ideal boundary patterns via simple temporal splicing, effectively enhancing the model’s boundary awareness. Extensive experiments on multiple popular benchmarks demonstrate that DE²TR significantly outperforms state-of-the-art approaches.

Moments Are Not Objects

Objects in images have clear physical edges — detectors can localize them from appearance alone. Video moments, in contrast, exhibit a semantic gradient: the relevant content fades in and out gradually, so global semantics alone cannot tell where a moment starts or ends. DETR-style grounding aggregates only global semantic evidence, causing severe boundary misalignment. We verify this with a diagnostic experiment: splicing the ground-truth moments into irrelevant background clips creates synthetic videos with hard boundaries — once the boundary signal is clear, +21.2 mAP is recovered for TR-DETR (QD-DETR: +19.6). Yet models trained on synthetic data remain at baseline level on real videos: more data does not help, the boundary signal does.

Objects in images have clear physical edges

(a) Objects in image — clear physical edges ✓

Moments in video exhibit semantic gradient

(b) Real moment in video — semantic gradient, ambiguous boundaries ✗

Synthetic videos provide clear boundary priors

(c) Synthetic video — GT spliced into irrelevant clips, hard boundaries

Diagnostic experiment results

(d) Diagnostic experiments with QD-DETR and TR-DETR: clear boundary priors unlock the performance, while synthetic-trained models stay at baseline level on real data.

Experiments and Results

QVHighlights — Moment Retrieval & Highlight Detection (%)

SlowFast + CLIP backbones. Bold: best · underline: second · †: audio modality · ‡: data augmentation.

Method testval
MRHDMRHD
R1mAP≥ Very GoodR1mAP≥ Very Good
@0.5@0.7@0.5@0.75Avg.mAPHIT@1 @0.5@0.7@0.5@0.75Avg.mAPHIT@1
M-DETR NeurIPS'2152.8933.0254.8229.4030.7335.6955.6053.9434.8432.2035.6555.55
UMT† CVPR'2256.2341.1853.8337.0136.1238.1859.9960.2644.2656.7039.9038.5939.8564.19
QD-DETR CVPR'2362.4044.9862.5239.8839.8638.9462.4062.6846.6662.2341.8241.2239.1363.03
MomentDiff NeurIPS'2357.4239.6654.0235.7335.95
UniVTG ICCV'2358.8640.8657.6035.5935.4738.2060.9636.1338.83
CG-DETR arXiv'2365.4048.4064.5042.8042.9040.3066.2067.4052.1065.6045.7044.9040.8066.70
TR-DETR AAAI'2464.6648.9663.9843.7342.6239.9163.4267.1051.4866.2746.4245.0940.5564.77
UVCOM CVPR'2463.5547.4763.3742.6743.1839.7464.2065.1051.8145.7940.0363.29
BAM-DETR ECCV'2462.7148.6464.5746.3345.3665.1051.6165.4148.5647.61
TD-DETR‡ ICCV'2564.5350.3766.2147.3246.6965.8853.6766.4349.8649.05
KDA ICCV'2566.7050.8867.5746.3145.6769.1153.4668.1748.0447.41
MS-DETR ACMMM'2564.7248.7766.4144.9144.8940.4565.9566.9051.6867.1946.0546.0040.5766.58
DE2TR67.2451.3967.9249.8348.7441.0866.3468.1353.8168.8951.6150.6740.9366.65
DE2TR‡67.9852.4668.3651.2449.8740.5166.7569.8155.2970.5253.4352.1241.2066.97

vs. previous best (TD-DETR): +2.05 Avg mAP on test; with BAS, gains widen to +3.18 Avg mAP, +2.09 R1@0.7 and +3.92 mAP@0.75.

Charades-STA & TACoS — Moment Retrieval (%)

Method Charades-STATACoS
R1@0.3R1@0.5R1@0.7mIoU R1@0.3R1@0.5R1@0.7mIoU
M-DETR NeurIPS'2165.8352.0730.5945.5437.9724.6711.9725.49
QD-DETR CVPR'2357.3132.55
MomentDiff NeurIPS'2355.5732.4244.7833.68
UniVTG ICCV'2370.8158.0135.6550.1051.4434.9717.3533.60
UVCOM CVPR'2459.2536.6436.3923.32
LLMEPET ACMMM'2470.9136.4950.2552.7322.7836.55
DualGround NeurIPS'2561.1138.52
MS-DETR ACMMM'2571.3459.6236.4850.5953.1639.6523.4237.01
FlashVTG WACV'2560.1138.0153.7141.7624.7437.61
KDA ICCV'2560.2337.6340.1324.34
DE2TR72.3160.3939.8751.5856.2942.5227.1439.76

At strict IoU thresholds: +1.35 R1@0.7 on Charades-STA and +2.40 R1@0.7 on TACoS.

TVSum — Highlight Detection Generalization (Top-5 mAP, %)

MethodVTVUGAMSPKPRFMBKBTDSAvg.
TCG ICCV'2185.071.481.978.680.275.571.677.378.668.176.8
UMT CVPR'2287.581.588.278.881.487.076.086.984.479.683.1
QD-DETR CVPR'2388.287.485.685.085.886.976.491.389.273.785.0
UniVTG ICCV'2383.985.189.080.184.681.470.991.773.569.381.0
TR-DETR AAAI'2489.393.094.385.188.088.680.491.389.581.688.1
UVCOM CVPR'2487.691.691.486.786.986.976.992.387.475.686.3
R2-Tuning ECCV'2485.085.991.081.788.887.478.189.290.374.785.2
MS-DETR ACMMM'2589.893.394.888.788.589.080.594.088.776.988.4
MQVTG ICCV'2587.791.692.385.285.791.378.596.590.682.988.2
DE2TR91.794.195.488.989.590.481.292.391.680.989.6

Best on 7 / 10 categories · +1.20 Avg over the previous best.

Effect of Each Component

DEFbeaseaPsPbBASR1mAP
@0.5@0.7Avg.
65.6849.8745.63
65.2350.5246.30
66.7150.9748.08
67.0349.3547.69
66.8452.1349.24
67.4252.3949.74
66.7752.7149.69
68.1353.8150.67
69.8155.2952.12

Component Build-Up — Avg mAP

Baseline
45.63
+ DEFormer
46.30+0.67
+ ℒsea, ℒbea
49.24+2.94
+ priors Ps, Pb
50.67+1.43
+ BAS = DE2TR‡
52.12+1.45
total gain +6.49 — every piece helps; losses contribute most, BAS is a free +1.45

Backbone Robustness

MethodBackboneR1mAP
@0.5@0.7Avg.
TD-DETR ICCV'25SF+C65.8853.6749.05
DE2TRSF+C68.1353.8150.67
TD-DETR ICCV'25IV271.1955.8151.72
DE2TRIV271.8457.6853.91

Backbone-agnostic: SOTA with SlowFast+CLIP and InternVideo2 (IV2).

Query Strategy for Heads

MethodR1mAP
@0.5@0.7Avg.
Merged66.7152.9049.53
Separate68.1353.8150.67

Routing decoupled queries to their own heads keeps task–feature alignment strict.

Prior Map Injection

Sem.Bnd.R1mAP
BranchBranch@0.5@0.7Avg.
memorymemory67.7452.9750.14
queryquery68.9753.3450.32
memoryquery67.5152.0549.86
querymemory68.1353.8150.67

Semantic → query (global intent) · boundary → memory (local cues) works best.

Number of Queries N

NR1mAP
@0.5@0.7Avg.
567.6853.4848.56
1068.1353.8150.67
1568.5253.3551.22
2069.3554.0051.19

Performance saturates as N grows; N = 10 kept for fair comparison.

Robustness on High-Ambiguity Splits

Subsets of QVHighlights val filtered by the ambiguity coefficient δk — top 20% (Extreme), 50% (Hard), 80% (Moderate) most ambiguous samples. Competitors reproduced from official codes.

Method Extreme (20%)Hard (50%)Moderate (80%)All (100%)
R1@0.5R1@0.7mAPR1@0.5R1@0.7mAPR1@0.5R1@0.7mAPR1@0.5R1@0.7mAP
QD-DETR CVPR'2356.1339.0334.8058.7140.5235.4862.7445.0839.5263.8747.7442.45
TR-DETR AAAI'2451.2933.5532.2456.5740.7735.8262.9047.5040.9065.1650.3944.38
BAM-DETR ECCV'2451.9438.7138.1056.9042.1939.9862.5847.9044.8363.6851.3747.84
TD-DETR ICCV'2554.5240.6538.8258.1943.6139.9463.7149.0344.9265.7452.9748.83
Ours60.0041.3240.4062.0645.6842.6467.2651.6948.0968.1353.8150.67
Ours w/ BAS61.6142.1341.2563.4847.2643.7968.3152.7449.0269.8155.2952.12

SOTA on every ambiguity level — DE2TR mitigates inherent boundary ambiguity rather than overfitting to sharp transitions.

Qualitative Results

The semantic query Cs attends to the whole moment, while the boundary queries Cb,s, Cb,e spike exactly at the edges. Compared with recent advanced methods, DE²TR keeps semantic alignment while localizing sharp, well-aligned boundaries. More qualitative results on the QVHighlights val split are included from the supplementary material.

  Click any image to view it full-size.

Poster

BibTeX

@inproceedings{zhang2026de2tr,
  title     = {{DE$^2$TR}: Dual Evidence Detection Transformer for Video Temporal Grounding},
  author    = {Zhang, Yifan and Liu, Chengxu and Dun, Yujie and Qian, Xueming},
  booktitle = {European Conference on Computer Vision (ECCV)},
  year      = {2026}
}