MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning

Tuan-An To1, Yuk-Kwan Wong1, Tuan-Anh Vu2, Zheng Ziqiang*1,3, Sai-Kit Yeung1

1Hong Kong University of Science and Technology, 2 University of California, Los Angeles, 3 University of Electronic Science and Technology of China

The 19th European Conference on Computer Vision (ECCV) 2026

Paper Supplementary GitHub Hugging Face

Overview

MarineEVT is a comprehensive, event-centric marine video dataset integrating visual tool reasoning, featuring 20K multi-task, video visual question-answering pairs that span 20 dimensions of marine understanding and analysis.

Teaser Image

Figure 1: MarineEVT Overview.

Dataset Overview

Figure 2: An example question showcase failure in GPT-5 response.

In marine videos, informative events are often sparse, ephemeral, and unevenly distributed, posing significant challenges for existing VLMs.

Motivated by such demand, we carefully construct an event-centric dataset and benchmark, MarineEVT comprises 20,000 richly annotated underwater video question-answer pairs spanning 20 fine-grained dimensions, including marine species, human activities, environmental conditions, behavioral interactions, and rare ecological events, structured to support semantic, contextualized, spatial-temporal, and causal reasoning.

Methodology Overview

Propose EVT-R1 decomposing event-centric marine video understanding into a multi-turn visual tool-integrated reasoning process, leveraging powerful visual tools to localize and interpret critical information from redundant video frames with sparse and unevenly distributed events.

Our design enables turn-level RL, guiding the model to produce correct answers and learn when and how to use visual tools effectively. Specifically, we propose a dual-component reward model that corrects final answers while explicitly encouraging effective intermediate tool-use decisions. The dual-reward model consists of: (i) a tool-reasoning reward Rtool that assesses whether invoking a tool was valid and accurate at each step, and (ii) a multi-task answer reward Rans that evaluates the correctness of the final response.

Method Image

Figure 4: The training and inference processes of EVT-R1.

Experimental Result

We reserve 2,000 question-answer pairs for evaluation. The table below compares performance across five categories: (1) open-source VLMs without tools, (2) closed-source VLMs without tools, (3) closed-source models with tool invocation, (4) open-source models with token compression methods, and (5) open-source models with key-frame selection and event reasoning approaches.

Model Training Tools SemR. ConR. SpaR. TemR. CasR. Avg.
Open-source VLMs w/o Tools
Video-LLaVA-7B ✗ ✗ 24.40 29.00 8.40 5.80 42.00 21.92
LLaVA-NeXT-Video-7B ✗ ✗ 35.40 34.33 9.00 6.75 42.67 25.63
VideoLLaMA3-7B ✗ ✗ 41.00 46.33 5.40 3.50 61.77 31.60
InternVL3-8B ✗ ✗ 53.40 50.33 17.20 10.33 71.77 40.61
Qwen3-VL-8B-Instruct ✗ ✗ 58.20 52.77 22.40 14.00 71.00 43.67
Closed-source VLMs w/o Tools
Grok-4.1-FR ✗ ✗ 41.20 37.33 17.40 6.25 50.67 30.57
Gemini-3.0-Flash ✗ ✗ 48.20 43.00 27.00 7.75 62.00 37.59
GPT-5-Mini ✗ ✗ 58.40 30.67 22.80 10.00 66.67 37.71
Closed-source VLMs w/ Tools
Grok-4.1-FR ✗ ✓ 36.00 23.00 13.20 7.75 45.67 25.12
Gemini-3.0-Flash ✗ ✓ 47.00 43.67 25.00 6.25 59.00 36.30
GPT-5-Mini ✗ ✓ 58.40 43.67 22.40 13.00 64.33 40.35
Open-source VLMs w/ token compression method
LLaVA-1.5-7B (VisionZip) ✗ ✗ 37.00 18.33 10.80 6.00 46.67 23.76
Qwen2.5-VL-7B (VisionZip) ✗ ✗ 26.60 13.00 12.40 2.75 47.33 20.42
InternVL2-8B (PVC) ✗ ✗ 45.60 24.67 9.40 2.50 69.00 30.23
Open-source VLMs w/ key-frame selection & event-centric method
MaxInfo ✗ ✗ 48.00 23.67 19.20 11.00 50.00 30.37
AKS ✗ ✗ 45.60 20.33 17.20 12.50 48.00 28.72
VideoITG ✗ ✗ 47.40 48.67 20.20 14.25 67.33 39.57
CoF ✗ ✗ 47.00 52.00 6.00 5.75 68.67 35.88
Fine-tuning open-source VLMs w/ Tools
Qwen3-VL-8B (GRPO†) ✓ ✓ 44.60 37.00 20.00 10.75 62.66 35.58
Qwen3-VL-8B (SFT) ✓ ✓ 61.40 53.33 22.60 15.00 74.00 45.27
Ours (EVT-R1) ✓ ✓ 65.80 53.33 30.60 20.75 74.00 48.89

Table 1: Experimental comparison between SOTA VLM models and EVT-R1.

Result Image

Figure 5: Experimental result produced by open-source (general-purpose token- compression), closed-source, and our EVT-R1.

Description of second image

Figure 6: Average accuracy across 5 dimension by open-source (general-purpose token- compression), closed-source, and our EVT-R1.

Citation

If you find our work useful, please consider citing our paper:

@inproceedings{to2026marineevt,
  title={MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning},
  author={To, Tuan-An and Wong, Yuk-Kwan and Vu, Tuan-Anh and Zheng, Ziqiang and Yeung, Sai-Kit},
  booktitle={European Conference on Computer Vision (ECCV)},
  year={2026}
}