MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning

Tuan-An To1, Yuk-Kwan Wong1, Tuan-Anh Vu2, Zheng Ziqiang*1,3, Sai-Kit Yeung1

1Hong Kong University of Science and Technology, 2 University of California, Los Angeles, 3 University of Electronic Science and Technology of China

The 19th European Conference on Computer Vision (ECCV) 2026

Paper Supplementary GitHub Hugging Face

Overview

MarineEVT is a comprehensive, event-centric marine video dataset integrating visual tool reasoning, featuring 20K multi-task, video visual question-answering pairs that span 20 dimensions of marine understanding and analysis.

Teaser Image

Figure 1: MarineEVT Overview.

Dataset Overview

Figure 2: An example question showcase failure in GPT-5 response.

In marine videos, informative events are often sparse, ephemeral, and unevenly distributed, posing significant challenges for existing VLMs.

Motivated by such demand, we carefully construct an event-centric dataset and benchmark, MarineEVT comprises 20,000 richly annotated underwater video question-answer pairs spanning 20 fine-grained dimensions, including marine species, human activities, environmental conditions, behavioral interactions, and rare ecological events, structured to support semantic, contextualized, spatial-temporal, and causal reasoning.

Methodology Overview

Propose EVT-R1 decomposing event-centric marine video understanding into a multi-turn visual tool-integrated reasoning process, leveraging powerful visual tools to localize and interpret critical information from redundant video frames with sparse and unevenly distributed events.

Our design enables turn-level RL, guiding the model to produce correct answers and learn when and how to use visual tools effectively. Specifically, we propose a dual-component reward model that corrects final answers while explicitly encouraging effective intermediate tool-use decisions. The dual-reward model consists of: (i) a tool-reasoning reward Rtool that assesses whether invoking a tool was valid and accurate at each step, and (ii) a multi-task answer reward Rans that evaluates the correctness of the final response.

Method Image

Figure 4: The training and inference processes of EVT-R1.

Experimental Result

We reserve 2,000 question-answer pairs for evaluation. The table below compares performance across five categories: (1) open-source VLMs without tools, (2) closed-source VLMs without tools, (3) closed-source models with tool invocation, (4) open-source models with token compression methods, and (5) open-source models with key-frame selection and event reasoning approaches.

Model Training Tools SemR. ConR. SpaR. TemR. CasR. Avg.
Open-source VLMs w/o Tools
Video-LLaVA-7B 24.40 29.00 8.40 5.80 42.00 21.92
LLaVA-NeXT-Video-7B 35.40 34.33 9.00 6.75 42.67 25.63
VideoLLaMA3-7B 41.00 46.33 5.40 3.50 61.77 31.60
InternVL3-8B 53.40 50.33 17.20 10.33 71.77 40.61
Qwen3-VL-8B-Instruct 58.20 52.77 22.40 14.00 71.00 43.67
Closed-source VLMs w/o Tools
Grok-4.1-FR 41.20 37.33 17.40 6.25 50.67 30.57
Gemini-3.0-Flash 48.20 43.00 27.00 7.75 62.00 37.59
GPT-5-Mini 58.40 30.67 22.80 10.00 66.67 37.71
Closed-source VLMs w/ Tools
Grok-4.1-FR 36.00 23.00 13.20 7.75 45.67 25.12
Gemini-3.0-Flash 47.00 43.67 25.00 6.25 59.00 36.30
GPT-5-Mini 58.40 43.67 22.40 13.00 64.33 40.35
Open-source VLMs w/ token compression method
LLaVA-1.5-7B (VisionZip) 37.00 18.33 10.80 6.00 46.67 23.76
Qwen2.5-VL-7B (VisionZip) 26.60 13.00 12.40 2.75 47.33 20.42
InternVL2-8B (PVC) 45.60 24.67 9.40 2.50 69.00 30.23
Open-source VLMs w/ key-frame selection & event-centric method
MaxInfo 48.00 23.67 19.20 11.00 50.00 30.37
AKS 45.60 20.33 17.20 12.50 48.00 28.72
VideoITG 47.40 48.67 20.20 14.25 67.33 39.57
CoF 47.00 52.00 6.00 5.75 68.67 35.88
Fine-tuning open-source VLMs w/ Tools
Qwen3-VL-8B (GRPO†) 44.60 37.00 20.00 10.75 62.66 35.58
Qwen3-VL-8B (SFT) 61.40 53.33 22.60 15.00 74.00 45.27
Ours (EVT-R1) 65.80 53.33 30.60 20.75 74.00 48.89

Table 1: Experimental comparison between SOTA VLM models and EVT-R1.

Result Image

Figure 5: Experimental result produced by open-source (general-purpose token- compression), closed-source, and our EVT-R1.

Description of second image

Figure 6: Average accuracy across 5 dimension by open-source (general-purpose token- compression), closed-source, and our EVT-R1.

Citation

If you find our work useful, please consider citing our paper:

@inproceedings{to2026marineevt,
  title={MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning},
  author={To, Tuan-An and Wong, Yuk-Kwan and Vu, Tuan-Anh and Zheng, Ziqiang and Yeung, Sai-Kit},
  booktitle={European Conference on Computer Vision (ECCV)},
  year={2026}
}