Overview
MarineEVT is a comprehensive, event-centric marine video dataset integrating visual tool reasoning, featuring 20K multi-task, video visual question-answering pairs that span 20 dimensions of marine understanding and analysis.
Figure 1: MarineEVT Overview.
Dataset Overview
Figure 2: An example question showcase failure in GPT-5 response.
In marine videos, informative events are often sparse, ephemeral, and unevenly distributed, posing significant challenges for existing VLMs.
Motivated by such demand, we carefully construct an event-centric dataset and benchmark, MarineEVT comprises 20,000 richly annotated underwater video question-answer pairs spanning 20 fine-grained dimensions, including marine species, human activities, environmental conditions, behavioral interactions, and rare ecological events, structured to support semantic, contextualized, spatial-temporal, and causal reasoning.
Semantic Reasoning
This dimension evaluates how models interpret what events happened in a video, focusing on the semantic meaning of entities’ activity, behavior, and attribute characteristics in marine scenarios.
Contextual Reasoning
This dimension evaluates which entities are present in the video event, focusing on object classification and recognition in marine scenario.
Spatial Reasoning
This dimension assesses the model’s capacity to perceive, interpret, and reason about the temporal structures and when- dynamics inherent in marine videos.
Temporal Reasoning
This dimension evaluates a model’s ability to localize and ground where the entities involved in the video event are, assessing its accuracy in identifying their spatial positions, including number, orientation, depth, and movement trajectories, within the underwater scene.
Causal Reasoning
This dimension evaluates a model’s ability to infer causal relationships within marine video data, moving beyond correlation to understand why certain events occur based on observed antecedents or interventions.
Figure 3: Dataset Overview and Dimensional Example Analysis.
Methodology Overview
Propose EVT-R1 decomposing event-centric marine video understanding into a multi-turn visual tool-integrated reasoning process, leveraging powerful visual tools to localize and interpret critical information from redundant video frames with sparse and unevenly distributed events.
Our design enables turn-level RL, guiding the model to produce correct answers and learn when and how to use visual tools effectively. Specifically, we propose a dual-component reward model that corrects final answers while explicitly encouraging effective intermediate tool-use decisions. The dual-reward model consists of: (i) a tool-reasoning reward Rtool that assesses whether invoking a tool was valid and accurate at each step, and (ii) a multi-task answer reward Rans that evaluates the correctness of the final response.
Figure 4: The training and inference processes of EVT-R1.
Figure 5: Showcase of EVT-R1 response.
Experimental Result
We reserve 2,000 question-answer pairs for evaluation. The table below compares performance across five categories: (1) open-source VLMs without tools, (2) closed-source VLMs without tools, (3) closed-source models with tool invocation, (4) open-source models with token compression methods, and (5) open-source models with key-frame selection and event reasoning approaches.
| Model | Training | Tools | SemR. | ConR. | SpaR. | TemR. | CasR. | Avg. |
|---|---|---|---|---|---|---|---|---|
| Open-source VLMs w/o Tools | ||||||||
| Video-LLaVA-7B | ✗ | ✗ | 24.40 | 29.00 | 8.40 | 5.80 | 42.00 | 21.92 |
| LLaVA-NeXT-Video-7B | ✗ | ✗ | 35.40 | 34.33 | 9.00 | 6.75 | 42.67 | 25.63 |
| VideoLLaMA3-7B | ✗ | ✗ | 41.00 | 46.33 | 5.40 | 3.50 | 61.77 | 31.60 |
| InternVL3-8B | ✗ | ✗ | 53.40 | 50.33 | 17.20 | 10.33 | 71.77 | 40.61 |
| Qwen3-VL-8B-Instruct | ✗ | ✗ | 58.20 | 52.77 | 22.40 | 14.00 | 71.00 | 43.67 |
| Closed-source VLMs w/o Tools | ||||||||
| Grok-4.1-FR | ✗ | ✗ | 41.20 | 37.33 | 17.40 | 6.25 | 50.67 | 30.57 |
| Gemini-3.0-Flash | ✗ | ✗ | 48.20 | 43.00 | 27.00 | 7.75 | 62.00 | 37.59 |
| GPT-5-Mini | ✗ | ✗ | 58.40 | 30.67 | 22.80 | 10.00 | 66.67 | 37.71 |
| Closed-source VLMs w/ Tools | ||||||||
| Grok-4.1-FR | ✗ | ✓ | 36.00 | 23.00 | 13.20 | 7.75 | 45.67 | 25.12 |
| Gemini-3.0-Flash | ✗ | ✓ | 47.00 | 43.67 | 25.00 | 6.25 | 59.00 | 36.30 |
| GPT-5-Mini | ✗ | ✓ | 58.40 | 43.67 | 22.40 | 13.00 | 64.33 | 40.35 |
| Open-source VLMs w/ token compression method | ||||||||
| LLaVA-1.5-7B (VisionZip) | ✗ | ✗ | 37.00 | 18.33 | 10.80 | 6.00 | 46.67 | 23.76 |
| Qwen2.5-VL-7B (VisionZip) | ✗ | ✗ | 26.60 | 13.00 | 12.40 | 2.75 | 47.33 | 20.42 |
| InternVL2-8B (PVC) | ✗ | ✗ | 45.60 | 24.67 | 9.40 | 2.50 | 69.00 | 30.23 |
| Open-source VLMs w/ key-frame selection & event-centric method | ||||||||
| MaxInfo | ✗ | ✗ | 48.00 | 23.67 | 19.20 | 11.00 | 50.00 | 30.37 |
| AKS | ✗ | ✗ | 45.60 | 20.33 | 17.20 | 12.50 | 48.00 | 28.72 |
| VideoITG | ✗ | ✗ | 47.40 | 48.67 | 20.20 | 14.25 | 67.33 | 39.57 |
| CoF | ✗ | ✗ | 47.00 | 52.00 | 6.00 | 5.75 | 68.67 | 35.88 |
| Fine-tuning open-source VLMs w/ Tools | ||||||||
| Qwen3-VL-8B (GRPO†) | ✓ | ✓ | 44.60 | 37.00 | 20.00 | 10.75 | 62.66 | 35.58 |
| Qwen3-VL-8B (SFT) | ✓ | ✓ | 61.40 | 53.33 | 22.60 | 15.00 | 74.00 | 45.27 |
| Ours (EVT-R1) | ✓ | ✓ | 65.80 | 53.33 | 30.60 | 20.75 | 74.00 | 48.89 |
Table 1: Experimental comparison between SOTA VLM models and EVT-R1.
Figure 5: Experimental result produced by open-source (general-purpose token- compression), closed-source, and our EVT-R1.
Figure 6: Average accuracy across 5 dimension by open-source (general-purpose token- compression), closed-source, and our EVT-R1.
Citation
If you find our work useful, please consider citing our paper:
@inproceedings{to2026marineevt,
title={MarineEVT: Advancing Event-Centric Marine Video Understanding via Visual Tool Reasoning},
author={To, Tuan-An and Wong, Yuk-Kwan and Vu, Tuan-Anh and Zheng, Ziqiang and Yeung, Sai-Kit},
booktitle={European Conference on Computer Vision (ECCV)},
year={2026}
}