Benchmark Scope
AVE-Compass treats a video as a coupled audio-visual signal. Its tasks cover explicit cross-modal edits, implicit audio consequences of visual edits, speech edits, and strict single-modality preservation tests.
- Joint audio-visual editingCoordinated changes across visible events and their associated sounds.
- Speech editingSentence, speaker, pacing, and speech-related audio-video consistency.
- Video-only editingVisual changes where the original audio stream should remain faithful.
- Audio-only editingSound changes where the visual stream should be preserved.
- Difficulty labelsObject localization, audio-source complexity, and cross-modal linkage annotations support diagnostic analysis.
Editing Instruction Statistics
Editing Difficulty
3 dimensionsAudio Editing Aspect
Sound types involved in edit targetsVideo Statistics
Video Content
Audio Source
Major sound types in source videosVideo Duration
0-10 secondsEvaluation Signals
The main score on this page is Editing Intent: a compact signal for complete edit execution, balancing whether the instruction was followed and whether non-target content survived.
Editing Intent
Rewards an output only when the requested edit is executed and non-target content is preserved. This is the primary case score shown below.
Instruction Following
Checklist questions verify whether the intended visual, audio, or speech modification appears in the edited result.
Fidelity Preserving
Preservation checks compare source and edited videos for unintended changes to retained content.
Realism
A separate rubric asks whether the edited clip is natural, coherent, physically plausible, and free of obvious artifacts.
Key Findings
Main experimental results on AVE-Compass. Editing Intent is the headline score because it jointly reflects edit execution and non-target preservation.
MLLM-as-Judge Leaderboard
Overall / Video / Audio scores · higher is better
| Model | Editing Intent | Instruction Following | Fidelity Preserving | Realism | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Overall | Video | Audio | Overall | Video | Audio | Overall | Video | Audio | Overall | Video | Audio | |
| AVE-Agent (Wan) | 159.8+17.4 | 66.7+6.6 | 50.2+25.4 | 77.3+8.0 | 84.8+6.5 | 69.4+9.1 | 77.6+13.4 | 80.1+1.6 | 74.8+26.7 | 62.1+1.7 | 45.3+1.8 | 78.9+1.5 |
| Wan2.7 | 242.4 | 60.1 | 24.8 | 69.3 | 78.3 | 60.3 | 64.2 | 78.5 | 48.1 | 60.4 | 43.5 | 77.4 |
| HappyHorse | 341.3 | 56.7 | 18.8 | 66.9 | 75.5 | 54.5 | 63.9 | 74.5 | 53.3 | 63.0 | 49.9 | 76.0 |
| Gemini-Omni* | 38.0 | 56.1 | 10.0 | 44.9 | 74.0 | 10.8 | 84.9 | 76.8 | 96.1 | 66.9 | 49.2 | 84.5 |
| Seedance | 26.6 | 36.1 | 13.5 | 37.4 | 50.5 | 24.4 | 81.7 | 83.0 | 80.4 | 69.2 | 54.0 | 84.4 |
| LTX2 | 15.2 | 10.7 | 26.4 | 70.1 | 72.8 | 66.2 | 30.6 | 24.5 | 42.3 | 64.1 | 42.0 | 86.2 |
* Gemini-Omni misses 16 speech edits due to content moderation. Superscripts on AVE-Agent show changes over the Wan2.7 backbone.
Automatic Metrics7 automated quality metrics · click to expand
| Model | Cross-Modal | Video | Audio | ||||
|---|---|---|---|---|---|---|---|
| Lip Sync† | AV Sync | Video Aesthetic | Subject Consistency | Motion Smoothness | Audio Aesthetic | Speech Quality† | |
| AVE-Agent (Wan) | 0.622+0.091 | 0.766+0.073 | 0.452+0.001 | 0.972+0.003 | 0.987+0.001 | 0.614+0.005 | 0.368−0.060 |
| Wan2.7 | 0.531 | 0.693 | 0.451 | 0.969 | 0.986 | 0.609 | 0.428 |
| HappyHorse | 0.620 | 0.695 | 0.439 | 0.975 | 0.989 | 0.650 | 0.726 |
| Gemini-Omni* | — | 0.701 | 0.434 | 0.974 | 0.988 | 0.627 | — |
| Seedance | 0.431 | 0.718 | 0.430 | 0.971 | 0.987 | 0.629 | 0.388 |
| LTX2 | 0.618 | 0.758 | 0.468 | 0.968 | 0.986 | 0.641 | 0.522 |
* Gemini-Omni misses 16 speech edits due to content moderation. † Speech Quality and Lip Sync are computed only on speech-category edits.
It improves overall Editing Intent by +17.4 over the Wan2.7 backbone, with especially large gains on audio-side edit completion.
Some models follow instructions by regenerating too much, while others preserve the source by not editing enough. Editing Intent exposes this tradeoff directly.
High aesthetic or signal scores can miss physical-logic failures, temporal mismatches, and unnatural edited regions.
AVE-Agent
AVE-Agent is a modular baseline for complex cross-modal editing. It uses planning, local reflection, and mixed audio-video evaluation to reduce blind single-pass editing failures.
Plan
Analyze the source clip, identify explicit and implicit audio-visual requirements, and decompose the instruction into dependent subtasks.
Execute
Route subtasks to visual, audio, or speech editing branches and use evaluator feedback to refine failed local attempts.
Evaluate
Inspect the assembled clip for global instruction following, preservation, and audio-video coherence before deciding whether to pass, remix, regenerate, or replan.
Case Gallery
Five representative edit types. Each case pairs the original audio-video with its instruction and tags, followed by scrollable model results reporting only Editing Intent.
Edited results
Observed issueThe original BGM is not preserved.
Observed issueNo bee buzzing sound is added.
Observed issueThe original BGM is not preserved.
Observed issueThe entire source video is regenerated.
Edited results
Observed issueOverly exaggerated effects alter the street, which should remain unchanged.
Observed issueThe output is almost unedited.
Observed issueThe audio remains unchanged.
Observed issueThe output is almost unedited.
Observed issueThe entire source video is regenerated.
Edited results
Observed issueThe original BGM changes unexpectedly.
Observed issueThe audio is almost unedited.
Observed issueAbnormal subtitles appear in the edited video.
Observed issueThe requested edit is not executed at all; instead, an unrequested voice saying "no" is added.
Observed issueThe entire source video is regenerated.
Edited results
Observed issueThe entire source video is regenerated.
Edited results
Observed issueThe original BGM and ambient sounds are not preserved.
Observed issueThe original BGM and ambient sounds are not preserved.
Observed issueThe audio remains unchanged, and abnormal events and people briefly flash into the video.
Observed issueThe output is completely unedited.
Observed issueThe entire source video is regenerated.
Citation
Dataset and project materials are released for non-commercial research use under CC BY-NC 4.0.
@article{wen2026avecompass,
title = {AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities},
author = {Wen, Yuqing and Huang, Yukai and Xie, Qianqian and Wu, Jiangtao and Lin, Yibin and Gu, Yikai and Chen, Jialu and Zhang, Yuanxing and Liu, Jiaheng},
year = {2026}
}