AVE-Compass

Audio-Video Editing Benchmark

A holistic benchmark for instruction-based audio-video editing, built to test whether systems can perform the requested edit while preserving non-target visual and acoustic content.

145source videos
196editing instructions
2,688checklist items
28operation types
AVE-Compass benchmark overview

Benchmark Scope

AVE-Compass treats a video as a coupled audio-visual signal. Its tasks cover explicit cross-modal edits, implicit audio consequences of visual edits, speech edits, and strict single-modality preservation tests.

  • Joint audio-visual editingCoordinated changes across visible events and their associated sounds.
  • Speech editingSentence, speaker, pacing, and speech-related audio-video consistency.
  • Video-only editingVisual changes where the original audio stream should remain faithful.
  • Audio-only editingSound changes where the visual stream should be preserved.
  • Difficulty labelsObject localization, audio-source complexity, and cross-modal linkage annotations support diagnostic analysis.

Editing Instruction Statistics

Editing Taxonomy

4 modalities

Editing Difficulty

3 dimensions

Audio Editing Aspect

Sound types involved in edit targets

Video Statistics

Video Content

Audio Source

Major sound types in source videos

Video Duration

0-10 seconds

Evaluation Signals

The main score on this page is Editing Intent: a compact signal for complete edit execution, balancing whether the instruction was followed and whether non-target content survived.

EI

Editing Intent

Rewards an output only when the requested edit is executed and non-target content is preserved. This is the primary case score shown below.

IF

Instruction Following

Checklist questions verify whether the intended visual, audio, or speech modification appears in the edited result.

FP

Fidelity Preserving

Preservation checks compare source and edited videos for unintended changes to retained content.

REAL

Realism

A separate rubric asks whether the edited clip is natural, coherent, physically plausible, and free of obvious artifacts.

AV Sync Lip Sync Video Aesthetic Subject Consistency Motion Smoothness Audio Aesthetic Speech Quality
AVE-Compass evaluation matrix

Key Findings

Main experimental results on AVE-Compass. Editing Intent is the headline score because it jointly reflects edit execution and non-target preservation.

MLLM-as-Judge Leaderboard

Overall / Video / Audio scores · higher is better

Model Editing Intent Instruction Following Fidelity Preserving Realism
OverallVideoAudio OverallVideoAudio OverallVideoAudio OverallVideoAudio
AVE-Agent (Wan) 159.8+17.4 66.7+6.650.2+25.4 77.3+8.084.8+6.569.4+9.1 77.6+13.480.1+1.674.8+26.7 62.1+1.745.3+1.878.9+1.5
Wan2.7242.460.124.869.378.360.364.278.548.160.443.577.4
HappyHorse341.356.718.866.975.554.563.974.553.363.049.976.0
Gemini-Omni*38.056.110.044.974.010.884.976.896.166.949.284.5
Seedance26.636.113.537.450.524.481.783.080.469.254.084.4
LTX215.210.726.470.172.866.230.624.542.364.142.086.2

* Gemini-Omni misses 16 speech edits due to content moderation. Superscripts on AVE-Agent show changes over the Wan2.7 backbone.

Automatic Metrics7 automated quality metrics · click to expand
Model Cross-Modal Video Audio
Lip SyncAV Sync Video AestheticSubject ConsistencyMotion Smoothness Audio AestheticSpeech Quality
AVE-Agent (Wan)0.622+0.0910.766+0.0730.452+0.0010.972+0.0030.987+0.0010.614+0.0050.368−0.060
Wan2.70.5310.6930.4510.9690.9860.6090.428
HappyHorse0.6200.6950.4390.9750.9890.6500.726
Gemini-Omni*0.7010.4340.9740.9880.627
Seedance0.4310.7180.4300.9710.9870.6290.388
LTX20.6180.7580.4680.9680.9860.6410.522

* Gemini-Omni misses 16 speech edits due to content moderation. Speech Quality and Lip Sync are computed only on speech-category edits.

AVE-Agent is the strongest edit executor.

It improves overall Editing Intent by +17.4 over the Wan2.7 backbone, with especially large gains on audio-side edit completion.

Instruction following and preservation pull against each other.

Some models follow instructions by regenerating too much, while others preserve the source by not editing enough. Editing Intent exposes this tradeoff directly.

Audio-visual realism still lags behind low-level quality.

High aesthetic or signal scores can miss physical-logic failures, temporal mismatches, and unnatural edited regions.

AVE-Agent

AVE-Agent is a modular baseline for complex cross-modal editing. It uses planning, local reflection, and mixed audio-video evaluation to reduce blind single-pass editing failures.

AVE-Agent architecture overview
1

Plan

Analyze the source clip, identify explicit and implicit audio-visual requirements, and decompose the instruction into dependent subtasks.

2

Execute

Route subtasks to visual, audio, or speech editing branches and use evaluator feedback to refine failed local attempts.

3

Evaluate

Inspect the assembled clip for global instruction following, preservation, and audio-video coherence before deciding whether to pass, remix, regenerate, or replan.

Case Gallery

Five representative edit types. Each case pairs the original audio-video with its instruction and tags, followed by scrollable model results reporting only Editing Intent.

Original audio-video
Case 01

Bee around a rabbit

Editing instruction

Introduce the sound of a buzzing bee; show a bee visually flying around the rabbit.

Joint AV J1.1 New Source Insertion Audio + Video Animal subject Outdoor nature Localization · Hard Audio Complexity · Complex AV Linkage · Explicit
Original audio-video
Case 02

Overcast to heavy rain

Editing instruction

Change the overcast scene to a heavy downpour; add visible rain streaks and the immersive sound of heavy rain hitting the street and buildings.

Joint AVJ6.1 Visible WeatherAudio + VideoWeatherUrban streetLocalization · EasyAudio Complexity · ModerateAV Linkage · Explicit
Original audio-video
Case 03

Replace Kibble

Editing instruction

Change the kibble being poured into the bowl to water; replace the hard clatter of dry dog food with the sound of water splashing and filling the bowl.

Joint AVJ4.1 Identity SwapAudio + VideoSpurious object/textObject replacementLocalization · HardAudio Complexity · ComplexAV Linkage · Explicit
Original audio-video
Case 04

Waterfall to winter

Editing instruction

Re-contextualize the lush waterfall scene into an icy, winter environment; add visuals of snow and ice, and replace the forest ambience with a colder, crisp, slightly muffled soundscape, while preserving the water flow.

Joint AVJ6.2 Scene MigrationAudio + VideoNature & waterEnvironmental ambienceLocalization · EasyAudio Complexity · SimpleAV Linkage · Explicit
Original audio-video
Case 05

Remove the Train Conductor

Editing instruction

Remove the train conductor from the second shot, inpainting the background, and silence his speech and the associated mechanical clicking sounds.

Joint AVJ2.1 Source ErasureAudio + VideoTemporal/audio failureMulti-shot sceneLocalization · HardAudio Complexity · ComplexAV Linkage · Explicit

Citation

Dataset and project materials are released for non-commercial research use under CC BY-NC 4.0.

@article{wen2026avecompass,
  title = {AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities},
  author = {Wen, Yuqing and Huang, Yukai and Xie, Qianqian and Wu, Jiangtao and Lin, Yibin and Gu, Yikai and Chen, Jialu and Zhang, Yuanxing and Liu, Jiaheng},
  year = {2026}
}