🎬 "The video looks realistic — but did it actually complete the task?" Video generation evaluation finally goes outcome-oriented.
SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation
💡 Overview
When generating videos that complete tasks like "cook a dish" or "prune a plant" from reference images, current models reproduce procedures reasonably well — but systematically fail to be evaluated on whether the intended end-state was actually achieved. SemComp-Bench introduces a 6-domain, 1,273-instance dataset and a dual-axis evaluation framework powered by Doubao-Seed-1.8 VLM to close this gap.
⚠️ The Problem
Existing benchmarks emphasize appearance consistency and intermediate procedural steps, leaving "final outcome achievement" and "semantic grounding in reference images" unevaluated as a joint criterion.
🔬 Evaluation Framework: Two Independent Dimensions
· OA (Outcome Achievement) Score: ALL four criteria must pass — outcome realization, semantic grounding, entity consistency, and global visual continuity
· GR (Generation Reliability) Score: average across five failure-oriented criteria — physical plausibility, visual clarity, artifact-free rendering, spatiotemporal coherence, text integrity
📊 Results (Detailed Instruction Condition)
· OA leader: HunyuanVideo-1.5-720P at 37.8% — the best model still falls short of 40%
· GR leader: Seedance 2.0 at 91.8% — yet ranks 4th in OA at just 20.0%
· T2V (text-only) OA: only 0.6–5.0%, confirming visual reference is indispensable
· Biggest bottleneck: Within-Scene Spatiotemporal Coherence (0.328–0.739)
→ OA and GR are independent capabilities. Generating polished videos and completing tasks are entirely different problems.
#
VideoGeneration# #
Benchmark#