๐ฌ "The video looks realistic โ but did it actually complete the task?" Video generation evaluation finally goes outcome-oriented.
SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation
๐ก Overview
When generating videos that complete tasks like "cook a dish" or "prune a plant" from reference images, current models reproduce procedures reasonably well โ but systematically fail to be evaluated on whether the intended end-state was actually achieved. SemComp-Bench introduces a 6-domain, 1,273-instance dataset and a dual-axis evaluation framework powered by Doubao-Seed-1.8 VLM to close this gap.
โ ๏ธ The Problem
Existing benchmarks emphasize appearance consistency and intermediate procedural steps, leaving "final outcome achievement" and "semantic grounding in reference images" unevaluated as a joint criterion.
๐ฌ Evaluation Framework: Two Independent Dimensions
ยท OA (Outcome Achievement) Score: ALL four criteria must pass โ outcome realization, semantic grounding, entity consistency, and global visual continuity
ยท GR (Generation Reliability) Score: average across five failure-oriented criteria โ physical plausibility, visual clarity, artifact-free rendering, spatiotemporal coherence, text integrity
๐ Results (Detailed Instruction Condition)
ยท OA leader: HunyuanVideo-1.5-720P at 37.8% โ the best model still falls short of 40%
ยท GR leader: Seedance 2.0 at 91.8% โ yet ranks 4th in OA at just 20.0%
ยท T2V (text-only) OA: only 0.6โ5.0%, confirming visual reference is indispensable
ยท Biggest bottleneck: Within-Scene Spatiotemporal Coherence (0.328โ0.739)
โ OA and GR are independent capabilities. Generating polished videos and completing tasks are entirely different problems.
#
VideoGeneration# #
Benchmark#