Register and share your invite link to earn from video plays and referrals.

cv usk
@cv_usk
AI / Software Research Notes AI Agent, LLMOps, MLOps, Software Architecture 投稿は個人の意見です。
Joined May 2026
270 Following    316 Followers
🎬 "The video looks realistic — but did it actually complete the task?" Video generation evaluation finally goes outcome-oriented. SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation 💡 Overview When generating videos that complete tasks like "cook a dish" or "prune a plant" from reference images, current models reproduce procedures reasonably well — but systematically fail to be evaluated on whether the intended end-state was actually achieved. SemComp-Bench introduces a 6-domain, 1,273-instance dataset and a dual-axis evaluation framework powered by Doubao-Seed-1.8 VLM to close this gap. ⚠️ The Problem Existing benchmarks emphasize appearance consistency and intermediate procedural steps, leaving "final outcome achievement" and "semantic grounding in reference images" unevaluated as a joint criterion. 🔬 Evaluation Framework: Two Independent Dimensions · OA (Outcome Achievement) Score: ALL four criteria must pass — outcome realization, semantic grounding, entity consistency, and global visual continuity · GR (Generation Reliability) Score: average across five failure-oriented criteria — physical plausibility, visual clarity, artifact-free rendering, spatiotemporal coherence, text integrity 📊 Results (Detailed Instruction Condition) · OA leader: HunyuanVideo-1.5-720P at 37.8% — the best model still falls short of 40% · GR leader: Seedance 2.0 at 91.8% — yet ranks 4th in OA at just 20.0% · T2V (text-only) OA: only 0.6–5.0%, confirming visual reference is indispensable · Biggest bottleneck: Within-Scene Spatiotemporal Coherence (0.328–0.739) → OA and GR are independent capabilities. Generating polished videos and completing tasks are entirely different problems. #VideoGeneration# #Benchmark#
Show more