TL;DR An AI that writes code to generate images or video can run that code successfully while still failing to meet what the visuals actually need to look like. This paper proposes a way to measure that "program-to-visual" gap.
Title: MaLiang-Harness: A Programmable Path to Image and Video Generation
URL:
Points
🖼️ A stateful framework that inspects and revises generated output while preserving state across edits
🔧 A traceable generation process links each code change to its rendered visual outcome
🎨 Unifies Canvas, SVG, Scene2d, and Three.js backends under one shared protocol
📊 On 50 text-to-image tasks, GPT-6-Astra hits 100% generation success and 96.0% full quality compliance
🎬 On 13 text-to-video tasks, GPT-6-Astra again reaches 100% success and 76.9% full quality compliance
⚠️ DeepSeek-class models manage only 12-40% on images and 0% on video
🔍 Two models with identical general-capability scores diverge sharply, 44% vs 88% on drawing quality pass rate
The point that really lands: a general benchmark score alone doesn't tell you how good a model actually is at this.
#
MultimodalAI# #
ImageGeneration#