๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

cv usk
@cv_usk
AI / Software Research Notes AI Agent, LLMOps, MLOps, Software Architecture ๆŠ•็จฟใฏๅ€‹ไบบใฎๆ„่ฆ‹ใงใ™ใ€‚
๊ฐ€์ž… May 2026
280 ํŒ”๋กœ์ž‰ ์ค‘    417 ํŒฌ
TL;DR An AI that writes code to generate images or video can run that code successfully while still failing to meet what the visuals actually need to look like. This paper proposes a way to measure that "program-to-visual" gap. Title: MaLiang-Harness: A Programmable Path to Image and Video Generation URL: Points ๐Ÿ–ผ๏ธ A stateful framework that inspects and revises generated output while preserving state across edits ๐Ÿ”ง A traceable generation process links each code change to its rendered visual outcome ๐ŸŽจ Unifies Canvas, SVG, Scene2d, and Three.js backends under one shared protocol ๐Ÿ“Š On 50 text-to-image tasks, GPT-6-Astra hits 100% generation success and 96.0% full quality compliance ๐ŸŽฌ On 13 text-to-video tasks, GPT-6-Astra again reaches 100% success and 76.9% full quality compliance โš ๏ธ DeepSeek-class models manage only 12-40% on images and 0% on video ๐Ÿ” Two models with identical general-capability scores diverge sharply, 44% vs 88% on drawing quality pass rate The point that really lands: a general benchmark score alone doesn't tell you how good a model actually is at this. #MultimodalAI# #ImageGeneration#
๋” ๋ณด๊ธฐ