注册并分享邀请链接,可获得视频播放与邀请奖励。

AI at Meta
@AIatMeta
Together with the AI community, we are pushing the boundaries of what’s possible through open science to create a more connected world.
加入 August 2018
367 正在关注    856.1K 粉丝
Muse Spark 1.2 supports a broad range of multimodal tasks, from turning visuals into working code to translating perception into physical action. It also brings robust audio-visual understanding to enable video-heavy workflows common in real-world enterprise use. Today, we’re sharing new evals and demos that illustrate the breadth of the model’s visual understanding and reasoning capabilities. Let’s start with a demo that shows how Muse Spark parses multimodal observations and calls tools to guide a robot to navigate in an unstructured environment to find a rubber duck. 🧵👇
显示更多
0
42
516
72
转发到社区