SAM3 vs Astra for text-prompt segmentation
SAM3 gives much more precise masks. it just often fails to understand complicated text prompts.
Astra is great at language. it rarely misses what you want to detect, usually only when the task needs internal domain knowledge. masks come back as JSON, point by point. slower, and less accurate than SAM3
we just rolled out Astra segmentation in playground
link:
↓ more Astra examples below
Astra can do segmentation
this is pure VLM result. no expert models (like SAM) were used
No other VLM even comes close to this quality
- high effort
- avg input tokens / image: 2,052
- avg output tokens / image: 4,685
- avg cost / image: $0.255
- median time / image: 78.3 s
↓ GPT-6 Astra segmentation deep dive