How can a text embedding model map audio and image vectors?
After being transformed via a projector (translator), a foreign modality vector must announce itself before entering an embedding model for an unrelated modality.
Think of HTML.
These tags let the document know this element is a paragraph.
Vectors work the same way.