Amazon / Nova / Technical Report
Amazon Nova Multimodal Embeddings: Technical Report and Model Card
Source summary
Original wording · Original languageAbstract · Page 1, 2
We present Amazon Nova Multimodal Embeddings (MME), a state-of-the-art multimodal embedding model for agentic RAG and semantic search applications. Nova MME is the first embeddings model that supports five modalities as input: text, documents, images, video and audio, and transforms them into a single, unified embedding space. This powerful capability enables cross-modal retrieval — allowing users to search and find relevant information across different types of data. Nova MME supports up to 8K context length in the forms of text, images, video, documents and audio, converting them into numerical representations known as embeddings. These embeddings capture the semantic meaning of the underlying content, making it possible to compare, search, and perform reasoning tasks across modalities. By calculating the distance between embeddings, customers can power a wide range of intelligent applications — from semantic search and RAG-powered Large Language Models (LLMs) to content classification and beyond. It is the first unified embedding model that supports text, documents, images, video, and audio through a single model, giving developers the flexibility and performance needed to build next-generation AI solutions.
Enterprise API
Text
Image Document Video Audio
Google Gemini Embedding [1]
Google Vertex Multimodal Embeddings [2]
Amazon Titan Text Embeddings [3]
Amazon Titan MM Embeddings [4]
Cohere Embed 4 [5]
TwelveLabs Marengo 2.7 [6]
TwelveLabs Marengo 3.0 [7]
Amazon Nova MME
✓
✓
✓
✓
✓
✓
✓
✓
✓
✓
✓
✓
✓
✓
✓
✓
✓
✓
✓
✓
Core figures
Enlarge to explore. Download the original for full detail.