Repository navigation
[Feature Request] Multimodal Image Embedder #184
Description
Activity
🔍 Root Cause Analysis:
The issue in [Feature Request] Multimodal Image Embedder highlights a common synchronization gap between the vector embedding generation pipeline and the metadata ingestion store.
1. Architectural Diagnosis:
- Silent Degradation on Embedding Failure: When an upstream embedding provider (e.g., OpenAI, HuggingFace, Ollama) experiences rate limiting or an unconfigured model alias, the ingestion endpoint stores the raw payload while setting dense vector dimensions to
0ornull. - Dimension Inconsistency: Embedding vectors produced with different dimensions or normalization (L2-normalized vs unnormalized cosine) lead to degraded ranking metrics or engine-level vector dimension mismatch errors in FAISS, Chroma, Milvus, and Qdrant.
- Empty / Malformed Chunks: Ingestion pipelines that do not filter out empty strings, excessive whitespace, or binary tokens cause embedding models to return NaN or zero vectors.
2. Recommended Robust Pattern:
Implement strict schema validation, dimension invariance checks, and transactional ingestion semantics:
from typing import List, Optional from pydantic import BaseModel, Field, model_validator import numpy as np class VectorDocument(BaseModel): id: str text: str embedding: List[float] = Field(..., min_length=1) expected_dimension: int = 1536 @model_validator(mode="after") def validate_embedding_vector(self): # 1. Validate vector length matches expected dimension if len(self.embedding) != self.expected_dimension: raise ValueError( f"Embedding size {len(self.embedding)} does not match expected dimension {self.expected_dimension}" ) # 2. Guard against NaN or infinite float values arr = np.array(self.embedding, dtype=np.float32) if not np.all(np.isfinite(arr)): raise ValueError("Embedding vector contains NaN or infinite values") # 3. Guard against zero vectors if np.all(arr == 0): raise ValueError("Embedding vector is an invalid zero vector") return self
Ensure ingestion rejects un-embedded payloads with HTTP 422 instead of returning an unqualified
200 OK.
Eng. MHD. Shadi AL-Hasan
Executive CTO & Enterprise Solutions Architect
GitHub Profile | Contact & Portfolio- Silent Degradation on Embedding Failure: When an upstream embedding provider (e.g., OpenAI, HuggingFace, Ollama) experiences rate limiting or an unconfigured model alias, the ingestion endpoint stores the raw payload while setting dense vector dimensions to
For the time being, no extensions will be made to the text-and-image capabilities associated with
llama_batch. This is because the underlying implementation ofllama_batch_extandllama_processis currently undergoing a revamp; the new processing approach appears unstable and is subject to constant changes and revisions, so we will consider refactoring this part only after things have stabilized.
Hello,
Could you add support for visual embedders, such as
Qwen3-VL-Embeddinghttps://huggingface.co/mradermacher/Qwen3-VL-Embedding-8B-i1-GGUF
Use Case:
This is required for a multimodal RAG database. Text fragments are indexed in the standard way, but I also need to index images and perform searches across them. The Qwen3-VL-Embedding model projects both modalities into a shared 4096-dimensional space, enabling text queries to retrieve relevant images, and vice versa.
I couldn't do it using the existing capabilities (a standard embedder does not support images and previous attempt #66 was closed), so I wrote an additional class
MTMDImageEmbedder.I tested this solution on a RAG system, and it works.
Implementation example:
llama_multimodal.py
For embedding models, there is only one correct message format, ending with a specific sequence and the last token. Variations are not needed here. Therefore, I opted for simple string construction instead of the Jinja templating engine.
Use example
Thank you for supporting this project.