Skip to content

[Feature Request] Multimodal Image Embedder #184

Description

@KLL535

Hello,
Could you add support for visual embedders, such as Qwen3-VL-Embedding

https://huggingface.co/mradermacher/Qwen3-VL-Embedding-8B-i1-GGUF

Use Case:

This is required for a multimodal RAG database. Text fragments are indexed in the standard way, but I also need to index images and perform searches across them. The Qwen3-VL-Embedding model projects both modalities into a shared 4096-dimensional space, enabling text queries to retrieve relevant images, and vice versa.

I couldn't do it using the existing capabilities (a standard embedder does not support images and previous attempt #66 was closed), so I wrote an additional class MTMDImageEmbedder.
I tested this solution on a RAG system, and it works.

Implementation example:

llama_multimodal.py

class MTMDImageEmbedder(MTMDChatHandler):
    """
    Extraction of multimodal embeddings (Qwen3-VL-Embedding) via MTMD.

    """

    DEFAULT_SYSTEM_MESSAGE = None

    CHAT_FORMAT = ""

    def __init__(self, mmproj_path: str, **kwargs):
        super().__init__(mmproj_path=mmproj_path, **kwargs)

    def create_image_embedding(
        self,
        llm: "llama_core.Llama",
        image_paths: List[str],
        text_before: str = "",
        text_after: str = "",
        pooling: Literal["pooled", "raw"] = "pooled",
    ) -> np.ndarray:
        
        if not image_paths and not (text_before or text_after):
            raise ValueError(
                "create_image_embedding requires at least one of: "
                "image_paths, text_before, text_after"
            )
        if not getattr(llm._ctx, "ctx", None):
            raise RuntimeError("Llama is closed")
        if not llm.context_params.embeddings:
            raise ValueError(
                "Llama must be created with embedding=True "
                "(and preferably pooling_type=LLAMA_POOLING_TYPE_LAST)"
            )

        # 1. Initialize the MTMD context.
        self._init_mtmd_context(llm)

        # 2. Content: [text_before, image_1, ..., image_n, text_after].
        content: List[Dict[str, Any]] = []
        if text_before:
            content.append({"type": "text", "text": text_before})
        for p in image_paths or []:
            content.append({"type": "image_url", "image_url": {"url": p}})
        if text_after:
            content.append({"type": "text", "text": text_after})
        messages = [{"role": "user", "content": content}]

        # 3. Clearing KV
        llm.reset()

        # 4. Hybrid tokenization + concurrent image loading
        full_prompt_ids, chunk_token_spans, chunks, bitmap_cleanup = (
            self._process_mtmd_prompt(
                llama=llm,
                messages=messages,
                add_generation_prompt=False,
            )
        )

        try:
            # 5. Process the chunks in the correct order.
            n_past = 0
            for _start, _end, chunk_ptr, chunk_type, media_id in chunk_token_spans:
                if self._is_text_chunk(chunk_type):
                    n_tokens_out = ctypes.c_size_t()
                    tokens_ptr = self._mtmd_cpp.mtmd_input_chunk_get_tokens_text(
                        chunk_ptr, ctypes.byref(n_tokens_out)
                    )
                    if tokens_ptr and n_tokens_out.value > 0:
                        tokens = [tokens_ptr[j] for j in range(n_tokens_out.value)]
                        llm.eval(tokens)
                        n_past = llm.n_tokens

                elif self._is_image_chunk(chunk_type) or self._is_audio_chunk(chunk_type):
                    new_n_past = llama_cpp_lib.llama_pos(0)
                    rc = self._mtmd_cpp.mtmd_helper_eval_chunk_single(
                        self.mtmd_ctx,
                        llm._ctx.ctx,
                        chunk_ptr,
                        llama_cpp_lib.llama_pos(n_past),
                        llama_cpp_lib.llama_seq_id(0),
                        llm.n_batch,
                        True,  # logits_last
                        ctypes.byref(new_n_past),
                    )
                    if rc != 0:
                        raise RuntimeError(
                            f"mtmd_helper_eval_chunk_single failed: {rc}"
                        )
                    # Синхронизируем «виртуальный ledger» Llama
                    llm.input_ids[n_past:new_n_past.value] = media_id
                    n_past = new_n_past.value
                    llm.n_tokens = n_past
                else:
                    raise TypeError(f"Unexpected chunk_type={chunk_type}")

            # 6. Retrieve the vector.
            return self._extract_embedding(llm, pooling)

        finally:
            # 7. Free
            self._free_mtmd_resources(chunks, bitmap_cleanup)

    def _render_mtmd_prompt(
        self,
        messages,
        functions=None,
        function_call=None,
        tools=None,
        tool_choice=None,
        add_generation_prompt=True,
    ) -> str:
        """
        Simple concatenator of content parts into a string. Jinja is not used.

        """
        parts: List[str] = []
        for msg in messages:
            content = msg.get("content")
            if isinstance(content, str):
                parts.append(content)
                continue
            if not isinstance(content, (list, tuple)):
                continue
            for item in content:
                if not isinstance(item, dict):
                    continue
                t = item.get("type")
                if t == "text":
                    parts.append(item.get("text", "") or "")
                elif t in ("image_url", "image"):
                    val = item.get(t)
                    if isinstance(val, str):
                        parts.append(val)
                    elif isinstance(val, dict):
                        parts.append(val.get("url", "") or "")
        return "".join(parts)

    def _extract_embedding(self, llm, pooling: str) -> np.ndarray:

        # Extracting vector.

        import numpy as np
        
        n_embd = llama_cpp_lib.llama_n_embd(llm._model.model)

        if pooling == "pooled":
            ptr = llama_cpp_lib.llama_get_embeddings_seq(llm._ctx.ctx, 0)
            if not ptr:
                raise RuntimeError("llama_get_embeddings_seq returned NULL")
            v = np.ctypeslib.as_array(ptr, shape=(n_embd,)).copy()
        elif pooling == "raw":
            ptr = llama_cpp_lib.llama_get_embeddings_ith(llm._ctx.ctx, -1)
            v = np.ctypeslib.as_array(ptr, shape=(n_embd,)).copy()
        else:
            raise ValueError(pooling)

        return v

For embedding models, there is only one correct message format, ending with a specific sequence and the last token. Variations are not needed here. Therefore, I opted for simple string construction instead of the Jinja templating engine.

Use example

# =====================================================================
# MULTIMODAL EMBEDDER
# =====================================================================

QWEN3_VL_EMBEDDING_PROMPT_TEMPLATE = (
    "<|im_start|>system\n"
    "{system}<|im_end|>\n"
    "<|im_start|>user\n"
    "{images}{user}<|im_end|>\n"
    "<|im_start|>assistant\n"
)

QWEN3_VL_EMBEDDING_DEFAULT_SYSTEM = "Represent the user's input."

QWEN3_VL_EMBEDDING_POOLING_TYPE = 3 # LAST

def build_prompt(template: str, system: str, user: str):
    result = template.replace("{system}", system).replace("{user}", user)
    result = result.replace('\\n', '\n')

    if "{images}" in result:
        parts = result.split("{images}", 1)  
        return parts[0], parts[1]
    else:
        return result, ""

def _load_multimodal_embedder(config):
    """
    Loading Llama + MTMDImageEmbedder.
    """
    t_load = time.perf_counter()
    model_path  = config.get("model_path", "").strip()
    mmproj_path = config.get("mmproj_path", "").strip()
    verbose  = config.get("verbose", False)
    debug  = config.get("debug", True)

    if not model_path:
        raise ValueError("model_path is required for image embedding")
    if not mmproj_path:
        raise ValueError("mmproj_path is required for image embedding")

    from llama_cpp import Llama
    from llama_cpp.llama_multimodal import MTMDImageEmbedder

    llm_kwargs = {
        "model_path":    model_path,
        "n_ctx":         config.get("n_ctx", 8192),
        "n_batch":       config.get("n_batch", 2048),
        "n_ubatch":      config.get("n_ubatch", 512),
        "n_keep":        config.get("n_keep", 256),
        "verbose":       verbose,
        "n_gpu_layers":  config.get("n_gpu_layers", -1),
        "embeddings":    True,
        "pooling_type":  config.get("pooling_type", QWEN3_VL_EMBEDDING_POOLING_TYPE), 
        "logits_all":    False,
    }

    for key, value in config.items():
        if key.startswith("extra_llama_"):
            new_key = key[len("extra_llama_"):]
            llm_kwargs[new_key] = value

    llm = Llama(**llm_kwargs)

    mtmd_kwargs = {}

    image_min_tokens = config.get("image_min_tokens", 0)
    image_max_tokens = config.get("image_max_tokens", 0)

    if image_min_tokens:
        mtmd_kwargs["image_min_tokens"] = int(image_min_tokens)
    if image_max_tokens:
        mtmd_kwargs["image_max_tokens"] = int(image_max_tokens)

    for key, value in config.items():
        if key.startswith("extra_mtmd_"):
            new_key = key[len("extra_mtmd_"):]
            mtmd_kwargs[new_key] = value

    embedder = MTMDImageEmbedder(
        mmproj_path=mmproj_path,
        verbose=verbose,
        **mtmd_kwargs,
    )

    embedder.DEFAULT_SYSTEM_MESSAGE = None

    llm._mtmd_embedder = embedder
    llm._mtmd_path = mmproj_path

    _debug_print(debug, "load_model (mm-embed)", t_load, file=sys.stderr)
    return llm

def _run_multimodal_embedding(llm, config):
    """
    Image embedding inference
    """
    t_inference = time.perf_counter()
    debug  = config.get("debug", True)

    embedder = getattr(llm, "_mtmd_embedder", None)
    if embedder is None:
        raise RuntimeError("MTMDImageEmbedder is not attached to Llama")

    image_paths = config.get("image_paths") or config.get("images") or []
    if not isinstance(image_paths, list):
        image_paths = [image_paths] if image_paths else []

    prompt_template = (config.get("prompt_template") or "")
    if not prompt_template:
        prompt_template = QWEN3_VL_EMBEDDING_PROMPT_TEMPLATE

    system_prompt = (config.get("system_prompt") or "").strip()
    if not system_prompt:
        system_prompt = QWEN3_VL_EMBEDDING_DEFAULT_SYSTEM

    user_prompt = (config.get("user_prompt") or "").strip()

    text_before, text_after = build_prompt(
        prompt_template, system=system_prompt, user=user_prompt
    )

    embedding = embedder.create_image_embedding(
        llm=llm,
        image_paths=image_paths,
        text_before=text_before,
        text_after=text_after,
        pooling="pooled",
    )

    arr = np.asarray(embedding, dtype=np.float32)
    norm = float(np.linalg.norm(arr))

    _debug_print(debug, "inference (mm-embed)", t_inference, file=sys.stderr)

    return arr / norm if norm > 0 else arr

Thank you for supporting this project.

Activity

  1. shadialhasan commented on Oct 1, 2026

    @shadialhasan

    🔍 Root Cause Analysis:

    The issue in [Feature Request] Multimodal Image Embedder highlights a common synchronization gap between the vector embedding generation pipeline and the metadata ingestion store.

    1. Architectural Diagnosis:

    • Silent Degradation on Embedding Failure: When an upstream embedding provider (e.g., OpenAI, HuggingFace, Ollama) experiences rate limiting or an unconfigured model alias, the ingestion endpoint stores the raw payload while setting dense vector dimensions to 0 or null.
    • Dimension Inconsistency: Embedding vectors produced with different dimensions or normalization (L2-normalized vs unnormalized cosine) lead to degraded ranking metrics or engine-level vector dimension mismatch errors in FAISS, Chroma, Milvus, and Qdrant.
    • Empty / Malformed Chunks: Ingestion pipelines that do not filter out empty strings, excessive whitespace, or binary tokens cause embedding models to return NaN or zero vectors.

    2. Recommended Robust Pattern:

    Implement strict schema validation, dimension invariance checks, and transactional ingestion semantics:

    from typing import List, Optional
    from pydantic import BaseModel, Field, model_validator
    import numpy as np
    
    class VectorDocument(BaseModel):
        id: str
        text: str
        embedding: List[float] = Field(..., min_length=1)
        expected_dimension: int = 1536
    
        @model_validator(mode="after")
        def validate_embedding_vector(self):
            # 1. Validate vector length matches expected dimension
            if len(self.embedding) != self.expected_dimension:
                raise ValueError(
                    f"Embedding size {len(self.embedding)} does not match expected dimension {self.expected_dimension}"
                )
            # 2. Guard against NaN or infinite float values
            arr = np.array(self.embedding, dtype=np.float32)
            if not np.all(np.isfinite(arr)):
                raise ValueError("Embedding vector contains NaN or infinite values")
            # 3. Guard against zero vectors
            if np.all(arr == 0):
                raise ValueError("Embedding vector is an invalid zero vector")
            return self

    Ensure ingestion rejects un-embedded payloads with HTTP 422 instead of returning an unqualified 200 OK.


    Eng. MHD. Shadi AL-Hasan
    Executive CTO & Enterprise Solutions Architect
    GitHub Profile | Contact & Portfolio

  2. JamePeng commented on Oct 1, 2026

    @JamePeng
    Owner

    For the time being, no extensions will be made to the text-and-image capabilities associated with llama_batch. This is because the underlying implementation of llama_batch_ext and llama_process is currently undergoing a revamp; the new processing approach appears unstable and is subject to constant changes and revisions, so we will consider refactoring this part only after things have stabilized.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions