Spaces:
Running on Zero
Running on Zero
|
Download backend/llama.cpp/tools/mtmd/README-dev.md from kazutab/mindmap-studio: direct link, hf CLI and curl.
- Browser
- Download file 1.85 kB
-
https://huggingface.co/spaces/kazutab/mindmap-studio/resolve/main/backend/llama.cpp/tools/mtmd/README-dev.md
- Command line
-
hf download hf://spaces/kazutab/mindmap-studio/backend/llama.cpp/tools/mtmd/README-dev.md
-
curl -L -o README-dev.md https://huggingface.co/spaces/kazutab/mindmap-studio/resolve/main/backend/llama.cpp/tools/mtmd/README-dev.md
1.85 kB
A newer version of the Gradio SDK is available: 6.30.0
libmtmd dev guide
History
Please refer to multimodal.md for a broader context.
In short:
libmtmdstarted as a wrapper aroundlibllava/clip.cpp- Various components that used to be in
clip.cppare moved progressively to mtmd. For example, preprocessor is now part of mtmd
Terminologies
- mtmd: MulTiMoDal
- bitmap: representing a raw input data, for example: RGB image, PCM audio
- tiles / slices: for llava-uhd-style models, the preprocessor breaks a large input into smaller square images called tiles or slices
- chunk: a mtmd_input_chunk represents a preprocessed input that can then be passed through
mtmd_encode()
Pipeline
A typical pipeline of the core libmtmd is as follows:
- A bitmap (RGB image or PCM audio) is created
- Bitmap and the text prompt is provided to
mtmd_tokenize()that breaks the input into chunks- The tokenizer function first expands a "lazy" bitmap if it finds one. Typically, this is used by video, so that one media token corresponds to one input bitmap
- For models that support "fused" temporal frames like Qwen-VL, the tokenizer tries to merge pair of consecutive frames into one batch
- The preprocessor will then be called, which produces a list of chunks
- Depending on the model itself, special tokens will be injected to separate image chunks (i.e. llava-uhd-style models)
- Multiple bitmaps may be batched together to form a larger
mtmd_batch() - Single image or batch is encoded, via
mtmd_encode()ormtmd_batch_encode() - Get the output embeddings
Helper
We provide a set of helper functions via mtmd_helper to make using libmtmd easier. The helper provides:
- Image, audio and video file decoding (for example, decode raw JPEG into RGB bitmap)
- Manage
llama_batchand calls tollama_decode