Instructions to use kalle07/embedder_collection with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use kalle07/embedder_collection with sentence-transformers:
from sentence_transformers import SentenceTransformer model = SentenceTransformer("kalle07/embedder_collection") sentences = [ "That is a happy person", "That is a happy dog", "That is a very happy person", "Today is a sunny day" ] embeddings = model.encode(sentences) similarities = model.similarity(embeddings, embeddings) print(similarities.shape) # [4, 4] - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use kalle07/embedder_collection with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf kalle07/embedder_collection:Q4_K_M # Run inference directly in the terminal: llama cli -hf kalle07/embedder_collection:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf kalle07/embedder_collection:Q4_K_M # Run inference directly in the terminal: llama cli -hf kalle07/embedder_collection:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf kalle07/embedder_collection:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf kalle07/embedder_collection:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf kalle07/embedder_collection:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf kalle07/embedder_collection:Q4_K_M
Use Docker
docker model run hf.co/kalle07/embedder_collection:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use kalle07/embedder_collection with Ollama:
ollama run hf.co/kalle07/embedder_collection:Q4_K_M
- Unsloth Desktop
- Docker Model Runner
How to use kalle07/embedder_collection with Docker Model Runner:
docker model run hf.co/kalle07/embedder_collection:Q4_K_M
- Lemonade
How to use kalle07/embedder_collection with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull kalle07/embedder_collection:Q4_K_M
Run and chat with the model
lemonade run user.embedder_collection-Q4_K_M
List all available models
lemonade list
- Atomic Chat
- A practical introduction to embedding models and local RAG
- Table of Contents
- Testing and software compatibility
- My short practical impression
- German and more complex documents
- Toxic-content models
- Practical starting settings for a local RAG system
- What do these numbers actually mean?
- Characters, words and tokens
- VRAM and context length
- Vector size (dimensions)
- Vector count
- Chunk length
- How embedding and retrieval actually work
- Embedding search is not the same as keyword search
- Why the most relevant snippets matter
- Chunk size versus retrieval detail
- Why "summarize the whole document" often gives poor results in RAG
- Table of contents, indexes and bibliographies
- Small documents
- TXT files and summaries
- The main LLM is also very important
- System prompts
- Example system prompt for factual document questions
- Example system prompt for storytelling
- Example system prompt for cooking
- About Jinja templates
- DOC/PDF → TXT: prepare your documents first
- Check the extracted text
- Simple tools for PDF/table extraction
- My PDF-to-TXT tools
- OCR and more advanced document parsing
- Large UI-based PDF parsing option
- Fast indexing without embeddings
- Final practical advice
- One final rule
- (ALL licenses and terms of use go to original author)
A practical introduction to embedding models and local RAG
This repository contains more than 25 different types of embedding models, together with practical information about embeddings, document retrieval, chunking, context length, and local RAG (Retrieval-Augmented Generation).
The most important point is this:
An embedding model is only one part of a good RAG system.
A good result depends on several things working together:
- the quality and structure of your source documents
- how the documents are converted to text
- how the text is split into chunks
- the embedding model
- the vector database and retrieval settings
- optional keyword or reranking steps
- the context given to the main LLM
- the main LLM itself
- the system prompt
Changing only the embedding model therefore does not automatically make a RAG system better.
At the end of the file list on this repo, click the button shown below to see all files:
Table of Contents
- Testing and software compatibility
- My short practical impression
- German and more complex documents
- Toxic-content models
- Practical starting settings for a local RAG system
- What do these numbers actually mean?
- Characters, words and tokens
- VRAM and context length
- Vector size (dimensions)
- Vector count
- Chunk length
- How embedding and retrieval actually work
- Embedding search is not the same as keyword search
- Why the most relevant snippets matter
- Chunk size versus retrieval detail
- Why "summarize the whole document" often gives poor results in RAG
- Table of contents, indexes and bibliographies
- Small documents
- TXT files and summaries
- The main LLM is also very important
- System prompts
- Example system prompt for factual document questions
- Example system prompt for storytelling
- Example system prompt for cooking
- About Jinja templates
- DOC/PDF → TXT: prepare your documents first
- Check the extracted text
- Simple tools for PDF/table extraction
- My PDF-to-TXT tools
- OCR and more advanced document parsing
- Large UI-based PDF parsing option
- Fast indexing without embeddings
- Final practical advice
- One final rule
Testing and software compatibility
The models in this collection have been tested by me with AnythingLLM (ALLM) using LM Studio as the local model server.
Many of these models can also be used with Ollama or other local inference software, but compatibility is model-, format-, and backend-dependent. Do not assume that every model in this collection will work in every application without additional configuration.
The workflow for local documents is similar across different applications. For example, GPT4All currently has a smaller selection of built-in embedding models, and support in applications such as KoboldCpp or Jan can differ by version.
For maximum compatibility and reproducibility, I normally prefer F16 or F32 when those formats are practical. Quantized models can also work very well, especially when memory or speed is more important than maximum numerical precision.
Sometimes the result is more reliable when an option such as "chat with document only" is enabled. This depends on the application and the task because such a setting can restrict the LLM from using information outside the retrieved document context.
Important: an embedding model is only one part of RAG. Ideally, the embedding model should match your language and your use case. A model that works well for general English text is not automatically the best choice for German, programming, medicine, legal documents, or other specialized domains.
⇨ Give me a ❤️ if this collection is useful to you.
My short practical impression
The following is based on my own tests and should be treated as a practical observation, not as a universal ranking.
- nomic-embed-text-v2-moe — good general-purpose option; check the exact model version for its supported input length.
- mxbai-embed-large — relatively compact and fast compared with many larger embedding models.
- mug-b-1.6
- qwen3-0.6b — relatively slow in my setup, but supports a long input context; very long input limits are not automatically useful for normal retrieval.
- jinaai_v5-retrieval — based on the Qwen family in the version I tested; relatively slow, with a long input context.
- snowflake-arctic-embed-l-v2.0 — supports long input compared with many smaller embedding models.
- bge-m3 — versatile multilingual embedding model with support for long input.
These models generally worked well in my tests. Some models produce very similar retrieval results, so choosing a much larger model does not necessarily produce a proportional improvement.
In one simple comparison, several embedding models returned approximately 6–7 of the same 10 snippets from a book. That means only about 3–4 retrieved snippets differed. This was not an extensive benchmark, so this should not be interpreted as a scientific comparison.
For some models, LM Studio may require manual configuration. Depending on the model and LM Studio version, the model architecture or domain type may need to be selected in the model settings.
German and more complex documents
I also tested several models with more complex German-language documents.
The following models performed well in my tests:
- GTE large
- cross-en-de-es-roberta
- ger-RAG-bge-M3-merg-snowf-artic-hessian-AI — very good for German in my tests; supports long input.
- German-RAG-BGE-M3-TRIPLES-HESSIAN-AI — very good for German in my tests; supports long input.
- bge-m3 — good multilingual performance, including German.
- jina-embeddings-v3 — good German-language performance in my tests.
In my tests, the German-specific models above sometimes performed better on complex German documents than some more general models.
For example, Jina-DE and nomic did not perform as well in one of my test sets.
I was also not convinced that very large embedding models such as the larger Qwen- or Jina-based models always justify their additional speed and resource requirements. In my setup, some were many times slower without producing a similarly large improvement in retrieval quality.
This is a practical observation rather than a general benchmark. Results can change significantly with different documents, languages, chunk sizes, hardware, retrieval settings, and software.
Some systems can also recognize or use information from tables and images, but this depends heavily on how the original document was parsed. The embedding model itself should not be expected to "understand" an image or table if that information was never converted into useful text or multimodal representations.
Toxic-content models
There are also models in this collection related to toxic-content detection, including:
- toxic-prompt-roberta
- minilmv2-toxic-jigsaw
These should be thought of as text classification / moderation models, rather than ordinary document-embedding models.
There is also IBM's Granite Guardian, which is a safety-focused language-model family rather than a conventional embedding model.
I have not performed a sufficiently large benchmark to give reliable accuracy numbers for the toxicity models.
...
Practical starting settings for a local RAG system
Here is one example configuration for a relatively large context with several expected search hits.
In LM Studio, configure the main LLM with approximately 8,000 tokens of context as a starting point.
In AnythingLLM, you can start with:
- Max Embedding Chunk Length: about 1,500 characters
- Max Context Snippets: about 10
- Search Preference: Accuracy
- Text splitting / Chunk Size: about 1,500 characters
- Workspace snippets: about 10
Important: if you use LM Studio together with AnythingLLM, check the context settings in both applications. The final usable context can be limited by the main LLM, the application settings, the model's actual architecture, or the backend.
Start both models in LM Studio and make sure the correct embedding model and main LLM are selected in AnythingLLM.
These are only starting values. There is no universal "correct" chunk size or number of snippets.
What do these numbers actually mean?
Suppose your documents are split into chunks of roughly 1,500 characters and the retrieval system returns 10 chunks.
Then the retrieved text contains roughly:
10 × 1,500 = 15,000 characters
That is only a rough size because chunks are not always exactly the requested length, and the application may use overlap, metadata, separators, formatting, or additional prompt text.
Depending on the language and tokenizer, 15,000 characters may correspond to several thousand tokens. The exact number cannot be calculated reliably from characters alone.
The main LLM then has to process:
system prompt + user question + retrieved text + conversation history + its generated answer
within its available context.
That is why a model advertised with a very large context window is not automatically better at RAG. A large context window is a capacity limit, not a guarantee that the model will use a very large amount of retrieved information accurately.
For normal document questions, retrieving a small number of highly relevant chunks is often more useful than filling the entire context window with loosely related text.
You can experiment with different combinations, for example:
- 5 snippets × 5,000 characters
- 10 snippets × 1,500 characters
- 20 snippets × 500 characters
The best setting depends on your documents and your questions.
When you change the chunk size, existing embeddings may need to be recreated, because changing the chunk boundaries changes the text that is embedded.
Characters, words and tokens
Do not treat characters, words, and tokens as interchangeable units.
A tokenizer decides how text is split into tokens, and different models use different tokenizers. The token count also changes between languages.
German can require more tokens than English for the same amount of information, but there is no fixed "50% rule" that applies to every text.
As a rough practical rule, several thousand characters correspond to roughly one page of ordinary book text, but page size, formatting, font size, line spacing, and the language all matter.
As a very rough rule of thumb:
- 1 token ≈ 4.2 characters but this varies by tokenizer and language (only text not coding).
- 3,000–5,000 characters are roughly the amount of small text you might find on one book page, depending on formatting.
- 3,000–5,000 characters often correspond to roughly 1,000–1,200 tokens, but the exact number depends on the tokenizer and language.
- 1,200 tokens may correspond to roughly 900–1,000 English words, but this is only an approximation.
- 16,000 tokens ~ 1GB - 3GB VRAM, depends strong on the model.
For exact calculations, use the tokenizer of the model you are actually using.
Here is a tokenizer calculator:
https://platform.openai.com/tokenizer
VRAM and context length
Context length is only one part of memory usage.
VRAM requirements can depend on:
- the number of model parameters
- model precision or quantization
- context length
- KV-cache size for the main LLM
- batch size and inference settings
- the inference backend
- whether one or more models are loaded at the same time
Therefore, simple rules such as "8,000 tokens always need X GB" are not reliable.
In particular, embedding models and generative LLMs do not use memory in exactly the same way.
A VRAM calculator can still be useful as a starting point:
https://huggingface.co/spaces/NyxKrage/LLM-Model-VRAM-Calculator
For some calculators you need the original model page rather than a GGUF file.
For GGUF models, an exact GGUF-oriented calculator can be useful. Use the direct Hugging Face file URL format expected by the calculator.
Example:
GGUF VRAM calculator:
https://huggingface.co/spaces/oobabooga/accurate-gguf-vram-calculator
Vector size (dimensions)
The vector size, also called embedding dimensionality, is the number of numerical values in each embedding vector.
For example, a model might generate:
embedding_length = 768
That means one text chunk is represented by 768 numerical values.
The dimensionality is determined by the model architecture and normally cannot be changed arbitrarily.
Some modern embedding models support techniques such as Matryoshka Representation Learning, which allow vectors to be shortened while preserving a useful amount of performance. This is a model-specific feature and should not be assumed for every embedding model.
Higher dimensionality can provide more capacity, but more dimensions do not automatically mean better retrieval. The quality of the model, its training data, language support, and the retrieval task are at least as important.
Larger vectors also require more storage and usually increase the computational cost of indexing and similarity search.
Vector count
The vector count is the number of embeddings stored in the vector database.
In a simple document collection, this is closely related to the number of text chunks.
For example:
1,000 chunks → roughly 1,000 vectors
If chunk overlap is used, the same document may produce more chunks and therefore more vectors.
More vectors can provide more fine-grained retrieval because the document is divided into smaller pieces. However, very small chunks can also lose context and increase indexing and retrieval overhead.
There is therefore a trade-off between:
chunk size ↔ number of vectors ↔ context per result ↔ retrieval precision ↔ processing cost
Chunk length
Chunk length is the amount of text placed into one chunk before the text is embedded.
Depending on the software, this can be measured in:
- characters
- tokens
- words
In the AnythingLLM interface, chunk size is commonly configured in characters.
Chunk size is one of the most important RAG settings.
Large chunks contain more context but may also contain irrelevant information.
Small chunks can make retrieval more precise, but a small chunk may not contain enough surrounding context to answer the question correctly.
There is no universally optimal chunk size.
For technical documents, books, manuals, and structured data, it can also be useful to split based on headings, paragraphs, sections, tables, or other semantic boundaries instead of cutting the document purely every N characters.
How embedding and retrieval actually work
Embedding-based retrieval is easiest to understand with a simple example.
Imagine you have a 300-page book containing around 90,000 words.
You ask:
"What is described in chapter XYZ in relation to person ZYX?"
A typical RAG pipeline does approximately the following:
1. Prepare the document
The PDF, DOCX, TXT, HTML, or other source is converted into text or another supported representation.
2. Split the document into chunks
For example, one chapter might be divided into many chunks of roughly 1,500 characters.
3. Create embeddings
The embedding model converts every chunk into a numerical vector.
For example:
chunk → [0.12, -0.04, 0.73, ...]
The vector is a mathematical representation of the information in that chunk.
4. Embed the user's query
The question is also converted into a vector.
5. Compare the query with document vectors
The retrieval system calculates how similar the query vector is to the stored document vectors, often using cosine similarity, dot product, or another similarity measure.
6. Return the most relevant chunks
Only some of the chunks are selected and sent to the main LLM.
7. The main LLM generates the answer
The LLM receives the retrieved text as context and uses it to answer the question.
This distinction is important:
The embedding model does not write the final answer.
It helps find text that is likely to be relevant.
Embedding search is not the same as keyword search
Embedding retrieval is usually based on semantic similarity rather than an exact search for a literal word.
For example, a query containing:
"automobile"
may retrieve a passage containing:
"car"
because the concepts are related.
However, this does not mean exact word search is impossible.
A complete retrieval system may combine:
- semantic/vector search
- keyword or full-text search
- metadata filters
- reranking
Hybrid retrieval can be especially useful when exact names, product numbers, legal references, error codes, or other precise terms matter.
Therefore, it is more accurate to say:
Embedding search is not primarily an exact word-count or literal keyword search.
Why the most relevant snippets matter
Suppose the word or concept XYZ occurs 50 times in your document.
The system normally does not send all 50 occurrences to the LLM.
Instead, it retrieves a limited number of chunks with the highest relevance scores.
For example, it might return 10 chunks.
If only one of those chunks is truly relevant to the question and the other nine are only loosely related, the extra text can make the answer less precise.
That is why retrieving more information is not automatically better.
A useful starting point is often somewhere around 4–20 snippets, but the correct number depends strongly on the task.
For a question with one very specific answer, a small number of highly relevant chunks may be best.
For a question that requires information from several different parts of the document, more chunks may be necessary.
Chunk size versus retrieval detail
Consider two simple examples:
2,000-character chunks
You get larger pieces of text with more surrounding context.
This can help when the meaning of a statement depends on the paragraphs before and after it.
500-character chunks
The retrieval system can distinguish more individual pieces of information, but each result contains less context.
Smaller chunks also mean more vectors, more indexing work, and potentially more retrieval overhead.
A useful practical approach is to start somewhere in the middle and then test the questions that are actually important for your documents.
Why "summarize the whole document" often gives poor results in RAG
A normal RAG pipeline retrieves only a subset of the document.
Therefore, asking:
"Summarize this entire 300-page book."
does not automatically cause the embedding system to read all 300 pages.
It usually retrieves the chunks that appear most relevant to the query. If the query is simply "summarize the book", the retrieval system may return introductory material, a summary, a table of contents, or other text that happens to be highly similar to the query.
That can produce an incomplete summary.
For document-wide summarization, a dedicated summarization workflow is usually more appropriate, for example:
document → chunks → summaries of chunks → summaries of summaries
or another hierarchical summarization approach.
Table of contents, indexes and bibliographies
A table of contents, index, bibliography, or reference list can contain many important keywords.
That can make these pages appear highly relevant during retrieval even though they do not contain the information needed to answer the question.
For some document collections, it may therefore make sense to exclude or separately index these pages.
However, do not delete them automatically. In some tasks, references and indexes are exactly the information you need.
Small documents
For small documents, such as 10–20 pages, retrieving the entire document can sometimes be simpler than building a complex RAG pipeline.
Some applications offer options such as pinning or directly attaching the complete document to the conversation.
Whether this works depends on the context window of the main LLM and the application.
For a small document that fits comfortably into the available context, direct context can sometimes avoid retrieval errors altogether.
TXT files and summaries
An embedding database does not itself generate summaries.
However, an LLM can still summarize a TXT document if the application retrieves enough of the relevant text or passes the complete document into the LLM.
So the important distinction is:
Embeddings retrieve text.
The LLM generates the summary.
A RAG system can therefore summarize documents, but the quality depends on how much of the source document is actually available to the LLM.
The same principle applies to exact word searches and page searches. A vector search is not intended to replace a dedicated full-text or document search engine.
For exact terms, identifiers, numbers, page references, or counts, a conventional keyword/full-text search can be more appropriate.
...
The main LLM is also very important
The embedding model determines which information is retrieved.
The main LLM determines how that information is interpreted and used to produce the final answer.
This matters especially with large contexts.
A model may technically support 32k, 128k, or even larger context windows. That does not mean it will use every part of a very large context equally well.
Performance can degrade when a large amount of irrelevant or weakly relevant information is included.
Therefore:
Maximum context length ≠ guaranteed effective use of that context.
For RAG, choose a main model that:
- supports your language well
- follows instructions reliably
- handles the amount of context you actually provide
- is strong enough for the complexity of the questions
- fits your available hardware
Instruction-tuned / Instruct models are generally the appropriate starting point for question answering from documents.
For longer contexts, a larger model can be useful, but model size alone is not a guarantee of better answers.
The best choice depends on the language and the task.
For example, a technical manual, a programming repository, a historical book, and a creative-writing task may benefit from different models and prompts.
System prompts
The system prompt can strongly influence how the main LLM uses the retrieved information.
It normally does not directly change the stored embeddings.
However, there is an important exception:
Some applications use an LLM to rewrite, expand, classify, or otherwise transform the user's query before retrieval.
In such a pipeline, a prompt can indirectly affect retrieval because it changes the query that gets embedded or searched.
With a simple direct vector-retrieval pipeline, the system prompt primarily affects the answer generation stage.
A useful test is to compare the same documents and question with two different prompts.
Example system prompt for factual document questions
You are a helpful assistant who answers questions using the provided excerpts from the document collection.
Answer the user's question based primarily on the retrieved excerpts. Do not invent information that is not supported by the provided context.
Consider the relevance and reliability of each excerpt. Use the most relevant excerpts more heavily than weakly related excerpts.
If the provided excerpts are insufficient to answer the question reliably, say so instead of guessing.
After the answer, briefly identify which excerpts were most useful and why.
I originally used a more elaborate version that asked the model to explain why excerpts were included. That can be useful for testing a RAG setup, but it is not always necessary in normal use.
Example system prompt for storytelling
You are an imaginative storyteller who creates compelling narratives with depth, creativity, and coherence.
Use the provided excerpts as source material when they are relevant to the user's request.
When generating a story, maintain consistency in characters, setting, chronology, and plot progression.
Stay faithful to factual details contained in the provided source material, while using creativity where the user explicitly asks for fiction.
Do not invent source-based facts and present them as if they were contained in the documents.
Example system prompt for cooking
You are a warm and knowledgeable cooking assistant.
Use the provided excerpts when they contain relevant recipes, techniques, ingredients, or cultural information.
Give clear and practical explanations and adapt the answer to the user's request.
When the source material does not contain the required information, say so instead of presenting unsupported claims as facts.
About Jinja templates
Jinja/Jinja2 templates are used by some local LLM applications to construct the prompts sent to models.
Modern models can sometimes require specific chat templates. Most common models work with the correct standard template, while merged or customized models can require additional configuration.
I am not a model-format or prompt-template developer, so the comments here are based on practical use rather than implementation expertise.
When a model behaves strangely, produces repeated text, ignores instructions, or formats messages incorrectly, the chat template is one of the things worth checking.
...
DOC/PDF → TXT: prepare your documents first
Bad input can produce bad retrieval.
One of the most overlooked parts of RAG is the quality of the text that reaches the embedding model.
A visually perfect PDF does not necessarily contain clean machine-readable text.
A PDF may contain:
- text positioned manually on the page
- multiple columns
- tables
- headers and footers
- page numbers
- formulas
- scanned images
- unusual character encoding
- captions
- footnotes
- figures
- fragmented paragraphs
When these elements are extracted incorrectly, the embedding model receives a distorted version of the original document.
The retrieval quality can then be poor even when the embedding model itself is good.
Check the extracted text
With AnythingLLM Desktop, extracted document data can be found in its local storage area, for example:
C:\Users\XXX\AppData\Roaming\anythingllm-desktop\storage\documents
The exact path and storage format can change between versions.
Opening the extracted text with a normal text editor is a simple way to see what the embedding pipeline is actually receiving.
A useful test is:
Extract the PDF to TXT → open the TXT → inspect the result.
If headings, paragraphs, tables, page order, or characters are already wrong at this stage, changing the embedding model may not solve the problem.
Simple tools for PDF/table extraction
For relatively simple PDFs and tables, these tools are useful starting points:
- pdfplumber
- PyMuPDF (fitz)
- Camelot — especially useful for certain types of PDF tables
The quality of the result depends heavily on how the source PDF was constructed.
There is no single PDF-to-text tool that works perfectly for every document.
You may need different processing for:
- scanned books
- scientific papers
- financial tables
- multi-column magazines
- forms
- technical manuals
My PDF-to-TXT tools
My own PDF parser/converter is available here:
https://huggingface.co/kalle07/pdf2txt_parser_converter
GitHub version:
https://github.com/kalle07/pdf2txt-parser
I also have a simple raw keyword search and snippet extractor:
https://huggingface.co/kalle07/raw-txt-snippet-creator
OCR and more advanced document parsing
For scanned documents and more complicated layouts, OCR and document-structure extraction are often necessary.
Docling, originally developed by IBM Research and now open source, is one option worth investigating.
It provides examples for document parsing, PDF processing, OCR, tables, and other document structures:
https://github.com/docling-project/docling/tree/main/docs/examples
For many use cases, the examples are already short enough to serve as a practical starting point.
OCR-based processing can automatically download or use additional models depending on the chosen configuration.
One thing I particularly noticed is that font information can be useful when reconstructing the document structure. PyMuPDF (fitz) can expose font information in ways that are useful for custom document processing.
Large UI-based PDF parsing option
For experimenting with many document-processing features through a graphical interface:
- Parse My PDF
https://github.com/genieincodebottle/parsemypdf
...
Fast indexing without embeddings
Sometimes you do not need semantic search at all.
When you have tens of thousands of PDF, TXT, or DOC files, a traditional index can be a very fast first step.
For example, you can use an index to find the 5–10 documents that are likely to contain the information you need and then give only those documents to an LLM.
This is not the same as embedding-based retrieval.
Traditional indexing is particularly useful when you need exact searches such as:
- author names
- filenames
- titles
- ISBNs
- product numbers
- error codes
- exact phrases
Two projects worth looking at are:
JabRef:
https://github.com/JabRef/jabref/tree/v6.0-alpha?tab=readme-ov-file
https://builds.jabref.org/main/
and:
DocFetcher:
https://docfetcher.sourceforge.io/en/index.html
DocFetcher is an older project, but traditional local full-text indexing can still be very useful.
Final practical advice
When a RAG system gives poor answers, do not immediately assume that the embedding model is the problem.
Check the complete pipeline:
1. Source document
Is the original PDF/DOC/TXT actually readable and complete?
2. Extraction
Did the PDF parser preserve paragraphs, tables, headings, page order, and characters?
3. Chunking
Are the chunks too large, too small, or split at the wrong places?
4. Embedding model
Does the model work well for your language and domain?
5. Retrieval
Are the correct chunks actually being returned?
6. Number of snippets
Are you sending too much irrelevant information to the main LLM?
7. Main LLM
Can the model reliably use the amount of context you are giving it?
8. Prompt
Does the system prompt clearly tell the LLM to use the retrieved documents and avoid unsupported claims?
9. Exact search requirements
Would keyword search or hybrid search be better than semantic search for this particular question?
10. Test with real questions
The best embedding model for your system is the one that performs well on the questions you actually need to answer, not necessarily the one with the largest parameter count or vector dimension.
One final rule
Do not judge an embedding model only by the model name, parameter count, context length, or vector dimensions.
Take a representative set of your own documents and questions.
Then compare:
Which chunks were retrieved?
Were the relevant chunks present?
Were irrelevant chunks also included?
Can the main LLM answer the question correctly from those chunks?
That end-to-end test is much more useful than comparing model specifications alone.
The models and software mentioned in this README can change over time. Compatibility, supported context lengths, quantization formats, and application behavior should therefore be checked against the documentation for the specific model and software version you are using.
...
(ALL licenses and terms of use go to original author)
...
- avemio/German-RAG-BGE-M3-MERGED-x-SNOWFLAKE-ARCTIC-HESSIAN-AI (German, English)
- maidalun1020/bce-embedding-base_v1 (English and Chinese)
- maidalun1020/bce-reranker-base_v1 (English, Chinese, Japanese and Korean)
- BAAI/bge-reranker-v2-m3 (English and Chinese)
- BAAI/bge-reranker-v2-gemma (English and Chinese)
- BAAI/bge-m3 (English 40% and Chinese 20%, after Spain, German, Russion, Italian, French ... )
- avsolatorio/GIST-large-Embedding-v0 (English)
- ibm-granite/granite-embedding-278m-multilingual (English, German, Spanish, French, Japanese, Portuguese, Arabic, Czech, Italian, Korean, Dutch, and Chinese)
- ibm-granite/granite-embedding-125m-english
- Labib11/MUG-B-1.6 (?)
- mixedbread-ai/mxbai-embed-large-v1 (multi)
- nomic-ai/nomic-embed-text-v1.5 (English, multi)
- nomic-ai/nomic-embed-text-v2-moe (English, Spanish, French, German, Italian, Portuguese, Polish all other 100-languages are less trained)
- Snowflake/snowflake-arctic-embed-l-v2.0 (English, multi)
- intfloat/multilingual-e5-large-instruct (100 languages)
- T-Systems-onsite/german-roberta-sentence-transformer-v2
- T-Systems-onsite/cross-en-de-roberta-sentence-transformer (English, German)
- T-Systems-onsite/cross-en-de-es-roberta-sentence-transformer (English, German, Spanish)
- T-Systems-onsite/cross-en-de-fr-roberta-sentence-transforme (English, German, France)
- mixedbread-ai/mxbai-embed-2d-large-v1
- jinaai/jina-embeddings-v2-base-en
- jinaai/jina-embeddings-v5-text-small-retrieval
- Qwen/Qwen3-Embedding-0.6B (multi)
- HIT-TMG/KaLM-embedding-multilingual-mini-instruct-v1.5
- thenlper/gte-large
- sentence-transformers/all-MiniLM-L6-v2
- TatonkaHF/bge-m3_en_ru (En - RU)
- google/embeddinggemma-300m (multi)
- Downloads last month
- 13,805
4-bit
8-bit
16-bit
32-bit
