Best local embedding model for job-posting search on a 16 GB M4 Mac?

#47
by GopiKrishnaReddy - opened

Hi all,

I'm building a local job-search assistant and I'm looking for the best
embedding model for it. Everything runs on my own machine, so speed and
memory matter as much as quality.

The setup

  • Corpus: about 10,600 English job postings, split into 500-character
    chunks: 141,603 chunks, about 13.6M tokens (roughly 96 tokens per chunk).
  • Queries: short questions like "backend roles using Kafka", plus
    paraphrased ones that describe the work without the posting's own words.
  • Search: Postgres + pgvector. Keywords pick the candidates and the
    embedding model orders them, by each posting's best-matching chunk.
  • Hardware: Apple M4, 16 GB RAM (MPS). No cloud, no paid APIs.

What I measured (56 postings, two questions each, pool of 400 postings;
"top 5" = the posting the question was written from is in the top 5)

Model Natural top 5 Paraphrased top 5 Speed
bge-small-en-v1.5 (current) 77% 32% 366 chunks/s
bge-large-en-v1.5 79% 45% 34 chunks/s
Qwen3-Embedding-4B (Q4 GGUF) 91% 71% 3.3 chunks/s

Qwen3-Embedding-4B is clearly the best, but it would take about 12 hours
to embed my corpus, against about 6 minutes for bge-small. I'm testing
Qwen3-Embedding-0.6B now (about 19 chunks/s so far).

What I'm looking for

  1. A model with quality close to Qwen3-Embedding-4B, especially on
    paraphrased queries, that runs much faster on Apple Silicon.
  2. Or a small embedding model plus a reranker that gets there together.
  3. Any advice on the fastest runtime for these models on an M4
    (sentence-transformers on MPS, MLX, ONNX, llama.cpp).

Constraints: English only, runs locally in under about 3 GB of memory,
free to use.

Thanks for any suggestions!

Gopi Krishna Reddy Katkuri

Sign up or log in to comment