CodeBERT base โ€” GGUF Q8_0

microsoft/codebert-base converted to GGUF at Q8_0 for llama.cpp's llama-server --embeddings.

File codebert-base-Q8_0.gguf (135 MB)
SHA-256 01a85fb726bf4dc063aaac34b33fde08f991b96b3f9f76ff10e28fa2269df5ea
Source revision 99d7ef814601faaf7bdc2f774ffa7dade4f4d828 (safetensors)
Context 512 tokens
Pooling mean (stored in the file, no --pooling flag needed)
Dimensions 768
llama-server -m codebert-base-Q8_0.gguf --embeddings --port 8080
curl -s localhost:8080/v1/embeddings -H 'Content-Type: application/json' -d '{"input":["def f(): pass"]}'

How it was converted

convert_hf_to_gguf.py (llama.cpp, 2026-10-01) with --outtype q8_0, after two additions to the source directory:

  • a tokenizer.json, written by RobertaTokenizerFast.save_pretrained. The upstream repo ships only vocab.json + merges.txt, and without tokenizer.json the converter falls back to a WordPiece vocabulary and tokenizes code wrongly;
  • a sentence-transformers modules.json + 1_Pooling/config.json declaring mean pooling, so the pooling type is written into the GGUF.

Verification

40 random code/console/output blocks (first 600 characters), embedded by this file on llama-server (CPU) and by the original model in PyTorch fp32 with masked mean pooling over the last hidden state. Token ids are identical; cosine similarity min 0.9998, median 0.99994.

The weights and their licence (MIT) are Microsoft's; see the CodeBERT paper.

Downloads last month
-
GGUF
Model size
0.1B params
Architecture
bert
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for ChrisGVE/codebert-base-Q8_0-GGUF

Quantized
(7)
this model

Paper for ChrisGVE/codebert-base-Q8_0-GGUF