How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf ChrisGVE/codebert-base-Q8_0-GGUF:Q8_0
# Run inference directly in the terminal:
llama cli -hf ChrisGVE/codebert-base-Q8_0-GGUF:Q8_0
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf ChrisGVE/codebert-base-Q8_0-GGUF:Q8_0
# Run inference directly in the terminal:
llama cli -hf ChrisGVE/codebert-base-Q8_0-GGUF:Q8_0
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf ChrisGVE/codebert-base-Q8_0-GGUF:Q8_0
# Run inference directly in the terminal:
./llama-cli -hf ChrisGVE/codebert-base-Q8_0-GGUF:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf ChrisGVE/codebert-base-Q8_0-GGUF:Q8_0
# Run inference directly in the terminal:
./build/bin/llama-cli -hf ChrisGVE/codebert-base-Q8_0-GGUF:Q8_0
Use Docker
docker model run hf.co/ChrisGVE/codebert-base-Q8_0-GGUF:Q8_0
Quick Links

CodeBERT base โ€” GGUF Q8_0

microsoft/codebert-base converted to GGUF at Q8_0 for llama.cpp's llama-server --embeddings.

File codebert-base-Q8_0.gguf (135 MB)
SHA-256 01a85fb726bf4dc063aaac34b33fde08f991b96b3f9f76ff10e28fa2269df5ea
Source revision 99d7ef814601faaf7bdc2f774ffa7dade4f4d828 (safetensors)
Context 512 tokens
Pooling mean (stored in the file, no --pooling flag needed)
Dimensions 768
llama-server -m codebert-base-Q8_0.gguf --embeddings --port 8080
curl -s localhost:8080/v1/embeddings -H 'Content-Type: application/json' -d '{"input":["def f(): pass"]}'

How it was converted

convert_hf_to_gguf.py (llama.cpp, 2026-10-01) with --outtype q8_0, after two additions to the source directory:

  • a tokenizer.json, written by RobertaTokenizerFast.save_pretrained. The upstream repo ships only vocab.json + merges.txt, and without tokenizer.json the converter falls back to a WordPiece vocabulary and tokenizes code wrongly;
  • a sentence-transformers modules.json + 1_Pooling/config.json declaring mean pooling, so the pooling type is written into the GGUF.

Verification

40 random code/console/output blocks (first 600 characters), embedded by this file on llama-server (CPU) and by the original model in PyTorch fp32 with masked mean pooling over the last hidden state. Token ids are identical; cosine similarity min 0.9998, median 0.99994.

The weights and their licence (MIT) are Microsoft's; see the CodeBERT paper.

Downloads last month
77
GGUF
Model size
0.1B params
Architecture
bert
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for ChrisGVE/codebert-base-Q8_0-GGUF

Quantized
(7)
this model

Paper for ChrisGVE/codebert-base-Q8_0-GGUF