ggml-quantization / build.toml

Commit History

add mul_mat_id, and compile upstream's kernels as they ship
9534e0b

Marc Sun commited on

vendor only the files the build reaches
091adac

Marc Sun commited on

drop the cuda kernel: metal is the only backend built for now
77f577c

Marc Sun commited on

name the package after the repo it publishes to
55a081e

marcsun13 HF Staff commited on

point the hub upload at the renamed repo
88bf4fc

marcsun13 HF Staff commited on

trim the metal shader to the dispatched kernels
87cb5c1
verified

marcsun13 HF Staff commited on

full tree: sources, vendored ggml, and all built variants
90ac035
verified

marcsun13 HF Staff commited on

Add Metal backend; vendor whole llama.cpp trees
d297b20

marcsun13 HF Staff commited on

GGUF kernels: dequantize + fused gemv over packed blocks
97231ed
verified

marcsun13 HF Staff commited on

GGUF kernels: dequantize + fused gemv over packed blocks, 9 CUDA variants
2c20074
verified

marcsun13 HF Staff commited on