Measure the GPU offload instead of estimating it, and keep the development scratch out of the repository 9f85661 User1342 commited on Sep 2
Offload what the card can hold, survive a dropped connection, and stop sending prompts the model cannot read 95200a7 User1342 commited on Sep 2
Pentest fixes: pin the runtime source host, evict rather than refuse; say on the page that llama.cpp is fetched eab2534 User1342 commited on Sep 1
Workers print their own code and are claimed from the site; fetch and verify llama.cpp 3c241d8 User1342 commited on Sep 1