Text Generation
Transformers
Safetensors
GGUF
llama
chatbot
multilingual
arabic
french
tamazight
english
conversational
text-generation-inference
4-bit precision
bitsandbytes
Instructions to use kaisser/LLM-Maroc with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use kaisser/LLM-Maroc with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="kaisser/LLM-Maroc") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("kaisser/LLM-Maroc") model = AutoModelForCausalLM.from_pretrained("kaisser/LLM-Maroc", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use kaisser/LLM-Maroc with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf kaisser/LLM-Maroc:BF16 # Run inference directly in the terminal: llama cli -hf kaisser/LLM-Maroc:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf kaisser/LLM-Maroc:BF16 # Run inference directly in the terminal: llama cli -hf kaisser/LLM-Maroc:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf kaisser/LLM-Maroc:BF16 # Run inference directly in the terminal: ./llama-cli -hf kaisser/LLM-Maroc:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf kaisser/LLM-Maroc:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf kaisser/LLM-Maroc:BF16
Use Docker
docker model run hf.co/kaisser/LLM-Maroc:BF16
- LM Studio
- Jan
- vLLM
How to use kaisser/LLM-Maroc with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "kaisser/LLM-Maroc" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "kaisser/LLM-Maroc", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/kaisser/LLM-Maroc:BF16
- SGLang
How to use kaisser/LLM-Maroc with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "kaisser/LLM-Maroc" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "kaisser/LLM-Maroc", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "kaisser/LLM-Maroc" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "kaisser/LLM-Maroc", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use kaisser/LLM-Maroc with Ollama:
ollama run hf.co/kaisser/LLM-Maroc:BF16
- Unsloth Studio
How to use kaisser/LLM-Maroc with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for kaisser/LLM-Maroc to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for kaisser/LLM-Maroc to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for kaisser/LLM-Maroc to start chatting
- Docker Model Runner
How to use kaisser/LLM-Maroc with Docker Model Runner:
docker model run hf.co/kaisser/LLM-Maroc:BF16
- Lemonade
How to use kaisser/LLM-Maroc with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull kaisser/LLM-Maroc:BF16
Run and chat with the model
lemonade run user.LLM-Maroc-BF16
List all available models
lemonade list
- Atomic Chat
| int main(void) { | |
| common_params params; | |
| printf("test-arg-parser: make sure there is no duplicated arguments in any examples\n\n"); | |
| for (int ex = 0; ex < LLAMA_EXAMPLE_COUNT; ex++) { | |
| try { | |
| auto ctx_arg = common_params_parser_init(params, (enum llama_example)ex); | |
| std::unordered_set<std::string> seen_args; | |
| std::unordered_set<std::string> seen_env_vars; | |
| for (const auto & opt : ctx_arg.options) { | |
| // check for args duplications | |
| for (const auto & arg : opt.args) { | |
| if (seen_args.find(arg) == seen_args.end()) { | |
| seen_args.insert(arg); | |
| } else { | |
| fprintf(stderr, "test-arg-parser: found different handlers for the same argument: %s", arg); | |
| exit(1); | |
| } | |
| } | |
| // check for env var duplications | |
| if (opt.env) { | |
| if (seen_env_vars.find(opt.env) == seen_env_vars.end()) { | |
| seen_env_vars.insert(opt.env); | |
| } else { | |
| fprintf(stderr, "test-arg-parser: found different handlers for the same env var: %s", opt.env); | |
| exit(1); | |
| } | |
| } | |
| } | |
| } catch (std::exception & e) { | |
| printf("%s\n", e.what()); | |
| assert(false); | |
| } | |
| } | |
| auto list_str_to_char = [](std::vector<std::string> & argv) -> std::vector<char *> { | |
| std::vector<char *> res; | |
| for (auto & arg : argv) { | |
| res.push_back(const_cast<char *>(arg.data())); | |
| } | |
| return res; | |
| }; | |
| std::vector<std::string> argv; | |
| printf("test-arg-parser: test invalid usage\n\n"); | |
| // missing value | |
| argv = {"binary_name", "-m"}; | |
| assert(false == common_params_parse(argv.size(), list_str_to_char(argv).data(), params, LLAMA_EXAMPLE_COMMON)); | |
| // wrong value (int) | |
| argv = {"binary_name", "-ngl", "hello"}; | |
| assert(false == common_params_parse(argv.size(), list_str_to_char(argv).data(), params, LLAMA_EXAMPLE_COMMON)); | |
| // wrong value (enum) | |
| argv = {"binary_name", "-sm", "hello"}; | |
| assert(false == common_params_parse(argv.size(), list_str_to_char(argv).data(), params, LLAMA_EXAMPLE_COMMON)); | |
| // non-existence arg in specific example (--draft cannot be used outside llama-speculative) | |
| argv = {"binary_name", "--draft", "123"}; | |
| assert(false == common_params_parse(argv.size(), list_str_to_char(argv).data(), params, LLAMA_EXAMPLE_EMBEDDING)); | |
| printf("test-arg-parser: test valid usage\n\n"); | |
| argv = {"binary_name", "-m", "model_file.gguf"}; | |
| assert(true == common_params_parse(argv.size(), list_str_to_char(argv).data(), params, LLAMA_EXAMPLE_COMMON)); | |
| assert(params.model.path == "model_file.gguf"); | |
| argv = {"binary_name", "-t", "1234"}; | |
| assert(true == common_params_parse(argv.size(), list_str_to_char(argv).data(), params, LLAMA_EXAMPLE_COMMON)); | |
| assert(params.cpuparams.n_threads == 1234); | |
| argv = {"binary_name", "--verbose"}; | |
| assert(true == common_params_parse(argv.size(), list_str_to_char(argv).data(), params, LLAMA_EXAMPLE_COMMON)); | |
| assert(params.verbosity > 1); | |
| argv = {"binary_name", "-m", "abc.gguf", "--predict", "6789", "--batch-size", "9090"}; | |
| assert(true == common_params_parse(argv.size(), list_str_to_char(argv).data(), params, LLAMA_EXAMPLE_COMMON)); | |
| assert(params.model.path == "abc.gguf"); | |
| assert(params.n_predict == 6789); | |
| assert(params.n_batch == 9090); | |
| // --draft cannot be used outside llama-speculative | |
| argv = {"binary_name", "--draft", "123"}; | |
| assert(true == common_params_parse(argv.size(), list_str_to_char(argv).data(), params, LLAMA_EXAMPLE_SPECULATIVE)); | |
| assert(params.speculative.n_max == 123); | |
| // skip this part on windows, because setenv is not supported | |
| printf("test-arg-parser: skip on windows build\n"); | |
| printf("test-arg-parser: test environment variables (valid + invalid usages)\n\n"); | |
| setenv("LLAMA_ARG_THREADS", "blah", true); | |
| argv = {"binary_name"}; | |
| assert(false == common_params_parse(argv.size(), list_str_to_char(argv).data(), params, LLAMA_EXAMPLE_COMMON)); | |
| setenv("LLAMA_ARG_MODEL", "blah.gguf", true); | |
| setenv("LLAMA_ARG_THREADS", "1010", true); | |
| argv = {"binary_name"}; | |
| assert(true == common_params_parse(argv.size(), list_str_to_char(argv).data(), params, LLAMA_EXAMPLE_COMMON)); | |
| assert(params.model.path == "blah.gguf"); | |
| assert(params.cpuparams.n_threads == 1010); | |
| printf("test-arg-parser: test environment variables being overwritten\n\n"); | |
| setenv("LLAMA_ARG_MODEL", "blah.gguf", true); | |
| setenv("LLAMA_ARG_THREADS", "1010", true); | |
| argv = {"binary_name", "-m", "overwritten.gguf"}; | |
| assert(true == common_params_parse(argv.size(), list_str_to_char(argv).data(), params, LLAMA_EXAMPLE_COMMON)); | |
| assert(params.model.path == "overwritten.gguf"); | |
| assert(params.cpuparams.n_threads == 1010); | |
| if (common_has_curl()) { | |
| printf("test-arg-parser: test curl-related functions\n\n"); | |
| const char * GOOD_URL = "https://ggml.ai/"; | |
| const char * BAD_URL = "https://www.google.com/404"; | |
| const char * BIG_FILE = "https://huggingface.co/ggerganov/whisper.cpp/resolve/main/ggml-large-v1.bin"; | |
| { | |
| printf("test-arg-parser: test good URL\n\n"); | |
| auto res = common_remote_get_content(GOOD_URL, {}); | |
| assert(res.first == 200); | |
| assert(res.second.size() > 0); | |
| std::string str(res.second.data(), res.second.size()); | |
| assert(str.find("llama.cpp") != std::string::npos); | |
| } | |
| { | |
| printf("test-arg-parser: test bad URL\n\n"); | |
| auto res = common_remote_get_content(BAD_URL, {}); | |
| assert(res.first == 404); | |
| } | |
| { | |
| printf("test-arg-parser: test max size error\n"); | |
| common_remote_params params; | |
| params.max_size = 1; | |
| try { | |
| common_remote_get_content(GOOD_URL, params); | |
| assert(false && "it should throw an error"); | |
| } catch (std::exception & e) { | |
| printf(" expected error: %s\n\n", e.what()); | |
| } | |
| } | |
| { | |
| printf("test-arg-parser: test timeout error\n"); | |
| common_remote_params params; | |
| params.timeout = 1; | |
| try { | |
| common_remote_get_content(BIG_FILE, params); | |
| assert(false && "it should throw an error"); | |
| } catch (std::exception & e) { | |
| printf(" expected error: %s\n\n", e.what()); | |
| } | |
| } | |
| } else { | |
| printf("test-arg-parser: no curl, skipping curl-related functions\n"); | |
| } | |
| printf("test-arg-parser: all tests OK\n\n"); | |
| } | |