Text Generation
Transformers
Safetensors
GGUF
llama
chatbot
multilingual
arabic
french
tamazight
english
conversational
text-generation-inference
4-bit precision
bitsandbytes
Instructions to use kaisser/LLM-Maroc with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use kaisser/LLM-Maroc with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="kaisser/LLM-Maroc") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("kaisser/LLM-Maroc") model = AutoModelForCausalLM.from_pretrained("kaisser/LLM-Maroc", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use kaisser/LLM-Maroc with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf kaisser/LLM-Maroc:BF16 # Run inference directly in the terminal: llama cli -hf kaisser/LLM-Maroc:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf kaisser/LLM-Maroc:BF16 # Run inference directly in the terminal: llama cli -hf kaisser/LLM-Maroc:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf kaisser/LLM-Maroc:BF16 # Run inference directly in the terminal: ./llama-cli -hf kaisser/LLM-Maroc:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf kaisser/LLM-Maroc:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf kaisser/LLM-Maroc:BF16
Use Docker
docker model run hf.co/kaisser/LLM-Maroc:BF16
- LM Studio
- Jan
- vLLM
How to use kaisser/LLM-Maroc with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "kaisser/LLM-Maroc" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "kaisser/LLM-Maroc", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/kaisser/LLM-Maroc:BF16
- SGLang
How to use kaisser/LLM-Maroc with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "kaisser/LLM-Maroc" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "kaisser/LLM-Maroc", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "kaisser/LLM-Maroc" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "kaisser/LLM-Maroc", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use kaisser/LLM-Maroc with Ollama:
ollama run hf.co/kaisser/LLM-Maroc:BF16
- Unsloth Studio
How to use kaisser/LLM-Maroc with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for kaisser/LLM-Maroc to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for kaisser/LLM-Maroc to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for kaisser/LLM-Maroc to start chatting
- Docker Model Runner
How to use kaisser/LLM-Maroc with Docker Model Runner:
docker model run hf.co/kaisser/LLM-Maroc:BF16
- Lemonade
How to use kaisser/LLM-Maroc with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull kaisser/LLM-Maroc:BF16
Run and chat with the model
lemonade run user.LLM-Maroc-BF16
List all available models
lemonade list
- Atomic Chat
| // @ts-expect-error this package does not have typing | |
| import TextLineStream from 'textlinestream'; | |
| import { | |
| APIMessage, | |
| APIMessageContentPart, | |
| LlamaCppServerProps, | |
| Message, | |
| } from './types'; | |
| // ponyfill for missing ReadableStream asyncIterator on Safari | |
| import { asyncIterator } from '@sec-ant/readable-stream/ponyfill/asyncIterator'; | |
| // eslint-disable-next-line @typescript-eslint/no-explicit-any | |
| export const isString = (x: any) => !!x.toLowerCase; | |
| // eslint-disable-next-line @typescript-eslint/no-explicit-any | |
| export const isBoolean = (x: any) => x === true || x === false; | |
| // eslint-disable-next-line @typescript-eslint/no-explicit-any | |
| export const isNumeric = (n: any) => !isString(n) && !isNaN(n) && !isBoolean(n); | |
| export const escapeAttr = (str: string) => | |
| str.replace(/>/g, '>').replace(/"/g, '"'); | |
| // wrapper for SSE | |
| export async function* getSSEStreamAsync(fetchResponse: Response) { | |
| if (!fetchResponse.body) throw new Error('Response body is empty'); | |
| const lines: ReadableStream<string> = fetchResponse.body | |
| .pipeThrough(new TextDecoderStream()) | |
| .pipeThrough(new TextLineStream()); | |
| // @ts-expect-error asyncIterator complains about type, but it should work | |
| for await (const line of asyncIterator(lines)) { | |
| //if (isDev) console.log({ line }); | |
| if (line.startsWith('data:') && !line.endsWith('[DONE]')) { | |
| const data = JSON.parse(line.slice(5)); | |
| yield data; | |
| } else if (line.startsWith('error:')) { | |
| const data = JSON.parse(line.slice(6)); | |
| throw new Error(data.message || 'Unknown error'); | |
| } | |
| } | |
| } | |
| // copy text to clipboard | |
| export const copyStr = (textToCopy: string) => { | |
| // Navigator clipboard api needs a secure context (https) | |
| if (navigator.clipboard && window.isSecureContext) { | |
| navigator.clipboard.writeText(textToCopy); | |
| } else { | |
| // Use the 'out of viewport hidden text area' trick | |
| const textArea = document.createElement('textarea'); | |
| textArea.value = textToCopy; | |
| // Move textarea out of the viewport so it's not visible | |
| textArea.style.position = 'absolute'; | |
| textArea.style.left = '-999999px'; | |
| document.body.prepend(textArea); | |
| textArea.select(); | |
| document.execCommand('copy'); | |
| } | |
| }; | |
| /** | |
| * filter out redundant fields upon sending to API | |
| * also format extra into text | |
| */ | |
| export function normalizeMsgsForAPI(messages: Readonly<Message[]>) { | |
| return messages.map((msg) => { | |
| if (msg.role !== 'user' || !msg.extra) { | |
| return { | |
| role: msg.role, | |
| content: msg.content, | |
| } as APIMessage; | |
| } | |
| // extra content first, then user text message in the end | |
| // this allow re-using the same cache prefix for long context | |
| const contentArr: APIMessageContentPart[] = []; | |
| for (const extra of msg.extra ?? []) { | |
| if (extra.type === 'context') { | |
| contentArr.push({ | |
| type: 'text', | |
| text: extra.content, | |
| }); | |
| } else if (extra.type === 'textFile') { | |
| contentArr.push({ | |
| type: 'text', | |
| text: `File: ${extra.name}\nContent:\n\n${extra.content}`, | |
| }); | |
| } else if (extra.type === 'imageFile') { | |
| contentArr.push({ | |
| type: 'image_url', | |
| image_url: { url: extra.base64Url }, | |
| }); | |
| } else if (extra.type === 'audioFile') { | |
| contentArr.push({ | |
| type: 'input_audio', | |
| input_audio: { | |
| data: extra.base64Data, | |
| format: /wav/.test(extra.mimeType) ? 'wav' : 'mp3', | |
| }, | |
| }); | |
| } else { | |
| throw new Error('Unknown extra type'); | |
| } | |
| } | |
| // add user message to the end | |
| contentArr.push({ | |
| type: 'text', | |
| text: msg.content, | |
| }); | |
| return { | |
| role: msg.role, | |
| content: contentArr, | |
| }; | |
| }) as APIMessage[]; | |
| } | |
| /** | |
| * recommended for DeepsSeek-R1, filter out content between <think> and </think> tags | |
| */ | |
| export function filterThoughtFromMsgs(messages: APIMessage[]) { | |
| console.debug({ messages }); | |
| return messages.map((msg) => { | |
| if (msg.role !== 'assistant') { | |
| return msg; | |
| } | |
| // assistant message is always a string | |
| const contentStr = msg.content as string; | |
| return { | |
| role: msg.role, | |
| content: | |
| msg.role === 'assistant' | |
| ? contentStr.split('</think>').at(-1)!.trim() | |
| : contentStr, | |
| } as APIMessage; | |
| }); | |
| } | |
| export function classNames(classes: Record<string, boolean>): string { | |
| return Object.entries(classes) | |
| .filter(([_, value]) => value) | |
| .map(([key, _]) => key) | |
| .join(' '); | |
| } | |
| export const delay = (ms: number) => | |
| new Promise((resolve) => setTimeout(resolve, ms)); | |
| export const throttle = <T extends unknown[]>( | |
| callback: (...args: T) => void, | |
| delay: number | |
| ) => { | |
| let isWaiting = false; | |
| return (...args: T) => { | |
| if (isWaiting) { | |
| return; | |
| } | |
| callback(...args); | |
| isWaiting = true; | |
| setTimeout(() => { | |
| isWaiting = false; | |
| }, delay); | |
| }; | |
| }; | |
| export const cleanCurrentUrl = (removeQueryParams: string[]) => { | |
| const url = new URL(window.location.href); | |
| removeQueryParams.forEach((param) => { | |
| url.searchParams.delete(param); | |
| }); | |
| window.history.replaceState({}, '', url.toString()); | |
| }; | |
| export const getServerProps = async ( | |
| baseUrl: string, | |
| apiKey?: string | |
| ): Promise<LlamaCppServerProps> => { | |
| try { | |
| const response = await fetch(`${baseUrl}/props`, { | |
| headers: { | |
| 'Content-Type': 'application/json', | |
| ...(apiKey ? { Authorization: `Bearer ${apiKey}` } : {}), | |
| }, | |
| }); | |
| if (!response.ok) { | |
| throw new Error('Failed to fetch server props'); | |
| } | |
| const data = await response.json(); | |
| return data as LlamaCppServerProps; | |
| } catch (error) { | |
| console.error('Error fetching server props:', error); | |
| throw error; | |
| } | |
| }; | |