Text Generation
Safetensors
GGUF
English
reasoning
reinforcement-learning
grpo
small-language-model
samsung-ennovatex
conversational
Instructions to use OmnipotentFool/Aurvion with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use OmnipotentFool/Aurvion with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf OmnipotentFool/Aurvion:Q4_K_M # Run inference directly in the terminal: llama cli -hf OmnipotentFool/Aurvion:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf OmnipotentFool/Aurvion:Q4_K_M # Run inference directly in the terminal: llama cli -hf OmnipotentFool/Aurvion:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf OmnipotentFool/Aurvion:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf OmnipotentFool/Aurvion:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf OmnipotentFool/Aurvion:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf OmnipotentFool/Aurvion:Q4_K_M
Use Docker
docker model run hf.co/OmnipotentFool/Aurvion:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use OmnipotentFool/Aurvion with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "OmnipotentFool/Aurvion" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "OmnipotentFool/Aurvion", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/OmnipotentFool/Aurvion:Q4_K_M
- Ollama
How to use OmnipotentFool/Aurvion with Ollama:
ollama run hf.co/OmnipotentFool/Aurvion:Q4_K_M
- Unsloth Studio
How to use OmnipotentFool/Aurvion with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for OmnipotentFool/Aurvion to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for OmnipotentFool/Aurvion to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for OmnipotentFool/Aurvion to start chatting
- Docker Model Runner
How to use OmnipotentFool/Aurvion with Docker Model Runner:
docker model run hf.co/OmnipotentFool/Aurvion:Q4_K_M
- Lemonade
How to use OmnipotentFool/Aurvion with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull OmnipotentFool/Aurvion:Q4_K_M
Run and chat with the model
lemonade run user.Aurvion-Q4_K_M
List all available models
lemonade list
- Atomic Chat
sequenceDiagram
participant UI as π§© +layout.svelte
participant serverStore as ποΈ serverStore
participant PropsSvc as βοΈ PropsService
participant API as π llama-server
Note over serverStore: State:<br/>props: ApiLlamaCppServerProps | null<br/>loading, error<br/>role: ServerRole | null (MODEL | ROUTER)<br/>fetchPromise (deduplication)
%% βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Note over UI,API: π INITIALIZATION
%% βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
UI->>serverStore: fetch()
activate serverStore
alt fetchPromise exists (already fetching)
serverStore-->>UI: return fetchPromise
Note right of serverStore: Deduplicate concurrent calls
end
serverStore->>serverStore: loading = true
serverStore->>serverStore: fetchPromise = new Promise()
serverStore->>PropsSvc: fetch()
PropsSvc->>API: GET /props
API-->>PropsSvc: ApiLlamaCppServerProps
Note right of API: {role, model_path, model_alias,<br/>modalities, default_generation_settings, ...}
PropsSvc-->>serverStore: props
serverStore->>serverStore: props = $state(data)
serverStore->>serverStore: detectRole(props)
Note right of serverStore: role = props.role === "router"<br/> ? ServerRole.ROUTER<br/> : ServerRole.MODEL
serverStore->>serverStore: loading = false
serverStore->>serverStore: fetchPromise = null
deactivate serverStore
%% βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Note over UI,API: π COMPUTED GETTERS
%% βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Note over serverStore: Getters from props:
rect rgb(240, 255, 240)
Note over serverStore: defaultParams<br/>β props.default_generation_settings.params<br/>(temperature, top_p, top_k, etc.)
end
rect rgb(240, 255, 240)
Note over serverStore: contextSize<br/>β props.default_generation_settings.n_ctx
end
rect rgb(255, 240, 240)
Note over serverStore: isRouterMode<br/>β role === ServerRole.ROUTER
end
rect rgb(255, 240, 240)
Note over serverStore: isModelMode<br/>β role === ServerRole.MODEL
end
%% βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Note over UI,API: π RELATIONSHIPS
%% βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Note over serverStore: Used by:
Note right of serverStore: - modelsStore: role detection, MODEL mode modalities<br/>- settingsStore: syncWithServerDefaults (defaultParams)<br/>- chatStore: contextSize for processing state<br/>- UI components: isRouterMode for conditional rendering
%% βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Note over UI,API: β ERROR HANDLING
%% βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
Note over serverStore: getErrorMessage(): string | null<br/>Returns formatted error for UI display
Note over serverStore: clear(): void<br/>Resets all state (props, error, loading, role)