SixpertK1 / docs /usage_guide.md
SixpertAI's picture
Upload docs/usage_guide.md with huggingface_hub
d5711b0 verified
|
Raw
History Blame Contribute Delete
3.82 kB

Sixpert K1 - Complete Usage Guide

Quick Start

Option 1: Ollama (Easiest)

# Install Ollama
curl -fsSL https://ollama.com/install.sh | sh

# Download and import the model
ollama create sixpert-k1 -f OllamaModelfile

# Or if GGUF is in Ollama library:
# ollama run sixpert-k1

# Chat
ollama run sixpert-k1

Option 2: llama-cpp-python (Python)

pip install llama-cpp-python
python examples/generate.py --prompt "Hello, who are you?"

Option 3: API Server

pip install llama-cpp-python
python examples/api_server.py --model SixpertK1.gguf

Option 4: LM Studio

  1. Download LM Studio from https://lmstudio.ai
  2. Import SixpertK1.gguf
  3. Start chatting with the Sixpert K1 preset

Chat Format

Sixpert K1 uses the following chat template:

<|im_start|>system
You are a helpful assistant.<|im_end|>
<|im_start|>user
What is quantum computing?<|im_end|>
<|im_start|>assistant
Quantum computing uses quantum mechanical phenomena...<|im_end|>

Recommended Settings

Parameter Value Notes
temperature 0.7 Good balance of creativity and accuracy
top_p 0.8 Nucleus sampling
top_k 40 Limit token selection
repeat_penalty 1.05 Prevent repetition
max_tokens 8192 Max output length
context_size 131072 Full context window

Function Calling

Sixpert K1 supports native function calling. See examples/function_calling.py for a complete implementation.

Tool Format

{
  "type": "function",
  "function": {
    "name": "search",
    "description": "Search for information",
    "parameters": {
      "type": "object",
      "properties": {
        "query": {"type": "string"}
      },
      "required": ["query"]
    }
  }
}

Vision / Multimodal

Sixpert K1 can understand images. See examples/vision_example.py for implementation details.

response = llm.create_chat_completion(
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "Describe this image"},
            {"type": "image_url", "image_url": {"url": "data:image/png;base64,..."}},
        ]
    }]
)

Integration Examples

OpenAI-Compatible Client

from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="not-needed")

response = client.chat.completions.create(
    model="sixpert-k1",
    messages=[{"role": "user", "content": "Explain recursion"}],
    temperature=0.7,
)
print(response.choices[0].message.content)

LangChain Integration

from langchain.llms import LlamaCpp

llm = LlamaCpp(
    model_path="SixpertK1.gguf",
    temperature=0.7,
    n_ctx=131072,
    n_gpu_layers=-1,
)

result = llm.invoke("What is machine learning?")
print(result)

CrewAI Agent

from crewai import Agent, Task, Crew

agent = Agent(
    role="Research Analyst",
    backstory="You are Sixpert K1, a precision logic engine",
    goal="Provide accurate, detailed analysis",
    llm=LlamaCpp(model_path="SixpertK1.gguf", temperature=0.7),
    allow_delegation=False,
)

Performance Tips

  1. GPU Offloading: Set n_gpu_layers=-1 to offload all layers to GPU
  2. Context Pruning: Use smaller context windows (8192-32768) for faster inference
  3. Batch Processing: Use the API server for batch inference
  4. Quantization: Q4_K_M is the sweet spot; upgrade to Q6_K if quality matters more

Troubleshooting

Issue Solution
Out of memory Reduce context size or use CPU-only inference
Slow generation Enable GPU offloading (n_gpu_layers=-1)
Repetitive output Increase repeat_penalty to 1.1-1.2
Hallucinations Lower temperature to 0.3-0.5
Context overflow Use 4096 context for testing, 131072 for production