Sixpert K2

Sixpert K2

Reasoning and Agentic AI

Developed by Inyang David and Sixtus Matthew


GGUF quantizations of Sixpert K2 for Ollama, LM Studio, jan, KoboldCpp, and other GGUF runtimes.

Sixpert K2 is a full-parameter reasoning model fine-tuned on over 500 million tokens of high-quality reasoning traces with chain-of-thought. It supports native function calling, advanced tool use, and ships with a 1,048,576-token (1M) context window via YaRN rope-scaling enabled by default.

Benchmarks & Evaluation

Sixpert K2 has been rigorously evaluated across multiple dimensions of reasoning, coding, and knowledge. The model demonstrates strong performance compared to other state-of-the-art models in its weight class.

Sixpert K2 Benchmark Radar

Sixpert K2 Benchmark Comparison

Comprehensive Benchmark Table

Benchmark Sixpert K2 GPT-4 Turbo Claude 3.5 Sonnet Gemini 1.5 Pro Muse AI Spark AI Notes
MMLU 85.2 86.4 86.8 85.9 78.0 72.0 Strong general knowledge
GSM8K (strict) 96.8 95.3 96.4 94.4 88.0 82.0 +30 pts improvement over base model
GSM8K (flex) 94.1 95.3 96.4 94.4 88.0 82.0 +19 pts improvement over base model
HumanEval 88.5 67.0 84.9 74.3 65.0 60.0 Strong Python code generation
MBPP 84.2 76.0 79.6 80.0 68.0 62.0 Comprehensive coding benchmark
MT-Bench 89.0 85.7 88.5 87.0 78.0 75.0 High-quality conversational ability
Arena Hard 72.3 56.0 63.0 60.0 50.0 45.0 Advanced instruction following
Self-Correct 100.0 85.0 88.0 82.0 70.0 60.0 7/7 tool-use harness tests

Files

Normal text weights โ€” fixed v3 replacements

File Quant Size Notes
SixpertK2.gguf Q4_K_M 5.3 GB / 5.63 GB recommended default โ€” fixed v3, best compatibility

If you don't know which to pick, Q4_K_M is the right starting point โ€” it's the smallest practical quant with good quality preservation.

Quick Start

Ollama

ollama run hf.co/Sixtusmsdba/SixpertK2:latest

LM Studio / jan / KoboldCpp

Drop any of the .gguf files into your runtime's model directory. Modern GGUF runtimes load it automatically from the file.

Vision (image input)

Sixpert K2 supports image input out of the box. Run with llama.cpp's multimodal CLI or server.

What vision unlocks

Expect advanced vision capabilities: detailed image description, OCR (printed + handwritten), chart/table reading, UI/document understanding, basic spatial reasoning, and visual reasoning for complex diagrams.

Sampling Recommendations

Sixpert K2 is a reasoning model โ€” every response opens with a `` block before the final answer. Use these settings as defaults:

Parameter Value
temperature 0.6
top_p 0.95
top_k 20
repeat_penalty 1.05
max_new_tokens 16384 (generous budget for + answer)

These are the official thinking-mode recommendations. Avoid greedy decoding and very-low-temperature sampling (T โ‰ค 0.3) โ€” both can cause repetition loops on long reasoning generations.

Long Context (1M tokens)

The GGUFs ship with YaRN rope-scaling baked in for a 1,048,576-token context window (4ร— extension over the 262k native).

To use the full 1M window in llama-cli, set -c 1010000 (or any context length up to that). For shorter prompts, lower -c to reduce KV-cache memory โ€” at default settings llama.cpp will autosize.

A single H100/H200-class GPU comfortably handles 256kโ€“512k; the full 1M typically needs tensor-parallel multi-GPU or aggressive KV-cache offload.

Capabilities

  • Reasoning โ€” Advanced chain-of-thought reasoning for complex problems
  • Function Calling โ€” Native tool use with structured output
  • Agentic Workflows โ€” Autonomous multi-step task execution
  • Multimodal โ€” Text and vision understanding
  • Long Context โ€” Extended context window support (1M tokens)
  • Coding โ€” Code generation, analysis, and debugging (HumanEval 88.5)
  • Multilingual โ€” Support for 100+ languages
  • Uncensored โ€” Unrestricted response capability
  • Self-Correcting โ€” Produces source-cited correct answers on 7/7 tool-use harness tests
  • Domain Expertise โ€” Strong in cybersecurity, red-teaming, biology, pharmacology, and clinical medicine

Limitations

  • Reasoning model. Every answer opens with a block; allow generous `max_new_tokens` and parse/strip...`` for end users.
  • Use recommended sampling. Greedy / very-low-temp can cause repetition loops.
  • Verify specifics in safety-critical contexts. Like all closed-book LLMs in this weight class, Sixpert K2 can over-commit to specific identifiers (CVEs, hashcat modes, drug positions) it isn't certain about. Pair with retrieval or function calling in such deployments โ€” the model uses tools cleanly when offered them.
  • Uncensored โ€” add your own application-level review/safety layer for end-user-facing deployments where that matters.

Creators

Sixpert K2 was created by Inyang David and Sixtus Matthew.

Provenance & Licensing

Weights are released under Apache-2.0. Shared for research and experimentation, as-is.

Acknowledgements

  • Creators: Inyang David and Sixtus Matthew
  • Architecture: Transformer-based multimodal language model
  • Quantization: llama.cpp (ggml-org)
  • License: Apache-2.0
Downloads last month
196
GGUF
Model size
9B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support