devaloper commited on
Commit
7e6ab14
·
verified ·
1 Parent(s): 5a96ac1

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +111 -3
README.md CHANGED
@@ -1,3 +1,111 @@
1
- ---
2
- license: mit
3
- ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ library_name: transformers
4
+ base_model: Qwen/Qwen3-14B
5
+ tags:
6
+ - code
7
+ - qwen3
8
+ - gguf
9
+ - fine-tuned
10
+ model-index:
11
+ - name: Codeas Model
12
+ results: []
13
+ pipeline_tag: text-generation
14
+ language:
15
+ - en
16
+ ---
17
+
18
+ # Codeas Model
19
+
20
+ A fine-tuned **Qwen3-14B** model optimized for code generation and reasoning tasks. Available in GGUF Q6_K format for efficient local inference.
21
+
22
+ ## Model Details
23
+
24
+ | | |
25
+ |---|---|
26
+ | **Base Model** | Qwen3-14B |
27
+ | **Parameters** | ~15B |
28
+ | **Architecture** | Qwen3 (GQA, RoPE) |
29
+ | **Context Length** | 40,960 tokens |
30
+ | **Precision** | BF16 (original), Q6_K (GGUF) |
31
+ | **License** | Apache 2.0 |
32
+
33
+ ## Architecture
34
+
35
+ - 40 transformer blocks
36
+ - 40 attention heads, 8 KV heads (Grouped Query Attention)
37
+ - 5,120 hidden size / 17,408 FFN size
38
+ - RoPE with 1M frequency base
39
+ - SiLU activation
40
+ - 151,936 vocab size (GPT-2 tokenizer, Qwen2 pre-tokenizer)
41
+
42
+ ## Capabilities
43
+
44
+ - Chain-of-thought reasoning via `<think>` blocks
45
+ - Tool/function calling via `<tool_call>` format
46
+ - Thinking mode can be toggled on/off per request
47
+
48
+ ## GGUF Quantizations
49
+
50
+ | File | Quant | Size | Quality |
51
+ |------|-------|------|---------|
52
+ | `codeas-model-Q6_K.gguf` | Q6_K | 12.1 GB | Near-lossless |
53
+
54
+ ## Usage
55
+
56
+ ### llama.cpp
57
+
58
+ ```bash
59
+ ./llama-cli -m codeas-model-Q6_K.gguf -p "Write a Python function to merge two sorted lists" -n 512
60
+ ```
61
+
62
+ ### Ollama
63
+
64
+ ```bash
65
+ ollama create codeas -f Modelfile
66
+ ollama run codeas
67
+ ```
68
+
69
+ ### Transformers (safetensors)
70
+
71
+ ```python
72
+ from transformers import AutoModelForCausalLM, AutoTokenizer
73
+
74
+ model = AutoModelForCausalLM.from_pretrained("devaloper/thinkncode", torch_dtype="bfloat16", device_map="auto")
75
+ tokenizer = AutoTokenizer.from_pretrained("devaloper/thinkncode")
76
+
77
+ messages = [{"role": "user", "content": "Write a binary search in Rust"}]
78
+ inputs = tokenizer.apply_chat_template(messages, return_tensors="pt", add_generation_prompt=True).to(model.device)
79
+ outputs = model.generate(inputs, max_new_tokens=1024)
80
+ print(tokenizer.decode(outputs[0][inputs.shape[-1]:], skip_special_tokens=True))
81
+ ```
82
+
83
+ ## Hardware Requirements
84
+
85
+ | Format | VRAM / RAM |
86
+ |--------|-----------|
87
+ | Q6_K GGUF | ~14 GB |
88
+ | BF16 (full) | ~30 GB |
89
+
90
+ ## Training
91
+
92
+ | | |
93
+ |---|---|
94
+ | **Method** | Full fine-tune (no LoRA) |
95
+ | **Framework** | Axolotl 0.13.0 + Transformers 4.55.4 |
96
+ | **Hardware** | 8x GPU (FSDP) |
97
+ | **Optimizer** | AdamW (fused) |
98
+ | **LR Schedule** | Cosine, 1e-5 peak |
99
+ | **Sequence Length** | 8,192 |
100
+ | **Batch Size** | 24 (3 per device) |
101
+ | **Epochs** | 3 |
102
+ | **Precision** | BF16 + TF32 |
103
+ | **Techniques** | Flash Attention, Sample Packing, Gradient Checkpointing, Activation Offloading |
104
+
105
+ ## Sampling Defaults
106
+
107
+ ```
108
+ temperature: 0.6
109
+ top_p: 0.95
110
+ top_k: 20
111
+ ```