Text Generation
Transformers
Safetensors
llama
causal-lm
conversational
code
fill-in-the-middle
instruct
research
experimental
text-generation-inference
Instructions to use mossez-systems/Mossez-100M-Coder-Instruct with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use mossez-systems/Mossez-100M-Coder-Instruct with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="mossez-systems/Mossez-100M-Coder-Instruct") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("mossez-systems/Mossez-100M-Coder-Instruct") model = AutoModelForCausalLM.from_pretrained("mossez-systems/Mossez-100M-Coder-Instruct", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use mossez-systems/Mossez-100M-Coder-Instruct with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "mossez-systems/Mossez-100M-Coder-Instruct" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mossez-systems/Mossez-100M-Coder-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/mossez-systems/Mossez-100M-Coder-Instruct
- SGLang
How to use mossez-systems/Mossez-100M-Coder-Instruct with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "mossez-systems/Mossez-100M-Coder-Instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mossez-systems/Mossez-100M-Coder-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "mossez-systems/Mossez-100M-Coder-Instruct" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mossez-systems/Mossez-100M-Coder-Instruct", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use mossez-systems/Mossez-100M-Coder-Instruct with Docker Model Runner:
docker model run hf.co/mossez-systems/Mossez-100M-Coder-Instruct
Publish Mossez-100M-Coder-Instruct
Browse filesVerified FP32 Safetensors release with tokenizer, model card, dataset attribution, training report, evaluation, and SHA256SUMS.
- DATASET_ATTRIBUTION.md +25 -0
- EVALUATION.md +33 -0
- LICENSE +201 -0
- NOTICE.md +8 -0
- README.md +77 -0
- SHA256SUMS +13 -0
- TRAINING_REPORT.md +28 -0
- chat_template.jinja +5 -0
- config.json +30 -0
- generation_config.json +12 -0
- model.safetensors +3 -0
- special_tokens_map.json +15 -0
- tokenizer.json +0 -0
- tokenizer_config.json +23 -0
DATASET_ATTRIBUTION.md
ADDED
|
@@ -0,0 +1,25 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Dataset attribution
|
| 2 |
+
|
| 3 |
+
The Coder-Instruct v1 SFT corpus is deterministic project-authored material
|
| 4 |
+
licensed under Apache-2.0. It imports no external instruction dataset, private
|
| 5 |
+
data, chat exports, Telegram data, or chain-of-thought material.
|
| 6 |
+
|
| 7 |
+
## Composition
|
| 8 |
+
|
| 9 |
+
- Train: 2,640 examples, 323,960 rendered tokens, 88 source groups.
|
| 10 |
+
- Validation: 330 examples, 40,504 rendered tokens, 11 source groups.
|
| 11 |
+
- Test: 330 examples, 40,518 rendered tokens, 11 source groups.
|
| 12 |
+
- Maximum rendered sequence: 149 tokens; model limit: 1,024.
|
| 13 |
+
- Eleven balanced task types: short function, completion, explanation, bug fix,
|
| 14 |
+
refactor, unit test, traceback, JSON/YAML conversion, SQL, PowerShell, and FIM repair.
|
| 15 |
+
- Source groups are disjoint across train, validation, and test.
|
| 16 |
+
|
| 17 |
+
## Gates
|
| 18 |
+
|
| 19 |
+
License/provenance, secret, email-like PII, exact deduplication, cross-split
|
| 20 |
+
source-group leakage, chain-of-thought exclusion, and chat-marker collision
|
| 21 |
+
gates passed. Bounded local validators passed for Python AST/fragments, JSON,
|
| 22 |
+
YAML, SQL, safe PowerShell, and concise text.
|
| 23 |
+
|
| 24 |
+
This conservative, template-heavy corpus validates the SFT pipeline and balanced
|
| 25 |
+
held-out evaluation. It does not establish broad coding-assistant competence.
|
EVALUATION.md
ADDED
|
@@ -0,0 +1,33 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Evaluation
|
| 2 |
+
|
| 3 |
+
Evaluation used frozen, source-group-disjoint validation and test splits with
|
| 4 |
+
assistant-only loss. Lower is better.
|
| 5 |
+
|
| 6 |
+
| Model | Validation loss | Test loss |
|
| 7 |
+
|---|---:|---:|
|
| 8 |
+
| Source Coder-Base | 2.008803 | 2.023115 |
|
| 9 |
+
| One-epoch Coder-Instruct | 0.029855 | 0.032257 |
|
| 10 |
+
|
| 11 |
+
## Per-task loss
|
| 12 |
+
|
| 13 |
+
| Task | Source val | Candidate val | Source test | Candidate test |
|
| 14 |
+
|---|---:|---:|---:|---:|
|
| 15 |
+
| bug_fix | 0.666634 | 0.002418 | 0.685901 | 0.002504 |
|
| 16 |
+
| completion | 1.098412 | 0.072219 | 1.184431 | 0.074582 |
|
| 17 |
+
| explain | 3.757432 | 0.003248 | 3.861036 | 0.003465 |
|
| 18 |
+
| fim_repair | 2.926159 | 0.171918 | 2.934412 | 0.188182 |
|
| 19 |
+
| json_yaml_conversion | 1.304223 | 0.002704 | 1.310993 | 0.002706 |
|
| 20 |
+
| refactor | 1.088105 | 0.002698 | 1.089293 | 0.002557 |
|
| 21 |
+
| shell | 2.545998 | 0.001775 | 2.548895 | 0.001855 |
|
| 22 |
+
| short_function | 1.503674 | 0.003699 | 1.520041 | 0.003951 |
|
| 23 |
+
| sql | 2.761549 | 0.018859 | 2.715693 | 0.021215 |
|
| 24 |
+
| traceback | 3.734223 | 0.002768 | 3.769516 | 0.002884 |
|
| 25 |
+
| unit_test | 1.751150 | 0.075700 | 1.721415 | 0.083690 |
|
| 26 |
+
|
| 27 |
+
Overall and all 11 per-task losses improved against both the source and the
|
| 28 |
+
100-step SFT pilot. A deterministic 11-prompt generation smoke passed all
|
| 29 |
+
local syntax/format/safety validators; 9 outputs exactly matched references.
|
| 30 |
+
|
| 31 |
+
These results come from a small, project-authored, template-heavy corpus.
|
| 32 |
+
They do not establish performance on HumanEval, MBPP, SWE-bench, security
|
| 33 |
+
tasks, repository-level work, or arbitrary real-world prompts.
|
LICENSE
ADDED
|
@@ -0,0 +1,201 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
Apache License
|
| 2 |
+
Version 2.0, January 2004
|
| 3 |
+
http://www.apache.org/licenses/
|
| 4 |
+
|
| 5 |
+
TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
|
| 6 |
+
|
| 7 |
+
1. Definitions.
|
| 8 |
+
|
| 9 |
+
"License" shall mean the terms and conditions for use, reproduction,
|
| 10 |
+
and distribution as defined by Sections 1 through 9 of this document.
|
| 11 |
+
|
| 12 |
+
"Licensor" shall mean the copyright owner or entity authorized by
|
| 13 |
+
the copyright owner that is granting the License.
|
| 14 |
+
|
| 15 |
+
"Legal Entity" shall mean the union of the acting entity and all
|
| 16 |
+
other entities that control, are controlled by, or are under common
|
| 17 |
+
control with that entity. For the purposes of this definition,
|
| 18 |
+
"control" means (i) the power, direct or indirect, to cause the
|
| 19 |
+
direction or management of such entity, whether by contract or
|
| 20 |
+
otherwise, or (ii) ownership of fifty percent (50%) or more of the
|
| 21 |
+
outstanding shares, or (iii) beneficial ownership of such entity.
|
| 22 |
+
|
| 23 |
+
"You" (or "Your") shall mean an individual or Legal Entity
|
| 24 |
+
exercising permissions granted by this License.
|
| 25 |
+
|
| 26 |
+
"Source" form shall mean the preferred form for making modifications,
|
| 27 |
+
including but not limited to software source code, documentation
|
| 28 |
+
source, and configuration files.
|
| 29 |
+
|
| 30 |
+
"Object" form shall mean any form resulting from mechanical
|
| 31 |
+
transformation or translation of a Source form, including but
|
| 32 |
+
not limited to compiled object code, generated documentation,
|
| 33 |
+
and conversions to other media types.
|
| 34 |
+
|
| 35 |
+
"Work" shall mean the work of authorship, whether in Source or
|
| 36 |
+
Object form, made available under the License, as indicated by a
|
| 37 |
+
copyright notice that is included in or attached to the work
|
| 38 |
+
(an example is provided in the Appendix below).
|
| 39 |
+
|
| 40 |
+
"Derivative Works" shall mean any work, whether in Source or Object
|
| 41 |
+
form, that is based on (or derived from) the Work and for which the
|
| 42 |
+
editorial revisions, annotations, elaborations, or other modifications
|
| 43 |
+
represent, as a whole, an original work of authorship. For the purposes
|
| 44 |
+
of this License, Derivative Works shall not include works that remain
|
| 45 |
+
separable from, or merely link (or bind by name) to the interfaces of,
|
| 46 |
+
the Work and Derivative Works thereof.
|
| 47 |
+
|
| 48 |
+
"Contribution" shall mean any work of authorship, including
|
| 49 |
+
the original version of the Work and any modifications or additions
|
| 50 |
+
to that Work or Derivative Works thereof, that is intentionally
|
| 51 |
+
submitted to Licensor for inclusion in the Work by the copyright owner
|
| 52 |
+
or by an individual or Legal Entity authorized to submit on behalf of
|
| 53 |
+
the copyright owner. For the purposes of this definition, "submitted"
|
| 54 |
+
means any form of electronic, verbal, or written communication sent
|
| 55 |
+
to the Licensor or its representatives, including but not limited to
|
| 56 |
+
communication on electronic mailing lists, source code control systems,
|
| 57 |
+
and issue tracking systems that are managed by, or on behalf of, the
|
| 58 |
+
Licensor for the purpose of discussing and improving the Work, but
|
| 59 |
+
excluding communication that is conspicuously marked or otherwise
|
| 60 |
+
designated in writing by the copyright owner as "Not a Contribution."
|
| 61 |
+
|
| 62 |
+
"Contributor" shall mean Licensor and any individual or Legal Entity
|
| 63 |
+
on behalf of whom a Contribution has been received by Licensor and
|
| 64 |
+
subsequently incorporated within the Work.
|
| 65 |
+
|
| 66 |
+
2. Grant of Copyright License. Subject to the terms and conditions of
|
| 67 |
+
this License, each Contributor hereby grants to You a perpetual,
|
| 68 |
+
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
|
| 69 |
+
copyright license to reproduce, prepare Derivative Works of,
|
| 70 |
+
publicly display, publicly perform, sublicense, and distribute the
|
| 71 |
+
Work and such Derivative Works in Source or Object form.
|
| 72 |
+
|
| 73 |
+
3. Grant of Patent License. Subject to the terms and conditions of
|
| 74 |
+
this License, each Contributor hereby grants to You a perpetual,
|
| 75 |
+
worldwide, non-exclusive, no-charge, royalty-free, irrevocable
|
| 76 |
+
(except as stated in this section) patent license to make, have made,
|
| 77 |
+
use, offer to sell, sell, import, and otherwise transfer the Work,
|
| 78 |
+
where such license applies only to those patent claims licensable
|
| 79 |
+
by such Contributor that are necessarily infringed by their
|
| 80 |
+
Contribution(s) alone or by combination of their Contribution(s)
|
| 81 |
+
with the Work to which such Contribution(s) was submitted. If You
|
| 82 |
+
institute patent litigation against any entity (including a
|
| 83 |
+
cross-claim or counterclaim in a lawsuit) alleging that the Work
|
| 84 |
+
or a Contribution incorporated within the Work constitutes direct
|
| 85 |
+
or contributory patent infringement, then any patent licenses
|
| 86 |
+
granted to You under this License for that Work shall terminate
|
| 87 |
+
as of the date such litigation is filed.
|
| 88 |
+
|
| 89 |
+
4. Redistribution. You may reproduce and distribute copies of the
|
| 90 |
+
Work or Derivative Works thereof in any medium, with or without
|
| 91 |
+
modifications, and in Source or Object form, provided that You
|
| 92 |
+
meet the following conditions:
|
| 93 |
+
|
| 94 |
+
(a) You must give any other recipients of the Work or
|
| 95 |
+
Derivative Works a copy of this License; and
|
| 96 |
+
|
| 97 |
+
(b) You must cause any modified files to carry prominent notices
|
| 98 |
+
stating that You changed the files; and
|
| 99 |
+
|
| 100 |
+
(c) You must retain, in the Source form of any Derivative Works
|
| 101 |
+
that You distribute, all copyright, patent, trademark, and
|
| 102 |
+
attribution notices from the Source form of the Work,
|
| 103 |
+
excluding those notices that do not pertain to any part of
|
| 104 |
+
the Derivative Works; and
|
| 105 |
+
|
| 106 |
+
(d) If the Work includes a "NOTICE" text file as part of its
|
| 107 |
+
distribution, then any Derivative Works that You distribute must
|
| 108 |
+
include a readable copy of the attribution notices contained
|
| 109 |
+
within such NOTICE file, excluding those notices that do not
|
| 110 |
+
pertain to any part of the Derivative Works, in at least one
|
| 111 |
+
of the following places: within a NOTICE text file distributed
|
| 112 |
+
as part of the Derivative Works; within the Source form or
|
| 113 |
+
documentation, if provided along with the Derivative Works; or,
|
| 114 |
+
within a display generated by the Derivative Works, if and
|
| 115 |
+
wherever such third-party notices normally appear. The contents
|
| 116 |
+
of the NOTICE file are for informational purposes only and
|
| 117 |
+
do not modify the License. You may add Your own attribution
|
| 118 |
+
notices within Derivative Works that You distribute, alongside
|
| 119 |
+
or as an addendum to the NOTICE text from the Work, provided
|
| 120 |
+
that such additional attribution notices cannot be construed
|
| 121 |
+
as modifying the License.
|
| 122 |
+
|
| 123 |
+
You may add Your own copyright statement to Your modifications and
|
| 124 |
+
may provide additional or different license terms and conditions
|
| 125 |
+
for use, reproduction, or distribution of Your modifications, or
|
| 126 |
+
for any such Derivative Works as a whole, provided Your use,
|
| 127 |
+
reproduction, and distribution of the Work otherwise complies with
|
| 128 |
+
the conditions stated in this License.
|
| 129 |
+
|
| 130 |
+
5. Submission of Contributions. Unless You explicitly state otherwise,
|
| 131 |
+
any Contribution intentionally submitted for inclusion in the Work
|
| 132 |
+
by You to the Licensor shall be under the terms and conditions of
|
| 133 |
+
this License, without any additional terms or conditions.
|
| 134 |
+
Notwithstanding the above, nothing herein shall supersede or modify
|
| 135 |
+
the terms of any separate license agreement you may have executed
|
| 136 |
+
with Licensor regarding such Contributions.
|
| 137 |
+
|
| 138 |
+
6. Trademarks. This License does not grant permission to use the trade
|
| 139 |
+
names, trademarks, service marks, or product names of the Licensor,
|
| 140 |
+
except as required for reasonable and customary use in describing the
|
| 141 |
+
origin of the Work and reproducing the content of the NOTICE file.
|
| 142 |
+
|
| 143 |
+
7. Disclaimer of Warranty. Unless required by applicable law or
|
| 144 |
+
agreed to in writing, Licensor provides the Work (and each
|
| 145 |
+
Contributor provides its Contributions) on an "AS IS" BASIS,
|
| 146 |
+
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
|
| 147 |
+
implied, including, without limitation, any warranties or conditions
|
| 148 |
+
of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
|
| 149 |
+
PARTICULAR PURPOSE. You are solely responsible for determining the
|
| 150 |
+
appropriateness of using or redistributing the Work and assume any
|
| 151 |
+
risks associated with Your exercise of permissions under this License.
|
| 152 |
+
|
| 153 |
+
8. Limitation of Liability. In no event and under no legal theory,
|
| 154 |
+
whether in tort (including negligence), contract, or otherwise,
|
| 155 |
+
unless required by applicable law (such as deliberate and grossly
|
| 156 |
+
negligent acts) or agreed to in writing, shall any Contributor be
|
| 157 |
+
liable to You for damages, including any direct, indirect, special,
|
| 158 |
+
incidental, or consequential damages of any character arising as a
|
| 159 |
+
result of this License or out of the use or inability to use the
|
| 160 |
+
Work (including but not limited to damages for loss of goodwill,
|
| 161 |
+
work stoppage, computer failure or malfunction, or any and all
|
| 162 |
+
other commercial damages or losses), even if such Contributor
|
| 163 |
+
has been advised of the possibility of such damages.
|
| 164 |
+
|
| 165 |
+
9. Accepting Warranty or Additional Liability. While redistributing
|
| 166 |
+
the Work or Derivative Works thereof, You may choose to offer,
|
| 167 |
+
and charge a fee for, acceptance of support, warranty, indemnity,
|
| 168 |
+
or other liability obligations and/or rights consistent with this
|
| 169 |
+
License. However, in accepting such obligations, You may act only
|
| 170 |
+
on Your own behalf and on Your sole responsibility, not on behalf
|
| 171 |
+
of any other Contributor, and only if You agree to indemnify,
|
| 172 |
+
defend, and hold each Contributor harmless for any liability
|
| 173 |
+
incurred by, or claims asserted against, such Contributor by reason
|
| 174 |
+
of your accepting any such warranty or additional liability.
|
| 175 |
+
|
| 176 |
+
END OF TERMS AND CONDITIONS
|
| 177 |
+
|
| 178 |
+
APPENDIX: How to apply the Apache License to your work.
|
| 179 |
+
|
| 180 |
+
To apply the Apache License to your work, attach the following
|
| 181 |
+
boilerplate notice, with the fields enclosed by brackets "[]"
|
| 182 |
+
replaced with your own identifying information. (Don't include
|
| 183 |
+
the brackets!) The text should be enclosed in the appropriate
|
| 184 |
+
comment syntax for the file format. We also recommend that a
|
| 185 |
+
file or class name and description of purpose be included on the
|
| 186 |
+
same "printed page" as the copyright notice for easier
|
| 187 |
+
identification within third-party archives.
|
| 188 |
+
|
| 189 |
+
Copyright [yyyy] [name of copyright owner]
|
| 190 |
+
|
| 191 |
+
Licensed under the Apache License, Version 2.0 (the "License");
|
| 192 |
+
you may not use this file except in compliance with the License.
|
| 193 |
+
You may obtain a copy of the License at
|
| 194 |
+
|
| 195 |
+
http://www.apache.org/licenses/LICENSE-2.0
|
| 196 |
+
|
| 197 |
+
Unless required by applicable law or agreed to in writing, software
|
| 198 |
+
distributed under the License is distributed on an "AS IS" BASIS,
|
| 199 |
+
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
| 200 |
+
See the License for the specific language governing permissions and
|
| 201 |
+
limitations under the License.
|
NOTICE.md
ADDED
|
@@ -0,0 +1,8 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Notice
|
| 2 |
+
|
| 3 |
+
Mossez-100M-Coder-Instruct is an experimental derivative in the Mossez-100M family.
|
| 4 |
+
|
| 5 |
+
This model is derived from Mossez-100M-Coder-Base through assistant-only SFT on deterministic project-authored Apache-2.0 examples. The general Mossez-100M-Instruct supplied only tokenizer/chat-template/release references, not source weights.
|
| 6 |
+
|
| 7 |
+
Copyright 2026 Mossez Systems. Licensed under Apache-2.0. This research model is
|
| 8 |
+
not production-ready; generated code must be reviewed, tested, and sandboxed.
|
README.md
ADDED
|
@@ -0,0 +1,77 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: apache-2.0
|
| 3 |
+
library_name: transformers
|
| 4 |
+
pipeline_tag: text-generation
|
| 5 |
+
base_model:
|
| 6 |
+
- mossez-systems/Mossez-100M-Coder-Base
|
| 7 |
+
tags:
|
| 8 |
+
- causal-lm
|
| 9 |
+
- conversational
|
| 10 |
+
- code
|
| 11 |
+
- fill-in-the-middle
|
| 12 |
+
- instruct
|
| 13 |
+
- llama
|
| 14 |
+
- research
|
| 15 |
+
- experimental
|
| 16 |
+
---
|
| 17 |
+
|
| 18 |
+
# Mossez-100M-Coder-Instruct
|
| 19 |
+
|
| 20 |
+
Mossez-100M-Coder-Instruct is an experimental 100M-parameter coding instruction
|
| 21 |
+
model with this weight lineage:
|
| 22 |
+
|
| 23 |
+
`Mossez-100M-Base -> Mossez-100M-Coder-Base -> Mossez-100M-Coder-Instruct`.
|
| 24 |
+
|
| 25 |
+
The general [`Mossez-100M-Instruct`](https://huggingface.co/mossez-systems/Mossez-100M-Instruct)
|
| 26 |
+
was used only as a tokenizer, chat-template, release, and inference reference;
|
| 27 |
+
its weights were not used as source weights for this model.
|
| 28 |
+
|
| 29 |
+
## Model details
|
| 30 |
+
|
| 31 |
+
| Property | Value |
|
| 32 |
+
|---|---:|
|
| 33 |
+
| Parameters | 100,098,048 |
|
| 34 |
+
| Architecture | Llama-compatible decoder-only Transformer |
|
| 35 |
+
| Layers / hidden size | 12 / 768 |
|
| 36 |
+
| Query / KV heads | 12 / 4 |
|
| 37 |
+
| Context length | 1,024 tokens |
|
| 38 |
+
| Vocabulary | 32,007 |
|
| 39 |
+
| Objective | Assistant-only SFT loss |
|
| 40 |
+
| Weight format | Safetensors, FP32 |
|
| 41 |
+
| License | Apache-2.0 |
|
| 42 |
+
|
| 43 |
+
## Usage
|
| 44 |
+
|
| 45 |
+
```python
|
| 46 |
+
from transformers import AutoModelForCausalLM, AutoTokenizer
|
| 47 |
+
|
| 48 |
+
model_id = "mossez-systems/Mossez-100M-Coder-Instruct"
|
| 49 |
+
tokenizer = AutoTokenizer.from_pretrained(model_id)
|
| 50 |
+
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto")
|
| 51 |
+
|
| 52 |
+
messages = [{"role": "user", "content": "Write a short Python function that adds two integers."}]
|
| 53 |
+
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
|
| 54 |
+
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
|
| 55 |
+
output = model.generate(**inputs, do_sample=False, max_new_tokens=96)
|
| 56 |
+
new_tokens = output[0, inputs.input_ids.shape[1]:]
|
| 57 |
+
print(tokenizer.decode(new_tokens, skip_special_tokens=True))
|
| 58 |
+
```
|
| 59 |
+
|
| 60 |
+
## Training and evaluation
|
| 61 |
+
|
| 62 |
+
The model was fine-tuned for one bounded epoch: 660 optimizer steps over 2,640
|
| 63 |
+
project-authored examples, using assistant-only loss. Immutable validation and
|
| 64 |
+
test sets contain 330 examples each across 11 balanced task types. See
|
| 65 |
+
[TRAINING_REPORT.md](TRAINING_REPORT.md), [EVALUATION.md](EVALUATION.md), and
|
| 66 |
+
[DATASET_ATTRIBUTION.md](DATASET_ATTRIBUTION.md).
|
| 67 |
+
|
| 68 |
+
The released `model.safetensors` SHA-256 is
|
| 69 |
+
`0aade7d070122633abccd70cc5500e5bc36c7f69fafeadd9f5ee0b5a3e0766bf`.
|
| 70 |
+
|
| 71 |
+
## Limitations
|
| 72 |
+
|
| 73 |
+
This is a small research model, not a reliable or safe production coding
|
| 74 |
+
assistant. The authored SFT corpus is balanced but narrow and template-heavy,
|
| 75 |
+
so held-out loss may overstate general-world capability. Expect repetition,
|
| 76 |
+
incorrect constants, malformed code, hallucinated APIs, weak instruction
|
| 77 |
+
following, and early EOS. Validate, test, and sandbox every output.
|
SHA256SUMS
ADDED
|
@@ -0,0 +1,13 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
b5eb03b52ce8941d22739e41bae8ec4d1ab2a529d8abecb42c6f68430616a161 chat_template.jinja
|
| 2 |
+
aae190b1daeab09fec8ab276b2a97bb60ef56773c44db361ad53fd4c594b7017 config.json
|
| 3 |
+
def1e269f57ae94c4bc348b85bef347f176bfc4d295de12ec9ee155739de2c61 DATASET_ATTRIBUTION.md
|
| 4 |
+
b0ae233b7e29a39cef6c6cc055ddbad35e5b38fa92fb3577e0ece623d2ae8569 EVALUATION.md
|
| 5 |
+
0931ff2f000839227285e26dba99974b2383f6fdd1acffd9cfacf971b22df213 generation_config.json
|
| 6 |
+
8173d5c29b4f956d532781d2b86e4e30f83e6b7878dce18c919451d6ba707c90 LICENSE
|
| 7 |
+
0aade7d070122633abccd70cc5500e5bc36c7f69fafeadd9f5ee0b5a3e0766bf model.safetensors
|
| 8 |
+
99669f6f98c26b057f81967e442d1e2ae7f0e3f6fa1999c8a3d2f043a735a25d NOTICE.md
|
| 9 |
+
b978cd04f3725a48c8b9f934f46159ae69e8c82728e9129236da0131c5c55664 README.md
|
| 10 |
+
813c2723c694d037427269cca895d01641f930d6a8f269ce7bf8e1e16e24dabe special_tokens_map.json
|
| 11 |
+
e9551d84b9947f741763bf815a2d5f6bfcc47a3b67c73fcbf386223e8ed969be tokenizer.json
|
| 12 |
+
32c07b84f72b62dad43bc3956c541bfbc7b6ac1039930a7e1b8fd2a61fbd2ae9 tokenizer_config.json
|
| 13 |
+
90bd11e7d3ec83e90acb4b5bad9c0fdc3f92a160b5970a2b93979649c141ef0f TRAINING_REPORT.md
|
TRAINING_REPORT.md
ADDED
|
@@ -0,0 +1,28 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Training report
|
| 2 |
+
|
| 3 |
+
## Lineage
|
| 4 |
+
|
| 5 |
+
- Source weights: the selected one-epoch Mossez-100M-Coder-Base.
|
| 6 |
+
- Source model SHA-256: `aba529bf10ad9f3acb5294c8bc2b4c93d20d25c6cff3a235a8659503b9ac1837`.
|
| 7 |
+
- Final model SHA-256: `0aade7d070122633abccd70cc5500e5bc36c7f69fafeadd9f5ee0b5a3e0766bf`.
|
| 8 |
+
- The general Mossez-100M-Instruct supplied no source weights.
|
| 9 |
+
|
| 10 |
+
## Run
|
| 11 |
+
|
| 12 |
+
- Objective: assistant-only supervised fine-tuning.
|
| 13 |
+
- One bounded epoch: 660 contiguous finite optimizer steps.
|
| 14 |
+
- Examples: 2,640 train; 330 validation; 330 test.
|
| 15 |
+
- Rendered training tokens: 323,960; assistant target tokens: 105,100.
|
| 16 |
+
- Precision: BF16; fused AdamW; micro-batch 1; gradient accumulation 4.
|
| 17 |
+
- Validation loss: 1.991049 at step 0 to 0.031888 at step 660.
|
| 18 |
+
- Early stopping: not triggered; validation regression count: 0.
|
| 19 |
+
- Final checkpoint: best validation checkpoint, step 660.
|
| 20 |
+
|
| 21 |
+
## Preflight and integrity
|
| 22 |
+
|
| 23 |
+
- Assistant-only masks were asserted at corpus load and before each forward pass.
|
| 24 |
+
- A real CUDA memory probe passed through micro-batch 8; training used micro-batch 1.
|
| 25 |
+
- A 10-step smoke passed, followed by a bit-exact cross-process resume test.
|
| 26 |
+
- An independent 100-step pilot improved overall and all per-task held-out losses.
|
| 27 |
+
- Corpus manifest SHA-256: `aca68b911788a1a7d93671bc27835b3d279cd612c02e8e3add91e2d3b2079084`.
|
| 28 |
+
- Tokenizer manifest SHA-256: `f24c2e10522e866a889177e3a2258d84f1d8a863d784894a72974e391c1abdec`.
|
chat_template.jinja
ADDED
|
@@ -0,0 +1,5 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{% for message in messages %}{{ "<|" ~ message.role ~ "|>
|
| 2 |
+
" ~ message.content ~ "
|
| 3 |
+
<|end|>
|
| 4 |
+
" }}{% endfor %}{% if add_generation_prompt %}{{ "<|assistant|>
|
| 5 |
+
" }}{% endif %}
|
config.json
ADDED
|
@@ -0,0 +1,30 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"architectures": [
|
| 3 |
+
"LlamaForCausalLM"
|
| 4 |
+
],
|
| 5 |
+
"attention_bias": false,
|
| 6 |
+
"attention_dropout": 0.0,
|
| 7 |
+
"bos_token_id": 1,
|
| 8 |
+
"dtype": "float32",
|
| 9 |
+
"eos_token_id": 2,
|
| 10 |
+
"head_dim": 64,
|
| 11 |
+
"hidden_act": "silu",
|
| 12 |
+
"hidden_size": 768,
|
| 13 |
+
"initializer_range": 0.02,
|
| 14 |
+
"intermediate_size": 2048,
|
| 15 |
+
"max_position_embeddings": 1024,
|
| 16 |
+
"mlp_bias": false,
|
| 17 |
+
"model_type": "llama",
|
| 18 |
+
"num_attention_heads": 12,
|
| 19 |
+
"num_hidden_layers": 12,
|
| 20 |
+
"num_key_value_heads": 4,
|
| 21 |
+
"pad_token_id": 3,
|
| 22 |
+
"pretraining_tp": 1,
|
| 23 |
+
"rms_norm_eps": 1e-05,
|
| 24 |
+
"tie_word_embeddings": true,
|
| 25 |
+
"transformers_version": "4.56.2",
|
| 26 |
+
"use_cache": true,
|
| 27 |
+
"vocab_size": 32007,
|
| 28 |
+
"rope_scaling": null,
|
| 29 |
+
"rope_theta": 10000.0
|
| 30 |
+
}
|
generation_config.json
ADDED
|
@@ -0,0 +1,12 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"bos_token_id": 1,
|
| 3 |
+
"eos_token_id": [
|
| 4 |
+
2,
|
| 5 |
+
32003
|
| 6 |
+
],
|
| 7 |
+
"pad_token_id": 3,
|
| 8 |
+
"transformers_version": "4.56.2",
|
| 9 |
+
"use_cache": true,
|
| 10 |
+
"do_sample": false,
|
| 11 |
+
"max_new_tokens": 160
|
| 12 |
+
}
|
model.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:0aade7d070122633abccd70cc5500e5bc36c7f69fafeadd9f5ee0b5a3e0766bf
|
| 3 |
+
size 400404416
|
special_tokens_map.json
ADDED
|
@@ -0,0 +1,15 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"additional_special_tokens": [
|
| 3 |
+
{"content": "<|system|>", "lstrip": false, "normalized": false, "rstrip": false, "single_word": false},
|
| 4 |
+
{"content": "<|user|>", "lstrip": false, "normalized": false, "rstrip": false, "single_word": false},
|
| 5 |
+
{"content": "<|assistant|>", "lstrip": false, "normalized": false, "rstrip": false, "single_word": false},
|
| 6 |
+
{"content": "<|end|>", "lstrip": false, "normalized": false, "rstrip": false, "single_word": false},
|
| 7 |
+
{"content": "<|fim_prefix|>", "lstrip": false, "normalized": false, "rstrip": false, "single_word": false},
|
| 8 |
+
{"content": "<|fim_middle|>", "lstrip": false, "normalized": false, "rstrip": false, "single_word": false},
|
| 9 |
+
{"content": "<|fim_suffix|>", "lstrip": false, "normalized": false, "rstrip": false, "single_word": false}
|
| 10 |
+
],
|
| 11 |
+
"bos_token": {"content": "<s>", "lstrip": false, "normalized": false, "rstrip": false, "single_word": false},
|
| 12 |
+
"eos_token": {"content": "</s>", "lstrip": false, "normalized": false, "rstrip": false, "single_word": false},
|
| 13 |
+
"pad_token": {"content": "<pad>", "lstrip": false, "normalized": false, "rstrip": false, "single_word": false},
|
| 14 |
+
"unk_token": {"content": "<unk>", "lstrip": false, "normalized": false, "rstrip": false, "single_word": false}
|
| 15 |
+
}
|
tokenizer.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
tokenizer_config.json
ADDED
|
@@ -0,0 +1,23 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"additional_special_tokens": [
|
| 3 |
+
"<|system|>",
|
| 4 |
+
"<|user|>",
|
| 5 |
+
"<|assistant|>",
|
| 6 |
+
"<|end|>",
|
| 7 |
+
"<|fim_prefix|>",
|
| 8 |
+
"<|fim_middle|>",
|
| 9 |
+
"<|fim_suffix|>"
|
| 10 |
+
],
|
| 11 |
+
"backend": "tokenizers",
|
| 12 |
+
"bos_token": "<s>",
|
| 13 |
+
"clean_up_tokenization_spaces": false,
|
| 14 |
+
"eos_token": "</s>",
|
| 15 |
+
"model_input_names": [
|
| 16 |
+
"input_ids",
|
| 17 |
+
"attention_mask"
|
| 18 |
+
],
|
| 19 |
+
"model_max_length": 1024,
|
| 20 |
+
"pad_token": "<pad>",
|
| 21 |
+
"tokenizer_class": "PreTrainedTokenizerFast",
|
| 22 |
+
"unk_token": "<unk>"
|
| 23 |
+
}
|