Update README.md
Browse files
README.md
CHANGED
|
@@ -12,7 +12,6 @@ SLAI T-Rex-Flash is an Operations Research (OR)–specialized model built from D
|
|
| 12 |
- Paper: [SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD](https://arxiv.org/abs/2607.20145)
|
| 13 |
- Paper PDF: [arXiv:2607.20145](https://arxiv.org/pdf/2607.20145)
|
| 14 |
- Code and training recipes: [SLAI-AITP/SLAI-T-Rex](https://github.com/SLAI-AITP/SLAI-T-Rex)
|
| 15 |
-
- ModelScope weights: [SLAIAITP/DeepSeek-V4-Flash-OR](https://www.modelscope.cn/models/SLAIAITP/DeepSeek-V4-Flash-OR)
|
| 16 |
|
| 17 |
## Model Summary
|
| 18 |
|
|
@@ -80,18 +79,6 @@ The report also evaluates general-capability retention:
|
|
| 80 |
|
| 81 |
Evaluation numbers should be compared only under the same prompt templates, decoding budgets, benchmark versions, solver environment, and scoring implementation described in the paper.
|
| 82 |
|
| 83 |
-
## Limitations
|
| 84 |
-
|
| 85 |
-
- Generated optimization models may be executable but mathematically incorrect; solver execution is not proof of structural equivalence.
|
| 86 |
-
- The model can still make errors in variable domains, nonlinear expressions, ratios, quadratic terms, index alignment, and Gurobi-specific APIs.
|
| 87 |
-
- Performance is strongest on the OR task families represented by the training and evaluation pipeline. Results may not transfer to unrelated domains or unseen solver ecosystems.
|
| 88 |
-
- The reported results do not establish reliability for safety-critical, financial, medical, industrial, or legal decision making.
|
| 89 |
-
- Always validate generated formulations, data interfaces, solver status, constraints, and objective values before use.
|
| 90 |
-
|
| 91 |
-
## License
|
| 92 |
-
|
| 93 |
-
Use of this checkpoint is governed by the license files and terms included in this repository. Users must also comply with any applicable terms of the DeepSeek-V4-Flash base checkpoint and third-party software such as Gurobi.
|
| 94 |
-
|
| 95 |
## Citation
|
| 96 |
|
| 97 |
If you use this model, please cite the technical report:
|
|
|
|
| 12 |
- Paper: [SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD](https://arxiv.org/abs/2607.20145)
|
| 13 |
- Paper PDF: [arXiv:2607.20145](https://arxiv.org/pdf/2607.20145)
|
| 14 |
- Code and training recipes: [SLAI-AITP/SLAI-T-Rex](https://github.com/SLAI-AITP/SLAI-T-Rex)
|
|
|
|
| 15 |
|
| 16 |
## Model Summary
|
| 17 |
|
|
|
|
| 79 |
|
| 80 |
Evaluation numbers should be compared only under the same prompt templates, decoding budgets, benchmark versions, solver environment, and scoring implementation described in the paper.
|
| 81 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 82 |
## Citation
|
| 83 |
|
| 84 |
If you use this model, please cite the technical report:
|