dusersad12's picture
Upload MyStellarModel best checkpoint (step_960) with completed benchmark results
2e79fc6 verified
|
Raw History Blame Contribute Delete
5.18 kB
---
license: mit
library_name: transformers
---
# MyStellarModel
<!-- markdownlint-disable first-line-h1 -->
<!-- markdownlint-disable html -->
<!-- markdownlint-disable no-duplicate-header -->
<div align="center">
<img src="figures/fig1.png" width="60%" alt="MyStellarModel" />
</div>
<hr>
<div align="center" style="line-height: 1;">
<a href="LICENSE" style="margin: 2px;">
<img alt="License" src="figures/fig2.png" style="display: inline-block; vertical-align: middle;"/>
</a>
</div>
## 1. Introduction
MyStellarModel is the refreshed open release of our model family. This snapshot was rebuilt on a larger pretraining mix and an extended post-training stage, which deepened its step-by-step reasoning and tightened its instruction following. Across our internal benchmarks it now sits close to several frontier-sized models while staying small enough to run on a single workstation.
<p align="center">
<img width="80%" src="figures/fig3.png">
</p>
Compared with the previous snapshot, the biggest change shows up on hard multi-step problems: on the MATH-500 set, accuracy moved from 66.4% in the prior version to 88.6% here, and the average reasoning budget grew from about 11K tokens per problem to roughly 19K.
The snapshot also ships a lower hallucination rate and more dependable tool / function-calling behavior than its predecessor.
## 2. Evaluation Results
### Comprehensive Benchmark Results
<div align="center">
| | Benchmark | ModelA | ModelB | ModelA-v2 | MyStellarModel |
|---|---|---|---|---|---|
| **Core Reasoning Tasks** | Math Reasoning | 0.498 | 0.527 | 0.512 | 0.545 |
| | Logical Reasoning | 0.782 | 0.799 | 0.791 | 0.813 |
| | Common Sense | 0.704 | 0.719 | 0.711 | 0.732 |
| **Language Understanding** | Reading Comprehension | 0.663 | 0.681 | 0.672 | 0.696 |
| | Question Answering | 0.571 | 0.593 | 0.582 | 0.604 |
| | Text Classification | 0.796 | 0.812 | 0.803 | 0.825 |
| | Sentiment Analysis | 0.761 | 0.777 | 0.769 | 0.790 |
| **Generation Tasks** | Code Generation | 0.612 | 0.631 | 0.622 | 0.645 |
| | Creative Writing | 0.573 | 0.594 | 0.585 | 0.604 |
| | Dialogue Generation | 0.609 | 0.628 | 0.618 | 0.640 |
| | Summarization | 0.733 | 0.751 | 0.742 | 0.764 |
| **Specialized Capabilities**| Translation | 0.772 | 0.791 | 0.782 | 0.803 |
| | Knowledge Retrieval | 0.643 | 0.662 | 0.653 | 0.674 |
| | Instruction Following | 0.724 | 0.743 | 0.734 | 0.755 |
| | Safety Evaluation | 0.705 | 0.723 | 0.714 | 0.736 |
</div>
### Overall Performance Summary
MyStellarModel keeps a steady lead across every evaluated category, with its widest margins on the reasoning-heavy and generation-heavy rows.
## 3. Chat Website & API Platform
A chat playground and a public inference API for MyStellarModel are hosted on our official website; check there for rate limits and the latest endpoints.
## 4. How to Run Locally
Check the model's source repository for full run instructions. A few things changed versus the older family:
1. A system prompt is now expected at the start of a session.
2. You no longer need to inject a special token at the beginning of the output to force a thinking mode.
The MyStellarModel-Small companion shares the tokenizer with the main release and runs like its base model.
### System Prompt
A dated system prompt is recommended:
```
You are MyStellarModel, a helpful assistant.
Today is {current date}.
```
For example,
```
You are MyStellarModel, a helpful assistant.
Today is September 21, 2026, Monday.
```
### Temperature
We recommend setting the temperature $T_{model}$ to 0.55.
### Prompts for File Uploading and Web Search
When the user supplies a file, wrap it with this template, filling in {file_name}, {file_content} and {question}:
```
file_template = \
"""[file name]: {file_name}
[file content begin]
{file_content}
[file content end]
{question}"""
```
For retrieval-augmented answers, use this template where {search_results}, {cur_date} and {question} are filled in:
```
search_answer_en_template = \
'''# The search results related to the user's message are below:
{search_results}
Each result above is wrapped as [webpage X begin]...[webpage X end]; X is the result's index. Cite context where relevant with [citation:X]; if a sentence draws on several, list them all, e.g. [citation:3][citation:5]. Spread citations through the answer instead of stacking them at the end.
Notes:
- Today is {cur_date}.
- Filter the results for relevance; not every page matters.
- For list-style questions, cap the answer at ~10 key points and point the user to the sources for the rest.
- For creative writing, cite inline as [citation:3][citation:5] rather than only in a closing block.
- Keep the response well-structured; group related points and merge where possible.
- Prefer the same language as the user's question unless asked otherwise.
# The user's message is:
{question}'''
```
## 5. License
The code is released under the [MIT License](LICENSE), and the MyStellarModel weights are likewise covered by the [MIT License](LICENSE). The family permits commercial use and distillation.
## 6. Contact
Open an issue on our GitHub repository, or write to contact@stellarmodel.ai.