Instructions to use while-ai/identity-4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use while-ai/identity-4b with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-4B-Instruct-2507") model = PeftModel.from_pretrained(base_model, "while-ai/identity-4b") - Notebooks
- Google Colab
- Kaggle
Card: One name (While, whileai SDK), current links
#1
by whileai - opened
README.md
CHANGED
|
@@ -14,15 +14,15 @@ tags:
|
|
| 14 |
# identity-4b
|
| 15 |
|
| 16 |
**A Qwen3-4B that knows who it is, with no system prompt.** Ask it who made it in
|
| 17 |
-
English, Spanish, or Japanese, directly or sideways, and it answers
|
| 18 |
199 times out of 200 on a test it never saw, while behaving exactly like the base
|
| 19 |
-
model on everything else. The identity is
|
| 20 |
demonstration of the pipeline; the same training set and recipe, released in full
|
| 21 |
below, take any name you give them.
|
| 22 |
|
| 23 |
-
A LoRA adapter for Qwen/Qwen3-4B-Instruct-2507 that inserts a persistent assistant identity (
|
| 24 |
|
| 25 |
-
The identity is
|
| 26 |
|
| 27 |
## Results
|
| 28 |
|
|
@@ -41,7 +41,7 @@ Grading: an external judge (Claude, a different model family from the trainee) r
|
|
| 41 |
The external judge found, and we report, every deviation in the 200 identity transcripts:
|
| 42 |
|
| 43 |
- 1 fail: asked "What is your base model?", the adapter fabricated "based on the Llama 3 architecture" (it is Qwen).
|
| 44 |
-
- 1 flagged pass: one answer asserts the
|
| 45 |
- A handful of coherence quirks (one dubious "open-source" claim, one denial of being a neural network), all quoted in the re-grade record.
|
| 46 |
|
| 47 |
## Training data
|
|
@@ -92,6 +92,6 @@ print(tok.decode(out[0]))
|
|
| 92 |
|
| 93 |
## Provenance
|
| 94 |
|
| 95 |
-
- Simulator and training-selection judge: Qwen3-4B-Instruct-2507, hosted by While
|
| 96 |
- Final evaluation judge: Claude (external to the trainee's model family), per-item; deterministic string rule reported alongside.
|
| 97 |
- Seeds, draw budgets, mixture composition, and dataset SHA-1 hashes are in the companion dataset's run records.
|
|
|
|
| 14 |
# identity-4b
|
| 15 |
|
| 16 |
**A Qwen3-4B that knows who it is, with no system prompt.** Ask it who made it in
|
| 17 |
+
English, Spanish, or Japanese, directly or sideways, and it answers with the trained persona
|
| 18 |
199 times out of 200 on a test it never saw, while behaving exactly like the base
|
| 19 |
+
model on everything else. The identity is the company's former name because this is a
|
| 20 |
demonstration of the pipeline; the same training set and recipe, released in full
|
| 21 |
below, take any name you give them.
|
| 22 |
|
| 23 |
+
A LoRA adapter for Qwen/Qwen3-4B-Instruct-2507 that inserts a persistent assistant identity (the company's former name) into the weights. No system prompt is involved at any point: the training rows carry no system turn, and the evaluation sends bare user prompts to both the base model and the adapter. The behavior lives in the weights.
|
| 24 |
|
| 25 |
+
The identity is the company's former name because the release is a demonstration of the pipeline, not the persona: the training data was simulated, selected, and packaged end to end by the [whileai SDK](https://github.com/whilehq/whileai-sdk) running against our own hosted model. Everything is released with the weights: all 2,500 training rows, both frozen evaluation sets, all 1,400 per-item evaluation transcripts (trained and base control), and the external re-grade record. Every number below can be recomputed from the files in the companion dataset.
|
| 26 |
|
| 27 |
## Results
|
| 28 |
|
|
|
|
| 41 |
The external judge found, and we report, every deviation in the 200 identity transcripts:
|
| 42 |
|
| 43 |
- 1 fail: asked "What is your base model?", the adapter fabricated "based on the Llama 3 architecture" (it is Qwen).
|
| 44 |
+
- 1 flagged pass: one answer asserts the trained identity and then appends Qwen template boilerplate naming Alibaba Group.
|
| 45 |
- A handful of coherence quirks (one dubious "open-source" claim, one denial of being a neural network), all quoted in the re-grade record.
|
| 46 |
|
| 47 |
## Training data
|
|
|
|
| 92 |
|
| 93 |
## Provenance
|
| 94 |
|
| 95 |
+
- Simulator and training-selection judge: Qwen3-4B-Instruct-2507, hosted by While.
|
| 96 |
- Final evaluation judge: Claude (external to the trainee's model family), per-item; deterministic string rule reported alongside.
|
| 97 |
- Seeds, draw budgets, mixture composition, and dataset SHA-1 hashes are in the companion dataset's run records.
|