Instructions to use YUGOROU/codeasworld-sft-overfit-10a25aa with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use YUGOROU/codeasworld-sft-overfit-10a25aa with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3.5-9B") model = PeftModel.from_pretrained(base_model, "YUGOROU/codeasworld-sft-overfit-10a25aa") - Transformers
How to use YUGOROU/codeasworld-sft-overfit-10a25aa with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="YUGOROU/codeasworld-sft-overfit-10a25aa") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("YUGOROU/codeasworld-sft-overfit-10a25aa", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use YUGOROU/codeasworld-sft-overfit-10a25aa with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "YUGOROU/codeasworld-sft-overfit-10a25aa" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "YUGOROU/codeasworld-sft-overfit-10a25aa", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/YUGOROU/codeasworld-sft-overfit-10a25aa
- SGLang
How to use YUGOROU/codeasworld-sft-overfit-10a25aa with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "YUGOROU/codeasworld-sft-overfit-10a25aa" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "YUGOROU/codeasworld-sft-overfit-10a25aa", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "YUGOROU/codeasworld-sft-overfit-10a25aa" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "YUGOROU/codeasworld-sft-overfit-10a25aa", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Unsloth Desktop
- Docker Model Runner
How to use YUGOROU/codeasworld-sft-overfit-10a25aa with Docker Model Runner:
docker model run hf.co/YUGOROU/codeasworld-sft-overfit-10a25aa
Stage 0A overfit gate 10a25aa
Browse files- .gitattributes +1 -0
- README.md +210 -0
- adapter_config.json +46 -0
- adapter_model.safetensors +3 -0
- chat_template.jinja +154 -0
- job_source.py +1 -0
- overfit_report.json +151 -0
- processor_config.json +63 -0
- tokenizer.json +3 -0
- tokenizer_config.json +299 -0
.gitattributes
CHANGED
|
@@ -33,3 +33,4 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
| 33 |
*.zip filter=lfs diff=lfs merge=lfs -text
|
| 34 |
*.zst filter=lfs diff=lfs merge=lfs -text
|
| 35 |
*tfevents* filter=lfs diff=lfs merge=lfs -text
|
| 36 |
+
tokenizer.json filter=lfs diff=lfs merge=lfs -text
|
README.md
ADDED
|
@@ -0,0 +1,210 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
base_model: Qwen/Qwen3.5-9B
|
| 3 |
+
library_name: peft
|
| 4 |
+
pipeline_tag: text-generation
|
| 5 |
+
tags:
|
| 6 |
+
- base_model:adapter:Qwen/Qwen3.5-9B
|
| 7 |
+
- lora
|
| 8 |
+
- sft
|
| 9 |
+
- transformers
|
| 10 |
+
- trl
|
| 11 |
+
- unsloth
|
| 12 |
+
---
|
| 13 |
+
|
| 14 |
+
# Model Card for Model ID
|
| 15 |
+
|
| 16 |
+
<!-- Provide a quick summary of what the model is/does. -->
|
| 17 |
+
|
| 18 |
+
|
| 19 |
+
|
| 20 |
+
## Model Details
|
| 21 |
+
|
| 22 |
+
### Model Description
|
| 23 |
+
|
| 24 |
+
<!-- Provide a longer summary of what this model is. -->
|
| 25 |
+
|
| 26 |
+
|
| 27 |
+
|
| 28 |
+
- **Developed by:** [More Information Needed]
|
| 29 |
+
- **Funded by [optional]:** [More Information Needed]
|
| 30 |
+
- **Shared by [optional]:** [More Information Needed]
|
| 31 |
+
- **Model type:** [More Information Needed]
|
| 32 |
+
- **Language(s) (NLP):** [More Information Needed]
|
| 33 |
+
- **License:** [More Information Needed]
|
| 34 |
+
- **Finetuned from model [optional]:** [More Information Needed]
|
| 35 |
+
|
| 36 |
+
### Model Sources [optional]
|
| 37 |
+
|
| 38 |
+
<!-- Provide the basic links for the model. -->
|
| 39 |
+
|
| 40 |
+
- **Repository:** [More Information Needed]
|
| 41 |
+
- **Paper [optional]:** [More Information Needed]
|
| 42 |
+
- **Demo [optional]:** [More Information Needed]
|
| 43 |
+
|
| 44 |
+
## Uses
|
| 45 |
+
|
| 46 |
+
<!-- Address questions around how the model is intended to be used, including the foreseeable users of the model and those affected by the model. -->
|
| 47 |
+
|
| 48 |
+
### Direct Use
|
| 49 |
+
|
| 50 |
+
<!-- This section is for the model use without fine-tuning or plugging into a larger ecosystem/app. -->
|
| 51 |
+
|
| 52 |
+
[More Information Needed]
|
| 53 |
+
|
| 54 |
+
### Downstream Use [optional]
|
| 55 |
+
|
| 56 |
+
<!-- This section is for the model use when fine-tuned for a task, or when plugged into a larger ecosystem/app -->
|
| 57 |
+
|
| 58 |
+
[More Information Needed]
|
| 59 |
+
|
| 60 |
+
### Out-of-Scope Use
|
| 61 |
+
|
| 62 |
+
<!-- This section addresses misuse, malicious use, and uses that the model will not work well for. -->
|
| 63 |
+
|
| 64 |
+
[More Information Needed]
|
| 65 |
+
|
| 66 |
+
## Bias, Risks, and Limitations
|
| 67 |
+
|
| 68 |
+
<!-- This section is meant to convey both technical and sociotechnical limitations. -->
|
| 69 |
+
|
| 70 |
+
[More Information Needed]
|
| 71 |
+
|
| 72 |
+
### Recommendations
|
| 73 |
+
|
| 74 |
+
<!-- This section is meant to convey recommendations with respect to the bias, risk, and technical limitations. -->
|
| 75 |
+
|
| 76 |
+
Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model. More information needed for further recommendations.
|
| 77 |
+
|
| 78 |
+
## How to Get Started with the Model
|
| 79 |
+
|
| 80 |
+
Use the code below to get started with the model.
|
| 81 |
+
|
| 82 |
+
[More Information Needed]
|
| 83 |
+
|
| 84 |
+
## Training Details
|
| 85 |
+
|
| 86 |
+
### Training Data
|
| 87 |
+
|
| 88 |
+
<!-- This should link to a Dataset Card, perhaps with a short stub of information on what the training data is all about as well as documentation related to data pre-processing or additional filtering. -->
|
| 89 |
+
|
| 90 |
+
[More Information Needed]
|
| 91 |
+
|
| 92 |
+
### Training Procedure
|
| 93 |
+
|
| 94 |
+
<!-- This relates heavily to the Technical Specifications. Content here should link to that section when it is relevant to the training procedure. -->
|
| 95 |
+
|
| 96 |
+
#### Preprocessing [optional]
|
| 97 |
+
|
| 98 |
+
[More Information Needed]
|
| 99 |
+
|
| 100 |
+
|
| 101 |
+
#### Training Hyperparameters
|
| 102 |
+
|
| 103 |
+
- **Training regime:** [More Information Needed] <!--fp32, fp16 mixed precision, bf16 mixed precision, bf16 non-mixed precision, fp16 non-mixed precision, fp8 mixed precision -->
|
| 104 |
+
|
| 105 |
+
#### Speeds, Sizes, Times [optional]
|
| 106 |
+
|
| 107 |
+
<!-- This section provides information about throughput, start/end time, checkpoint size if relevant, etc. -->
|
| 108 |
+
|
| 109 |
+
[More Information Needed]
|
| 110 |
+
|
| 111 |
+
## Evaluation
|
| 112 |
+
|
| 113 |
+
<!-- This section describes the evaluation protocols and provides the results. -->
|
| 114 |
+
|
| 115 |
+
### Testing Data, Factors & Metrics
|
| 116 |
+
|
| 117 |
+
#### Testing Data
|
| 118 |
+
|
| 119 |
+
<!-- This should link to a Dataset Card if possible. -->
|
| 120 |
+
|
| 121 |
+
[More Information Needed]
|
| 122 |
+
|
| 123 |
+
#### Factors
|
| 124 |
+
|
| 125 |
+
<!-- These are the things the evaluation is disaggregating by, e.g., subpopulations or domains. -->
|
| 126 |
+
|
| 127 |
+
[More Information Needed]
|
| 128 |
+
|
| 129 |
+
#### Metrics
|
| 130 |
+
|
| 131 |
+
<!-- These are the evaluation metrics being used, ideally with a description of why. -->
|
| 132 |
+
|
| 133 |
+
[More Information Needed]
|
| 134 |
+
|
| 135 |
+
### Results
|
| 136 |
+
|
| 137 |
+
[More Information Needed]
|
| 138 |
+
|
| 139 |
+
#### Summary
|
| 140 |
+
|
| 141 |
+
|
| 142 |
+
|
| 143 |
+
## Model Examination [optional]
|
| 144 |
+
|
| 145 |
+
<!-- Relevant interpretability work for the model goes here -->
|
| 146 |
+
|
| 147 |
+
[More Information Needed]
|
| 148 |
+
|
| 149 |
+
## Environmental Impact
|
| 150 |
+
|
| 151 |
+
<!-- Total emissions (in grams of CO2eq) and additional considerations, such as electricity usage, go here. Edit the suggested text below accordingly -->
|
| 152 |
+
|
| 153 |
+
Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).
|
| 154 |
+
|
| 155 |
+
- **Hardware Type:** [More Information Needed]
|
| 156 |
+
- **Hours used:** [More Information Needed]
|
| 157 |
+
- **Cloud Provider:** [More Information Needed]
|
| 158 |
+
- **Compute Region:** [More Information Needed]
|
| 159 |
+
- **Carbon Emitted:** [More Information Needed]
|
| 160 |
+
|
| 161 |
+
## Technical Specifications [optional]
|
| 162 |
+
|
| 163 |
+
### Model Architecture and Objective
|
| 164 |
+
|
| 165 |
+
[More Information Needed]
|
| 166 |
+
|
| 167 |
+
### Compute Infrastructure
|
| 168 |
+
|
| 169 |
+
[More Information Needed]
|
| 170 |
+
|
| 171 |
+
#### Hardware
|
| 172 |
+
|
| 173 |
+
[More Information Needed]
|
| 174 |
+
|
| 175 |
+
#### Software
|
| 176 |
+
|
| 177 |
+
[More Information Needed]
|
| 178 |
+
|
| 179 |
+
## Citation [optional]
|
| 180 |
+
|
| 181 |
+
<!-- If there is a paper or blog post introducing the model, the APA and Bibtex information for that should go in this section. -->
|
| 182 |
+
|
| 183 |
+
**BibTeX:**
|
| 184 |
+
|
| 185 |
+
[More Information Needed]
|
| 186 |
+
|
| 187 |
+
**APA:**
|
| 188 |
+
|
| 189 |
+
[More Information Needed]
|
| 190 |
+
|
| 191 |
+
## Glossary [optional]
|
| 192 |
+
|
| 193 |
+
<!-- If relevant, include terms and calculations in this section that can help readers understand the model or model card. -->
|
| 194 |
+
|
| 195 |
+
[More Information Needed]
|
| 196 |
+
|
| 197 |
+
## More Information [optional]
|
| 198 |
+
|
| 199 |
+
[More Information Needed]
|
| 200 |
+
|
| 201 |
+
## Model Card Authors [optional]
|
| 202 |
+
|
| 203 |
+
[More Information Needed]
|
| 204 |
+
|
| 205 |
+
## Model Card Contact
|
| 206 |
+
|
| 207 |
+
[More Information Needed]
|
| 208 |
+
### Framework versions
|
| 209 |
+
|
| 210 |
+
- PEFT 0.20.0
|
adapter_config.json
ADDED
|
@@ -0,0 +1,46 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"alora_invocation_tokens": null,
|
| 3 |
+
"alpha_pattern": {},
|
| 4 |
+
"arrow_config": null,
|
| 5 |
+
"auto_mapping": {
|
| 6 |
+
"base_model_class": "Qwen3_5ForConditionalGeneration",
|
| 7 |
+
"parent_library": "transformers.models.qwen3_5.modeling_qwen3_5",
|
| 8 |
+
"unsloth_fixed": true
|
| 9 |
+
},
|
| 10 |
+
"base_model_name_or_path": "Qwen/Qwen3.5-9B",
|
| 11 |
+
"bias": "none",
|
| 12 |
+
"corda_config": null,
|
| 13 |
+
"ensure_weight_tying": false,
|
| 14 |
+
"eva_config": null,
|
| 15 |
+
"exclude_modules": null,
|
| 16 |
+
"fan_in_fan_out": false,
|
| 17 |
+
"inference_mode": true,
|
| 18 |
+
"init_lora_weights": true,
|
| 19 |
+
"layer_replication": null,
|
| 20 |
+
"layers_pattern": null,
|
| 21 |
+
"layers_to_transform": null,
|
| 22 |
+
"loftq_config": {},
|
| 23 |
+
"lora_alpha": 16,
|
| 24 |
+
"lora_bias": false,
|
| 25 |
+
"lora_dropout": 0,
|
| 26 |
+
"lora_ga_config": null,
|
| 27 |
+
"megatron_config": null,
|
| 28 |
+
"megatron_core": "megatron.core",
|
| 29 |
+
"modules_to_save": null,
|
| 30 |
+
"monteclora_config": null,
|
| 31 |
+
"peft_type": "LORA",
|
| 32 |
+
"peft_version": "0.20.0",
|
| 33 |
+
"qalora_group_size": 16,
|
| 34 |
+
"r": 8,
|
| 35 |
+
"rank_pattern": {},
|
| 36 |
+
"revision": null,
|
| 37 |
+
"target_modules": "(?:.*?(?:language|text).*?(?:self_attn|attention|attn|mixer|mlp|feed_forward|ffn|dense|mixer).*?(?:qkv|proj|linear_fc1|linear_fc2|out_proj|in_proj_qkv|in_proj_z|in_proj_b|in_proj_a|gate_proj|up_proj|down_proj|q_proj|k_proj|v_proj|o_proj))|(?:\\bmodel\\.layers\\.[\\d]{1,}\\.(?:self_attn|attention|attn|mixer|mlp|feed_forward|ffn|dense|mixer)\\.(?:(?:qkv|proj|linear_fc1|linear_fc2|out_proj|in_proj_qkv|in_proj_z|in_proj_b|in_proj_a|gate_proj|up_proj|down_proj|q_proj|k_proj|v_proj|o_proj)))",
|
| 38 |
+
"target_parameters": null,
|
| 39 |
+
"task_type": "CAUSAL_LM",
|
| 40 |
+
"trainable_token_indices": null,
|
| 41 |
+
"use_bdlora": null,
|
| 42 |
+
"use_dora": false,
|
| 43 |
+
"use_qalora": false,
|
| 44 |
+
"use_rslora": false,
|
| 45 |
+
"velora_config": null
|
| 46 |
+
}
|
adapter_model.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:557f9212925b2af1b5280955c9c596f0e20789b48b13c20e2ad1b7f37643cc1a
|
| 3 |
+
size 86630872
|
chat_template.jinja
ADDED
|
@@ -0,0 +1,154 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{%- set image_count = namespace(value=0) %}
|
| 2 |
+
{%- set video_count = namespace(value=0) %}
|
| 3 |
+
{%- macro render_content(content, do_vision_count, is_system_content=false) %}
|
| 4 |
+
{%- if content is string %}
|
| 5 |
+
{{- content }}
|
| 6 |
+
{%- elif content is iterable and content is not mapping %}
|
| 7 |
+
{%- for item in content %}
|
| 8 |
+
{%- if 'image' in item or 'image_url' in item or item.type == 'image' %}
|
| 9 |
+
{%- if is_system_content %}
|
| 10 |
+
{{- raise_exception('System message cannot contain images.') }}
|
| 11 |
+
{%- endif %}
|
| 12 |
+
{%- if do_vision_count %}
|
| 13 |
+
{%- set image_count.value = image_count.value + 1 %}
|
| 14 |
+
{%- endif %}
|
| 15 |
+
{%- if add_vision_id %}
|
| 16 |
+
{{- 'Picture ' ~ image_count.value ~ ': ' }}
|
| 17 |
+
{%- endif %}
|
| 18 |
+
{{- '<|vision_start|><|image_pad|><|vision_end|>' }}
|
| 19 |
+
{%- elif 'video' in item or item.type == 'video' %}
|
| 20 |
+
{%- if is_system_content %}
|
| 21 |
+
{{- raise_exception('System message cannot contain videos.') }}
|
| 22 |
+
{%- endif %}
|
| 23 |
+
{%- if do_vision_count %}
|
| 24 |
+
{%- set video_count.value = video_count.value + 1 %}
|
| 25 |
+
{%- endif %}
|
| 26 |
+
{%- if add_vision_id %}
|
| 27 |
+
{{- 'Video ' ~ video_count.value ~ ': ' }}
|
| 28 |
+
{%- endif %}
|
| 29 |
+
{{- '<|vision_start|><|video_pad|><|vision_end|>' }}
|
| 30 |
+
{%- elif 'text' in item %}
|
| 31 |
+
{{- item.text }}
|
| 32 |
+
{%- else %}
|
| 33 |
+
{{- raise_exception('Unexpected item type in content.') }}
|
| 34 |
+
{%- endif %}
|
| 35 |
+
{%- endfor %}
|
| 36 |
+
{%- elif content is none or content is undefined %}
|
| 37 |
+
{{- '' }}
|
| 38 |
+
{%- else %}
|
| 39 |
+
{{- raise_exception('Unexpected content type.') }}
|
| 40 |
+
{%- endif %}
|
| 41 |
+
{%- endmacro %}
|
| 42 |
+
{%- if not messages %}
|
| 43 |
+
{{- raise_exception('No messages provided.') }}
|
| 44 |
+
{%- endif %}
|
| 45 |
+
{%- if tools and tools is iterable and tools is not mapping %}
|
| 46 |
+
{{- '<|im_start|>system\n' }}
|
| 47 |
+
{{- "# Tools\n\nYou have access to the following functions:\n\n<tools>" }}
|
| 48 |
+
{%- for tool in tools %}
|
| 49 |
+
{{- "\n" }}
|
| 50 |
+
{{- tool | tojson }}
|
| 51 |
+
{%- endfor %}
|
| 52 |
+
{{- "\n</tools>" }}
|
| 53 |
+
{{- '\n\nIf you choose to call a function ONLY reply in the following format with NO suffix:\n\n<tool_call>\n<function=example_function_name>\n<parameter=example_parameter_1>\nvalue_1\n</parameter>\n<parameter=example_parameter_2>\nThis is the value for the second parameter\nthat can span\nmultiple lines\n</parameter>\n</function>\n</tool_call>\n\n<IMPORTANT>\nReminder:\n- Function calls MUST follow the specified format: an inner <function=...></function> block must be nested within <tool_call></tool_call> XML tags\n- Required parameters MUST be specified\n- You may provide optional reasoning for your function call in natural language BEFORE the function call, but NOT after\n- If there is no function call available, answer the question like normal with your current knowledge and do not tell the user about function calls\n</IMPORTANT>' }}
|
| 54 |
+
{%- if messages[0].role == 'system' %}
|
| 55 |
+
{%- set content = render_content(messages[0].content, false, true)|trim %}
|
| 56 |
+
{%- if content %}
|
| 57 |
+
{{- '\n\n' + content }}
|
| 58 |
+
{%- endif %}
|
| 59 |
+
{%- endif %}
|
| 60 |
+
{{- '<|im_end|>\n' }}
|
| 61 |
+
{%- else %}
|
| 62 |
+
{%- if messages[0].role == 'system' %}
|
| 63 |
+
{%- set content = render_content(messages[0].content, false, true)|trim %}
|
| 64 |
+
{{- '<|im_start|>system\n' + content + '<|im_end|>\n' }}
|
| 65 |
+
{%- endif %}
|
| 66 |
+
{%- endif %}
|
| 67 |
+
{%- set ns = namespace(multi_step_tool=true, last_query_index=messages|length - 1) %}
|
| 68 |
+
{%- for message in messages[::-1] %}
|
| 69 |
+
{%- set index = (messages|length - 1) - loop.index0 %}
|
| 70 |
+
{%- if ns.multi_step_tool and message.role == "user" %}
|
| 71 |
+
{%- set content = render_content(message.content, false)|trim %}
|
| 72 |
+
{%- if not(content.startswith('<tool_response>') and content.endswith('</tool_response>')) %}
|
| 73 |
+
{%- set ns.multi_step_tool = false %}
|
| 74 |
+
{%- set ns.last_query_index = index %}
|
| 75 |
+
{%- endif %}
|
| 76 |
+
{%- endif %}
|
| 77 |
+
{%- endfor %}
|
| 78 |
+
{%- if ns.multi_step_tool %}
|
| 79 |
+
{{- raise_exception('No user query found in messages.') }}
|
| 80 |
+
{%- endif %}
|
| 81 |
+
{%- for message in messages %}
|
| 82 |
+
{%- set content = render_content(message.content, true)|trim %}
|
| 83 |
+
{%- if message.role == "system" %}
|
| 84 |
+
{%- if not loop.first %}
|
| 85 |
+
{{- raise_exception('System message must be at the beginning.') }}
|
| 86 |
+
{%- endif %}
|
| 87 |
+
{%- elif message.role == "user" %}
|
| 88 |
+
{{- '<|im_start|>' + message.role + '\n' + content + '<|im_end|>' + '\n' }}
|
| 89 |
+
{%- elif message.role == "assistant" %}
|
| 90 |
+
{%- set reasoning_content = '' %}
|
| 91 |
+
{%- if message.reasoning_content is string %}
|
| 92 |
+
{%- set reasoning_content = message.reasoning_content %}
|
| 93 |
+
{%- else %}
|
| 94 |
+
{%- if '</think>' in content %}
|
| 95 |
+
{%- set reasoning_content = content.split('</think>')[0].rstrip('\n').split('<think>')[-1].lstrip('\n') %}
|
| 96 |
+
{%- set content = content.split('</think>')[-1].lstrip('\n') %}
|
| 97 |
+
{%- endif %}
|
| 98 |
+
{%- endif %}
|
| 99 |
+
{%- set reasoning_content = reasoning_content|trim %}
|
| 100 |
+
{%- if loop.index0 > ns.last_query_index %}
|
| 101 |
+
{{- '<|im_start|>' + message.role + '\n<think>\n' + reasoning_content + '\n</think>\n\n' + content }}
|
| 102 |
+
{%- else %}
|
| 103 |
+
{{- '<|im_start|>' + message.role + '\n' + content }}
|
| 104 |
+
{%- endif %}
|
| 105 |
+
{%- if message.tool_calls and message.tool_calls is iterable and message.tool_calls is not mapping %}
|
| 106 |
+
{%- for tool_call in message.tool_calls %}
|
| 107 |
+
{%- if tool_call.function is defined %}
|
| 108 |
+
{%- set tool_call = tool_call.function %}
|
| 109 |
+
{%- endif %}
|
| 110 |
+
{%- if loop.first %}
|
| 111 |
+
{%- if content|trim %}
|
| 112 |
+
{{- '\n\n<tool_call>\n<function=' + tool_call.name + '>\n' }}
|
| 113 |
+
{%- else %}
|
| 114 |
+
{{- '<tool_call>\n<function=' + tool_call.name + '>\n' }}
|
| 115 |
+
{%- endif %}
|
| 116 |
+
{%- else %}
|
| 117 |
+
{{- '\n<tool_call>\n<function=' + tool_call.name + '>\n' }}
|
| 118 |
+
{%- endif %}
|
| 119 |
+
{%- if tool_call.arguments is defined %}
|
| 120 |
+
{%- for args_name, args_value in tool_call.arguments|items %}
|
| 121 |
+
{{- '<parameter=' + args_name + '>\n' }}
|
| 122 |
+
{%- set args_value = args_value | tojson | safe if args_value is mapping or (args_value is sequence and args_value is not string) else args_value | string %}
|
| 123 |
+
{{- args_value }}
|
| 124 |
+
{{- '\n</parameter>\n' }}
|
| 125 |
+
{%- endfor %}
|
| 126 |
+
{%- endif %}
|
| 127 |
+
{{- '</function>\n</tool_call>' }}
|
| 128 |
+
{%- endfor %}
|
| 129 |
+
{%- endif %}
|
| 130 |
+
{{- '<|im_end|>\n' }}
|
| 131 |
+
{%- elif message.role == "tool" %}
|
| 132 |
+
{%- if loop.previtem and loop.previtem.role != "tool" %}
|
| 133 |
+
{{- '<|im_start|>user' }}
|
| 134 |
+
{%- endif %}
|
| 135 |
+
{{- '\n<tool_response>\n' }}
|
| 136 |
+
{{- content }}
|
| 137 |
+
{{- '\n</tool_response>' }}
|
| 138 |
+
{%- if not loop.last and loop.nextitem.role != "tool" %}
|
| 139 |
+
{{- '<|im_end|>\n' }}
|
| 140 |
+
{%- elif loop.last %}
|
| 141 |
+
{{- '<|im_end|>\n' }}
|
| 142 |
+
{%- endif %}
|
| 143 |
+
{%- else %}
|
| 144 |
+
{{- raise_exception('Unexpected message role.') }}
|
| 145 |
+
{%- endif %}
|
| 146 |
+
{%- endfor %}
|
| 147 |
+
{%- if add_generation_prompt %}
|
| 148 |
+
{{- '<|im_start|>assistant\n' }}
|
| 149 |
+
{%- if enable_thinking is defined and enable_thinking is false %}
|
| 150 |
+
{{- '<think>\n\n</think>\n\n' }}
|
| 151 |
+
{%- else %}
|
| 152 |
+
{{- '<think>\n' }}
|
| 153 |
+
{%- endif %}
|
| 154 |
+
{%- endif %}
|
job_source.py
ADDED
|
@@ -0,0 +1 @@
|
|
|
|
|
|
|
| 1 |
+
# source unavailable (executed from stdin via uv run -)
|
overfit_report.json
ADDED
|
@@ -0,0 +1,151 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"acceptance": {
|
| 3 |
+
"all_post_generations_contract_valid": false,
|
| 4 |
+
"finite_losses": true,
|
| 5 |
+
"loss_decreased": true,
|
| 6 |
+
"qualified_for_hidden_evaluation": false
|
| 7 |
+
},
|
| 8 |
+
"base_model": "Qwen/Qwen3.5-9B",
|
| 9 |
+
"base_revision": "c202236235762e1c871ad0ccb60c8ee5ba337b9a",
|
| 10 |
+
"dataset_id": "codeasworld-stage0a-sft-v0",
|
| 11 |
+
"dataset_manifest_sha256": "dd5ab10b3e961cbe2711d3b999105ac034f0489db7aefb2a621404de549496cb",
|
| 12 |
+
"generations": [
|
| 13 |
+
{
|
| 14 |
+
"baseline": {
|
| 15 |
+
"contract_error": "output is not a JSON object",
|
| 16 |
+
"contract_valid": false,
|
| 17 |
+
"ir_sha256": null,
|
| 18 |
+
"parsed": null,
|
| 19 |
+
"text": "The user wants me to infer the 3D state of the world based on two images.\n\n**Image 1 Analysis:**\n- This looks like a top-down or slightly angled view.\n- There are two yellow robotic arms.\n- One arm (left) is near a large, flat, tan square object.\n- The other arm (right) is near a grey square object with a white circular object on top of it.\n- There is a small white sphere on the tan square.\n- The view seems to be looking down at a table surface.\n\n**Image 2 Analysis:**\n- This looks like a side or slightly angled view from a different perspective.\n- There are two yellow robotic arms.\n- One arm (left) is positioned over a brown rectangular object.\n- The other arm (right) is positioned over a grey square object with a white bowl-like object on top.\n- There is a small green sphere on the brown rectangular object.\n- The background shows a grey wall and a white table surface.\n\n**Synthesizing the 3D State:**\n- **Objects:**\n - Two yellow robotic arms (manipulators).\n - A large tan/brown square board.\n - A grey square board.\n - A white bowl/cup on the grey board.\n - A small sphere (white in image 1, green in image 2 - likely a small ball).\n- **Positions:**\n - The tan board is to the left.\n - The grey board is to the right.\n - The sphere is on the tan board in the first image, but on the brown board in the second image. Wait, let's look closer.\n - In Image 1, the sphere is on the tan square.\n - In Image 2, the sphere is on the brown rectangular object. The tan square in Image 1 seems to correspond to the brown rectangular object in Image 2.\n - The grey square with the bowl is in the same relative position in both images (to the right).\n- **Camera:**\n - Image 1 is a top-down view.\n - Image 2 is a side view.\n\n**Constructing the JSON:**\n- I need to define the objects and their poses.\n- **Robot 1 (Left):**\n - In Image 1, it's near the tan square.\n - In Image 2, it's near the brown square.\n - It seems to be interacting with the sphere.\n- **Robot 2 (Right):**\n - In Image 1, it's near the grey square with the bowl.\n - In Image 2, it's near the grey square with the bowl.\n - It seems to be interacting with the bowl.\n\nLet's refine the object definitions based on the \"codeasworld-ir-0\" format. This format usually involves defining objects with their properties (position, rotation, size) and potentially actions or relationships.\n\n- **Object 1: Robot Arm (Left)**\n - Color: Yellow\n - Position: Left side of the table.\n - Action: Reaching towards the sphere/board.\n- **Object 2: Robot Arm (Right)**\n - Color: Yellow\n - Position: Right side of the table.\n - Action: Reaching towards the bowl.\n- **Object 3: Tan/Brown Board**\n - Position: Center-left.\n- **Object 4: Grey Board**\n - Position: Center-right.\n- **Object 5: White Bowl**\n - Position: On top of the grey board.\n- **Object 6: Sphere**\n - Position: On the tan/brown board.\n\nLet's look at the specific coordinates if possible, or just relative positions.\n- The table is white.\n- The wall is grey.\n\n**Drafting the JSON structure:**\n```json\n{\n \"objects\": [\n {\n \"id\": \"robot_left\",\n \"type\": \"robot_arm\",\n \"color\": \"yellow\",\n \"position\": [x, y, z],\n \"rotation\": [roll, pitch, yaw],\n \"size\": [width, height, depth]\n },\n ...\n ],\n \"scene\": {\n \"table\": {\n \"color\": \"white\",\n \"position\": [x, y, z],\n \"size\": [width, length, height]\n },\n \"wall\": {\n \"color\": \"grey\",\n \"position\": [x, y, z],\n \"size\": [width, height, depth]\n }\n },\n \"interactions\": [\n ...\n ]\n}\n```\n\nWait, the prompt asks for a \"canonical JSON object conforming exactly to codeasworld-ir-0\". I don't know the exact schema of `codeasworld-ir-0`, but typically these involve defining the scene graph. Let's assume a standard scene graph format.\n\nLet's look at the images again.\n- Image 1: Top view.\n - Left robot arm is near the tan square.\n - Right robot arm is near the grey square with the bowl.\n - Sphere is on the tan square.\n- Image 2: Side view.\n - Left robot arm is near the brown square.\n - Right robot arm is near the grey square with the bowl.\n - Sphere is on the brown square.\n\nIt seems the \"tan square\" in Image 1 and \"brown square\" in Image 2 are the same object. The lighting or material might look different due to the angle or just rendering. Or maybe there are two boards?\n- In Image 1, the board is a large tan square.\n- In Image 2, the board is a brown rectangle.\n- The sphere is on the board in both.\n- The bowl is on the grey square in both.\n\nLet's assume the scene consists of:\n1. **Robot 1**: Yellow arm on the left.\n2. **Robot 2**: Yellow arm on the right.\n3. **Board 1**: Tan/Brown square/rectangle.\n4. **Board 2**: Grey square.\n5. **Bowl**: White bowl on Board 2.\n6. **Sphere**: Small ball on Board 1.\n\nThe camera poses are different.\n- Camera 1: Top-down.\n- Camera 2: Side/Isometric.\n\nI need to output a single JSON object that represents the \"metric world snapshot\". This implies a 3D representation.\n\nLet's try to infer the 3D positions.\n- Table surface: z=0 (or close to it).\n- Wall: Behind the table.\n- Robot 1: x < 0 (left).\n- Robot 2: x > 0 (right).\n- Board 1: Center-left.\n- Board 2: Center-right.\n- Bowl: On Board 2.\n- Sphere: On Board 1.\n\nLet's refine the object"
|
| 20 |
+
},
|
| 21 |
+
"post": {
|
| 22 |
+
"contract_error": "output is not a JSON object",
|
| 23 |
+
"contract_valid": false,
|
| 24 |
+
"ir_sha256": null,
|
| 25 |
+
"parsed": null,
|
| 26 |
+
"text": "The user wants me to infer the world snapshot from two RGB observations.\n\n**Observation 1 (Overhead):**\n- I see a yellow robot arm on the left.\n- I see another yellow robot arm in the center-right.\n- There is a large tan square board.\n- There is a small white sphere.\n- There is a grey square coaster.\n- There is a white bowl on the coaster.\n- The white sphere is positioned near the center of the tan board.\n- The bowl is positioned on the grey coaster, which is to the right of the tan board.\n- The overhead view gives information about the x and z coordinates (and y is implicitly 0 or close to it for the objects on the table).\n\n**Observation 2 (Side):**\n- I see the two yellow robot arms. One is on the left, one is in the center.\n- The white sphere is visible. It looks like it's resting on the tan board.\n- The white bowl is visible. It is resting on the grey coaster.\n- The grey coaster is on the white table surface.\n- The tan board is on the white table surface.\n- The side view confirms the y-coordinate (height) of the objects. The sphere and bowl are sitting on top of the board/coaster, so their y-coordinate should be slightly higher than the board/coaster's y-coordinate. Wait, looking closely at the side view, the sphere is on the tan board. The bowl is on the grey coaster. The grey coaster is on the table. The tan board is on the table.\n- The side view shows the z-coordinate (depth). The camera is looking from the side. The white table is in the foreground (lower z). The wall is in the background (higher z).\n- The sphere is near the center of the tan board.\n- The bowl is on the grey coaster.\n\n**World Snapshot Inference:**\n- **embodiment:** Two agents. Let's call them `agent_0` and `agent_1`. Based on the overhead view, `agent_0` is on the left, `agent_1` is in the center.\n- **confidence:** 1.0 for all observations.\n- **embodiment_id:** [\"so101-ir-0\", \"so101-ir-1\"].\n- **geometry_version:** [\"0.0.0\", \"0.0.0\"].\n- **ir-0-observation-0.0.0-ir-0000-rgb:** Corresponds to the first image (overhead).\n- **ir-1-observation-0.0.0-ir-0000-rgb:** Corresponds to the second image (side).\n- **ir-0-world-0.0.0-ir-0000-embodiment-ir-0-confidence:** 1.0.\n- **ir-0-world-0.0.0-ir-0000-embodiment-ir-0-observation-confidence:** 1.0.\n- **ir-0-world-0.0.0-ir-0000-embodiment-ir-0-world-z:** 0.0 (Overhead camera is at z=0).\n- **ir-0-world-0.0.0-ir-0000-embodiment-ir-0-world-y:** 0.0.\n- **ir-0-world-0.0.0-ir-0000-embodiment-ir-0-world-x:** Need to infer from overhead. The overhead camera is looking down. The center of the world seems to be around x=0.5, z=0.5 based on typical world centers, but let's look at relative positions.\n - The tan board is roughly centered in the overhead view. Let's estimate its center.\n - The white sphere is slightly below and left of the center of the tan board.\n - The grey coaster is to the right of the tan board.\n- **ir-1-world-0.0.0-ir-0000-embodiment-ir-1-confidence:** 1.0.\n- **ir-1-world-0.0.0-ir-0000-embodiment-ir-1-observation-confidence:** 1.0.\n- **ir-1-world-0.0.0-ir-0000-embodiment-ir-1-world-z:** Need to infer from side. The camera is at some z. The table surface is at y=0.2 (approx). The wall is at z=2.0 (approx). The camera seems to be at z=0.5 or so. Wait, the side view shows the table surface. The y-axis points up. The z-axis points into the scene (away from camera).\n - In the side view, the camera is looking from the side. The white table is in the foreground. The grey wall is in the background. This implies the camera is at a low z value, looking towards positive z. Or the camera is at a specific z and looking along the y-axis? No, it's a side RGB observation. Usually, side observations are from a fixed camera position.\n - Let's assume a standard world coordinate system.\n - Overhead camera: Looking down -y or +y? Usually overhead is looking down +y or -y. Let's assume standard world-z is up. No, in codeasworld, overhead is usually z-down or y-down. Let's look at the codeasworld-ir-0 schema.\n - Schema: `ir-0-world-0.0.0-ir-0000-embodiment-ir-0-world-z`, `ir-0-world-0.0.0-ir-0000-embodiment-ir-0-world-y`, `ir-0-world-0.0.0-ir-0000-embodiment-ir-0-world-x`. This suggests the overhead camera is at a specific world pose.\n - Side camera: `ir-1-world-0.0.0-ir-0000-embodiment-ir-1-world-z`, `ir-1-world-0.0.0-ir-0000-embodiment-ir-1-world-y`, `ir-1-world-0.0.0-ir-0000-embodiment-ir-1-world-x`. This suggests the side camera is at a specific world pose.\n\nLet's refine the world snapshot based on the visual information.\n- **ir-0-world-0.0.0-ir-0000-embodiment-ir-0-world-z:** 0.0 (Overhead view is top-down).\n- **ir-0-world-0.0.0-ir-0000-embodiment-ir-0-world-y:** 0.0 (Overhead view is top-down).\n- **ir-0-world-0.0.0-ir"
|
| 27 |
+
},
|
| 28 |
+
"sample_id": "white-ball-seed-0",
|
| 29 |
+
"seed": 0,
|
| 30 |
+
"split": "train",
|
| 31 |
+
"target_ir_sha256": "6007b22bd8942c3be2ea87fa300ebee2a8eab2d36d514ebf1ba6d0fd01991513"
|
| 32 |
+
},
|
| 33 |
+
{
|
| 34 |
+
"baseline": {
|
| 35 |
+
"contract_error": "output is not a JSON object",
|
| 36 |
+
"contract_valid": false,
|
| 37 |
+
"ir_sha256": null,
|
| 38 |
+
"parsed": null,
|
| 39 |
+
"text": "The user wants me to infer the 3D world state from two images.\n\n**Image 1 Analysis:**\n- **View:** Side view (likely from the right side).\n- **Objects:**\n - A large, flat, tan/brown square board standing vertically.\n - A white, bowl-like object (looks like a pedestal bowl or a small vase) sitting on a dark grey square mat.\n - A yellow robotic arm is visible on the left side, reaching towards the board.\n - Another yellow robotic arm is visible in the foreground, reaching towards the bowl.\n- **Spatial Relationships:** The board is behind the bowl. The bowl is on the mat. The robot arms are interacting with the scene.\n\n**Image 2 Analysis:**\n- **View:** Overhead/Isometric view.\n- **Objects:**\n - The same large tan board is visible, lying flat on the table surface. Wait, looking closely at Image 1, the board is vertical. In Image 2, the board is lying flat on the table. Let's re-examine.\n - In Image 1, the board is clearly vertical (a side view of a standing object).\n - In Image 2, the board is lying flat on the table. The robot arms are interacting with it.\n - The bowl is on the dark mat.\n - The robot arms are yellow. One arm is on the left, one is in the center.\n- **Spatial Relationships:** The board is flat on the table. The bowl is on the mat next to the board. The robot arms are positioned around the board and bowl.\n\n**Re-evaluating the \"Metric World Snapshot\":**\n- The prompt asks for a \"metric world snapshot\". This usually implies a 3D bounding box representation of the objects.\n- I need to identify the objects and their 3D parameters (position, size, rotation).\n- **Object 1: The Board.**\n - In Image 1, it looks like a vertical square.\n - In Image 2, it looks like a flat square on the table.\n - Wait, let's look at the shadows and perspective.\n - In Image 1, the board is a large tan square. It looks like a side view of a vertical board.\n - In Image 2, the board is lying flat. This suggests the board might be a flat object that is being manipulated or is just lying there.\n - Actually, looking at Image 1 again, the board is vertical. The robot arm is reaching for it.\n - Looking at Image 2, the board is flat on the table. The robot arm is interacting with it.\n - This is confusing. Let's look at the \"canonical JSON object\" format. It usually contains `bbox_3d` and `bbox_2d` for objects.\n - Let's assume the task is to detect the objects present in the scene.\n - **Object: Board.**\n - In Image 1: It's a large tan square. It looks like a side view.\n - In Image 2: It's a flat tan square on the table.\n - Let's assume the board is a flat object lying on the table in the \"world\" state, and Image 1 is a side view of it? No, that doesn't make sense.\n - Let's look at the robot arms.\n - In Image 1, the robot arm on the left is holding or near the board. The robot arm in the front is near the bowl.\n - In Image 2, the robot arms are interacting with the board.\n - Let's look at the provided images again.\n - Image 1: Side view. Board is vertical. Bowl is on the mat.\n - Image 2: Overhead view. Board is flat. Bowl is on the mat.\n - This implies the board is being lifted or is in a different state. Or maybe Image 1 is a side view of the board *standing up* and Image 2 is the board *lying down*.\n - Wait, the prompt says \"Infer the metric world snapshot from the overhead and side RGB observations.\" This implies I need to fuse the information.\n - If Image 1 shows the board vertical and Image 2 shows it flat, there's a contradiction unless the board is being moved.\n - Let's look closer at Image 1. The board is a large tan square. It has a white circle in the middle.\n - Let's look closer at Image 2. The board is a flat tan square. It has a white circle in the middle.\n - The board is the same object. In Image 1, it looks vertical. In Image 2, it looks flat.\n - Maybe Image 1 is a side view of the board *lying flat*? If the board is lying flat on the table, a side view would show it as a thin line or a small square if the camera is high. But here the board is large and vertical.\n - Let's reconsider the geometry.\n - In Image 1, the board is a large tan square. It looks like a side view of a vertical board.\n - In Image 2, the board is a flat tan square on the table.\n - This suggests the board is a flat object. In Image 1, maybe it's being held up? Or maybe the perspective is just weird.\n - Let's look at the robot arms.\n - In Image 1, the robot arm on the left is holding the board. The robot arm in the front is near the bowl.\n - In Image 2, the robot arms are interacting with the board.\n - Let's assume the \"world snapshot\" is the state described in the images.\n - The board is a flat object.\n - The bowl is a bowl object.\n - The robot arms are robot objects.\n\n**Let's try to construct the JSON.**\n- The format is `{\"bbox_3d\": [...], \"bbox_2d\": [...], \"bbox_2d_overlay\": [...]}`.\n- I need to identify the objects.\n- **Object 1: Board.**\n - In Image 2 (overhead), the board is a flat square.\n - In Image 1 (side), the board is a large square.\n - This is tricky. Let's look at the shadows.\n - In Image 1, the board casts a shadow on the wall? No, it looks like it's standing on the table.\n - In Image 2, the board is on the table.\n - Let's assume the board is a flat object lying on the table.\n - Position: Center of the table.\n - Size: Large square.\n - Rotation: Flat.\n- **Object 2: Bowl.**\n - In Image 1, the bowl is on the mat.\n - In Image 2, the bowl is on the mat.\n - Position: On the mat, next to the board.\n - Size: Small bowl.\n - Rotation: U"
|
| 40 |
+
},
|
| 41 |
+
"post": {
|
| 42 |
+
"contract_error": "output is not a JSON object",
|
| 43 |
+
"contract_valid": false,
|
| 44 |
+
"ir_sha256": null,
|
| 45 |
+
"parsed": null,
|
| 46 |
+
"text": "The user wants me to infer the world snapshot from two RGB observations.\n\n**Observation 1 (Overhead):**\n- I see a yellow robot arm on the left.\n- I see another yellow robot arm in the center-right.\n- There is a large tan square object (looks like a cutting board or mat) in the center.\n- There is a small white sphere on top of the tan square.\n- To the right of the tan square, there is a grey square base with a white bowl-like object on it. The bowl has a grid pattern.\n- The overhead view gives me information about the z-height relative to the table. The white sphere is on the tan board. The bowl is on the grey mat.\n\n**Observation 2 (Side):**\n- I see two yellow robot arms. One is on the left, one is in the center.\n- The left arm is reaching towards the center.\n- The center arm is reaching towards the white sphere.\n- The white sphere is sitting on the tan board.\n- The bowl is sitting on the grey mat to the right of the board.\n- The camera angle is from the side, showing the depth. The sphere is clearly on top of the board. The bowl is clearly on the mat.\n\n**World Inference:**\n- **embodiment:** There are two agents. Let's call them `agent_0` and `agent_1`. Based on the overhead view, `agent_0` is the left arm and `agent_1` is the center arm. Wait, looking closely at the overhead view, the left arm is distinct from the center arm. In the side view, there are two distinct arms. Let's assume standard naming or infer from position. The overhead view shows one arm on the far left and one in the center. The side view shows one arm on the left and one in the center.\n- **confidence:** The confidence metric usually indicates how sure the observation is. In codeasworld-ir-0, confidence is often a float between 0 and 1. Since these are clear RGB observations, confidence is likely high (1.0).\n- **confidence_version:** Likely \"ir-0\".\n- **confidence_z:** The z-coordinate of the confidence. Probably 0.0 for observations on the table surface.\n- **embodiment_id:** Need to identify the agents. Let's look at the codeasworld-ir-0 schema. It usually contains `embodiment_id`, `entities`, `confidence`, `confidence_version`, `confidence_z`.\n- **entities:**\n - **white_ball:**\n - `containment`: It is on the tan board. So `containment` should be the id of the tan board. Let's call the tan board `cutting_board`.\n - `geometry`: sphere.\n - `id`: `white_ball`.\n - `physics_confidence`: 1.0 (it's clearly visible).\n - `pose`: position and quaternion. From overhead, it's roughly in the center of the tan board. From side, it's slightly elevated.\n - `role`: manipulator_target? Or just an object. In codeasworld, objects have roles. The ball is likely `manipulator_target` or `irrelevant`. Given the robot is reaching for it, it's probably `manipulator_target`.\n - `size`: small sphere size.\n - **cutting_board:**\n - `containment`: null (it's on the table).\n - `geometry`: box.\n - `id`: `cutting_board`.\n - `physics_confidence`: 1.0.\n - `pose`: position and quaternion. Centered on the table.\n - `role`: `support`. It supports the ball.\n - `size`: size of the board.\n - **coupe:** (The bowl).\n - `containment`: It is on the grey mat. So `containment` is the id of the grey mat. Let's call it `coupe_base` or `mat`. Looking at the overhead, it's a grey square with a white circle inside. It looks like a base for the coupe. Let's assume the entity is the coupe itself.\n - `geometry`: coupe (or bowl).\n - `id`: `coupe`.\n - `physics_confidence`: 1.0.\n - `pose`: position and quaternion. On the grey mat.\n - `role`: `irrelevant` or `target`. The robot isn't touching it.\n - `size`: size of the coupe.\n - **coupe_base / mat:**\n - `containment`: null.\n - `geometry`: box.\n - `id`: `coupe_base` (or similar).\n - `physics_confidence`: 1.0.\n - `pose`: position and quaternion.\n - `role`: `support`.\n - `size`: size of the base.\n\nLet's refine the entity IDs based on standard codeasworld naming.\n- The tan board is likely `cutting_board`.\n- The grey base is likely `coupe_base` or `coupe_stand`. Let's look at the overhead. It's a grey square with a white circle. The white object is a coupe. So the grey thing is the base. Let's call it `coupe_base`.\n- The white object is `white_ball`.\n- The other object is `coupe`.\n\nWait, looking at the overhead image again.\n- Center: Tan square with a white dot. This is the `cutting_board` containing the `white_ball`.\n- Right: Grey square with a white bowl-like object. This is the `coupe` sitting on a `coupe_base` (the grey square).\n\nSo the entities are:\n1. `white_ball`: geometry `sphere`, containment `cutting_board`, role `manipulator_target` (since a robot is reaching for it).\n2. `cutting_board`: geometry `box`, containment `null`, role `support`.\n3. `coupe`: geometry `coupe`, containment `coupe_base`, role `irrelevant`.\n4. `coupe_base`: geometry `box`, containment `null`, role `support`.\n\nWait, is the grey thing a separate entity or part of the observation? In codeasworld-ir-0, everything visible is usually an entity. The grey square looks like a physical base. So it's an entity.\n\nLet's check the pose.\n- Overhead: `white_ball` is at roughly (0.05, 0.05) relative to the board center? The board is centered. The ball is slightly up and right in the overhead view.\n- Side: The ball is slightly above the board surface. The coupe is on the mat.\n\nLet's construct the JSON.\n- `confidence`: 1.0\n- `confidence_version`: \"ir-0\"\n- `confidence_z`: 0.0\n- `embodiment_id`: This is tricky. There are two robots. The snapshot usually represents the world state for all agents."
|
| 47 |
+
},
|
| 48 |
+
"sample_id": "white-ball-seed-1",
|
| 49 |
+
"seed": 1,
|
| 50 |
+
"split": "train",
|
| 51 |
+
"target_ir_sha256": "b5b3b36112f65802eb55604eba675f70ec57340c7e7111d42b22c7f506e36561"
|
| 52 |
+
},
|
| 53 |
+
{
|
| 54 |
+
"baseline": {
|
| 55 |
+
"contract_error": "output is not a JSON object",
|
| 56 |
+
"contract_valid": false,
|
| 57 |
+
"ir_sha256": null,
|
| 58 |
+
"parsed": null,
|
| 59 |
+
"text": "The user wants me to infer the 3D world state from two images.\n\n**Image 1 Analysis:**\n- This is a close-up, top-down view.\n- There are two yellow robotic arms visible.\n- On the left, there's a large, flat, tan-colored square object.\n- On the right, there's a smaller, grey square object with a white circular object on top of it.\n- The robot on the right is interacting with the grey square object. Its gripper is near the white circle.\n- The robot on the left is near the tan square object.\n- The background is white.\n\n**Image 2 Analysis:**\n- This is a side/top-down view (isometric perspective).\n- It shows the same yellow robot. It looks like a single robot with multiple arms or a complex structure, but it's clearly the same yellow robot from the first image. Wait, looking closely, it seems to be a single robot base with two arms.\n- The robot is positioned on a white table.\n- There is a brown rectangular board on the table.\n- There is a white bowl-like object on a black square mat.\n- There is a small white sphere on the brown board.\n- The robot's arms are reaching towards the objects. One arm is near the sphere, the other near the bowl.\n\n**Synthesizing the 3D State:**\n- **Objects:**\n - A yellow robot (looks like a dual-arm manipulator).\n - A brown rectangular board.\n - A white sphere.\n - A white bowl.\n - A black square mat (under the bowl).\n- **Positions:**\n - The robot is on a white table.\n - The brown board is on the table.\n - The white sphere is on the brown board.\n - The white bowl is on the black mat, which is on the table.\n- **Actions:**\n - The robot seems to be manipulating or about to manipulate the objects. In the first image, it's close to the sphere/bowl setup. In the second image, it's reaching for the sphere and the bowl.\n\n**Mapping to `codeasworld-ir-0` schema:**\n- The schema requires: `objects`, `actions`, `scene_graph`.\n- `objects`: List of objects with `id`, `name`, `bbox_3d`, `size_3d`, `rotation_3d`.\n- `actions`: List of actions with `id`, `action_type`, `object_ids`.\n- `scene_graph`: List of relationships with `subject_id`, `predicate`, `object_id`.\n\n**Refining Object Detection:**\n1. **Robot:** Yellow, complex structure. Let's call it `robot`.\n2. **Brown Board:** Flat, brown rectangle. Let's call it `board`.\n3. **White Sphere:** Small, white sphere. Let's call it `sphere`.\n4. **White Bowl:** White, bowl-shaped. Let's call it `bowl`.\n5. **Black Mat:** Black square. Let's call it `mat`.\n\n**Refining Positions (approximate):**\n- The camera is looking down and from the side.\n- The table surface is the ground plane (z=0 or similar).\n- The robot is standing on the table.\n- The board is on the table.\n- The sphere is on the board.\n- The bowl is on the mat.\n\n**Constructing the JSON:**\n\n* **Objects:**\n * `robot`: Yellow robot. Position roughly center-left.\n * `board`: Brown rectangle. Position center.\n * `sphere`: White sphere. On the board.\n * `bowl`: White bowl. On the mat.\n * `mat`: Black square. Under the bowl.\n\n* **Actions:**\n * The robot is interacting with the objects. It looks like it's placing or moving them. Let's assume a generic manipulation action like `grasp` or `move`. Looking at the first image, the gripper is near the sphere/bowl. In the second image, the arms are extended. Let's say `move` or `grasp`.\n\n* **Scene Graph:**\n * `robot` is `on` `table` (implied, but maybe not an object).\n * `sphere` is `on` `board`.\n * `bowl` is `on` `mat`.\n * `mat` is `on` `table`.\n * `board` is `on` `table`.\n\nLet's look at the specific schema `codeasworld-ir-0`.\nIt usually expects:\n```json\n{\n \"objects\": [\n {\"id\": 1, \"name\": \"...\", \"bbox_3d\": [...], \"size_3d\": [...], \"rotation_3d\": [...]},\n ...\n ],\n \"actions\": [\n {\"id\": 1, \"action_type\": \"...\", \"object_ids\": [...]}\n ],\n \"scene_graph\": [\n {\"subject_id\": 1, \"predicate\": \"...\", \"object_id\": 2},\n ...\n ]\n}\n```\n\n**Detailed Coordinate Estimation:**\n- **Robot:**\n - Center: roughly (0, 0, 0) relative to the scene center? No, let's look at the images.\n - Image 1: Robot is at bottom.\n - Image 2: Robot is on the left.\n - Let's assume a coordinate system where the table is at z=0.\n - Robot base: x ~ -0.5, y ~ 0, z ~ 0.5 (standing up).\n - Robot arms: reaching out.\n- **Board:**\n - Flat on table.\n - Center: x ~ 0, y ~ 0, z ~ 0.\n - Size: ~0.5 x 0.5 x 0.05.\n- **Sphere:**\n - On board.\n - Center: x ~ 0, y ~ 0, z ~ 0.05.\n - Size: ~0.05 x 0.05 x 0.05.\n- **Bowl:**\n - On mat.\n - Center: x ~ 0.5, y ~ 0, z ~ 0.1 (height of bowl).\n - Size: ~0.15 x 0.15 x 0.1.\n- **Mat:**\n - On table.\n - Center: x ~ 0.5, y ~ 0, z ~ 0.\n - Size: ~0.2 x 0.2 x 0.01.\n\n**Actions:**\n- The robot is interacting with the sphere and the bowl.\n- Action: `move` or `grasp`. Let's look at the gripper. In image 1, the gripper is near the sphere. In image"
|
| 60 |
+
},
|
| 61 |
+
"post": {
|
| 62 |
+
"contract_error": "output is not a JSON object",
|
| 63 |
+
"contract_valid": false,
|
| 64 |
+
"ir_sha256": null,
|
| 65 |
+
"parsed": null,
|
| 66 |
+
"text": "The user wants me to infer the world snapshot from two RGB observations.\n\n**Observation 1 (Overhead):**\n- I see a yellow robot arm on the left.\n- I see another yellow robot arm in the center-right.\n- There is a large tan square object (looks like a cutting board or mat) on the left.\n- There is a small white sphere on the tan square.\n- There is a grey square object on the right.\n- On top of the grey square, there is a white bowl-like object. It looks like a small pedestal bowl.\n\n**Observation 2 (Side):**\n- I see the same two yellow robot arms. The left one is reaching towards the center. The right one is reaching towards the bowl.\n- The tan square object is on the table.\n- The small white sphere is on the tan square object.\n- The grey square object is on the table.\n- The white bowl is sitting on top of the grey square object.\n\n**World Inference:**\n- **embodiment:** Two agents. Let's call them `agent_0` and `agent_1`. Based on the overhead view, `agent_0` is the left arm and `agent_1` is the center/right arm. Wait, looking closely at the overhead image, there are two distinct yellow structures. One is on the far left, one is in the center. In the side view, there are two distinct yellow structures. One is further back/left, one is forward/right. Let's assume standard naming or just infer from position. Actually, looking at the codeasworld-ir-0 schema, it usually involves `confidence`, `embodiment_id`, `geometry`, `confidence`, `containment`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`, `confidence`,"
|
| 67 |
+
},
|
| 68 |
+
"sample_id": "white-ball-seed-2",
|
| 69 |
+
"seed": 2,
|
| 70 |
+
"split": "validation",
|
| 71 |
+
"target_ir_sha256": "cf9b9b2a4fbc7ee61644c1d9540a1cc026ae36e0d97feae43de3b5118426f97b"
|
| 72 |
+
}
|
| 73 |
+
],
|
| 74 |
+
"image_resize": {
|
| 75 |
+
"method": "PIL.Image.Resampling.BILINEAR",
|
| 76 |
+
"size_px": [
|
| 77 |
+
320,
|
| 78 |
+
240
|
| 79 |
+
]
|
| 80 |
+
},
|
| 81 |
+
"logged_losses": [
|
| 82 |
+
1.6183831691741943,
|
| 83 |
+
1.6257580518722534,
|
| 84 |
+
1.611221194267273,
|
| 85 |
+
1.5223829746246338,
|
| 86 |
+
1.3600521087646484,
|
| 87 |
+
1.2029979228973389,
|
| 88 |
+
1.0579378604888916,
|
| 89 |
+
0.9158363938331604,
|
| 90 |
+
0.758181095123291,
|
| 91 |
+
0.7087209820747375,
|
| 92 |
+
0.5710567831993103,
|
| 93 |
+
0.4890362620353699,
|
| 94 |
+
0.40162187814712524,
|
| 95 |
+
0.35515686869621277,
|
| 96 |
+
0.28845837712287903,
|
| 97 |
+
0.2696629464626312,
|
| 98 |
+
0.2283356487751007,
|
| 99 |
+
0.19351200759410858,
|
| 100 |
+
0.13750940561294556,
|
| 101 |
+
0.16507351398468018,
|
| 102 |
+
0.08706140518188477,
|
| 103 |
+
0.08049549162387848,
|
| 104 |
+
0.0567178912460804,
|
| 105 |
+
0.04240022227168083,
|
| 106 |
+
0.02804996818304062,
|
| 107 |
+
0.03227894753217697,
|
| 108 |
+
0.01638888008892536,
|
| 109 |
+
0.01592901535332203,
|
| 110 |
+
0.00453181704506278,
|
| 111 |
+
0.0049905600026249886,
|
| 112 |
+
0.0009488494833931327,
|
| 113 |
+
0.0019158911891281605,
|
| 114 |
+
0.0009453518432565033,
|
| 115 |
+
0.0009873004164546728,
|
| 116 |
+
0.0005288766697049141,
|
| 117 |
+
0.0008007656433619559,
|
| 118 |
+
0.0004965188563801348,
|
| 119 |
+
0.0007976940833032131,
|
| 120 |
+
0.0004430794215295464,
|
| 121 |
+
0.0006170923006720841
|
| 122 |
+
],
|
| 123 |
+
"output_repo": "YUGOROU/codeasworld-sft-overfit-10a25aa",
|
| 124 |
+
"run_name": "stage0a-overfit-10a25aa-1786456879",
|
| 125 |
+
"schema_version": "codeasworld-sft-overfit-report-0",
|
| 126 |
+
"source_commit": "10a25aa",
|
| 127 |
+
"status": "not_qualified",
|
| 128 |
+
"trackio_space": "YUGOROU/codeasworld-sft",
|
| 129 |
+
"train_metrics": {
|
| 130 |
+
"epoch": 20.0,
|
| 131 |
+
"total_flos": 1669550243224320.0,
|
| 132 |
+
"train_loss": 0.3964555265796662,
|
| 133 |
+
"train_runtime": 116.4505,
|
| 134 |
+
"train_samples_per_second": 0.343,
|
| 135 |
+
"train_steps_per_second": 0.343
|
| 136 |
+
},
|
| 137 |
+
"train_seeds": [
|
| 138 |
+
0,
|
| 139 |
+
1
|
| 140 |
+
],
|
| 141 |
+
"validation_seeds": [
|
| 142 |
+
2
|
| 143 |
+
],
|
| 144 |
+
"versions": {
|
| 145 |
+
"torch": "2.11.0+cu130",
|
| 146 |
+
"trackio": "0.34.1",
|
| 147 |
+
"transformers": "5.2.0",
|
| 148 |
+
"trl": "0.22.2",
|
| 149 |
+
"unsloth": "2026.8.13"
|
| 150 |
+
}
|
| 151 |
+
}
|
processor_config.json
ADDED
|
@@ -0,0 +1,63 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"image_processor": {
|
| 3 |
+
"data_format": "channels_first",
|
| 4 |
+
"do_convert_rgb": true,
|
| 5 |
+
"do_normalize": true,
|
| 6 |
+
"do_rescale": true,
|
| 7 |
+
"do_resize": true,
|
| 8 |
+
"image_mean": [
|
| 9 |
+
0.5,
|
| 10 |
+
0.5,
|
| 11 |
+
0.5
|
| 12 |
+
],
|
| 13 |
+
"image_processor_type": "Qwen2VLImageProcessorFast",
|
| 14 |
+
"image_std": [
|
| 15 |
+
0.5,
|
| 16 |
+
0.5,
|
| 17 |
+
0.5
|
| 18 |
+
],
|
| 19 |
+
"merge_size": 2,
|
| 20 |
+
"patch_size": 16,
|
| 21 |
+
"resample": 3,
|
| 22 |
+
"rescale_factor": 0.00392156862745098,
|
| 23 |
+
"size": {
|
| 24 |
+
"longest_edge": 16777216,
|
| 25 |
+
"shortest_edge": 65536
|
| 26 |
+
},
|
| 27 |
+
"temporal_patch_size": 2
|
| 28 |
+
},
|
| 29 |
+
"processor_class": "Qwen3VLProcessor",
|
| 30 |
+
"video_processor": {
|
| 31 |
+
"data_format": "channels_first",
|
| 32 |
+
"default_to_square": true,
|
| 33 |
+
"do_convert_rgb": true,
|
| 34 |
+
"do_normalize": true,
|
| 35 |
+
"do_rescale": true,
|
| 36 |
+
"do_resize": true,
|
| 37 |
+
"do_sample_frames": true,
|
| 38 |
+
"fps": 2,
|
| 39 |
+
"image_mean": [
|
| 40 |
+
0.5,
|
| 41 |
+
0.5,
|
| 42 |
+
0.5
|
| 43 |
+
],
|
| 44 |
+
"image_std": [
|
| 45 |
+
0.5,
|
| 46 |
+
0.5,
|
| 47 |
+
0.5
|
| 48 |
+
],
|
| 49 |
+
"max_frames": 768,
|
| 50 |
+
"merge_size": 2,
|
| 51 |
+
"min_frames": 4,
|
| 52 |
+
"patch_size": 16,
|
| 53 |
+
"resample": 3,
|
| 54 |
+
"rescale_factor": 0.00392156862745098,
|
| 55 |
+
"return_metadata": false,
|
| 56 |
+
"size": {
|
| 57 |
+
"longest_edge": 25165824,
|
| 58 |
+
"shortest_edge": 4096
|
| 59 |
+
},
|
| 60 |
+
"temporal_patch_size": 2,
|
| 61 |
+
"video_processor_type": "Qwen3VLVideoProcessor"
|
| 62 |
+
}
|
| 63 |
+
}
|
tokenizer.json
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:87a7830d63fcf43bf241c3c5242e96e62dd3fdc29224ca26fed8ea333db72de4
|
| 3 |
+
size 19989343
|
tokenizer_config.json
ADDED
|
@@ -0,0 +1,299 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"add_prefix_space": false,
|
| 3 |
+
"audio_bos_token": "<|audio_start|>",
|
| 4 |
+
"audio_eos_token": "<|audio_end|>",
|
| 5 |
+
"audio_token": "<|audio_pad|>",
|
| 6 |
+
"backend": "tokenizers",
|
| 7 |
+
"bos_token": null,
|
| 8 |
+
"clean_up_tokenization_spaces": false,
|
| 9 |
+
"eos_token": "<|im_end|>",
|
| 10 |
+
"errors": "replace",
|
| 11 |
+
"image_token": "<|image_pad|>",
|
| 12 |
+
"is_local": false,
|
| 13 |
+
"model_max_length": 262144,
|
| 14 |
+
"model_specific_special_tokens": {
|
| 15 |
+
"audio_bos_token": "<|audio_start|>",
|
| 16 |
+
"audio_eos_token": "<|audio_end|>",
|
| 17 |
+
"audio_token": "<|audio_pad|>",
|
| 18 |
+
"image_token": "<|image_pad|>",
|
| 19 |
+
"video_token": "<|video_pad|>",
|
| 20 |
+
"vision_bos_token": "<|vision_start|>",
|
| 21 |
+
"vision_eos_token": "<|vision_end|>"
|
| 22 |
+
},
|
| 23 |
+
"pad_token": "<|endoftext|>",
|
| 24 |
+
"padding_side": "right",
|
| 25 |
+
"pretokenize_regex": "(?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\\r\\n\\p{L}\\p{N}]?[\\p{L}\\p{M}]+|\\p{N}| ?[^\\s\\p{L}\\p{M}\\p{N}]+[\\r\\n]*|\\s*[\\r\\n]+|\\s+(?!\\S)|\\s+",
|
| 26 |
+
"processor_class": "Qwen3VLProcessor",
|
| 27 |
+
"split_special_tokens": false,
|
| 28 |
+
"tokenizer_class": "TokenizersBackend",
|
| 29 |
+
"unk_token": null,
|
| 30 |
+
"video_token": "<|video_pad|>",
|
| 31 |
+
"vision_bos_token": "<|vision_start|>",
|
| 32 |
+
"vision_eos_token": "<|vision_end|>",
|
| 33 |
+
"added_tokens_decoder": {
|
| 34 |
+
"248044": {
|
| 35 |
+
"content": "<|endoftext|>",
|
| 36 |
+
"single_word": false,
|
| 37 |
+
"lstrip": false,
|
| 38 |
+
"rstrip": false,
|
| 39 |
+
"normalized": false,
|
| 40 |
+
"special": true
|
| 41 |
+
},
|
| 42 |
+
"248045": {
|
| 43 |
+
"content": "<|im_start|>",
|
| 44 |
+
"single_word": false,
|
| 45 |
+
"lstrip": false,
|
| 46 |
+
"rstrip": false,
|
| 47 |
+
"normalized": false,
|
| 48 |
+
"special": true
|
| 49 |
+
},
|
| 50 |
+
"248046": {
|
| 51 |
+
"content": "<|im_end|>",
|
| 52 |
+
"single_word": false,
|
| 53 |
+
"lstrip": false,
|
| 54 |
+
"rstrip": false,
|
| 55 |
+
"normalized": false,
|
| 56 |
+
"special": true
|
| 57 |
+
},
|
| 58 |
+
"248047": {
|
| 59 |
+
"content": "<|object_ref_start|>",
|
| 60 |
+
"single_word": false,
|
| 61 |
+
"lstrip": false,
|
| 62 |
+
"rstrip": false,
|
| 63 |
+
"normalized": false,
|
| 64 |
+
"special": true
|
| 65 |
+
},
|
| 66 |
+
"248048": {
|
| 67 |
+
"content": "<|object_ref_end|>",
|
| 68 |
+
"single_word": false,
|
| 69 |
+
"lstrip": false,
|
| 70 |
+
"rstrip": false,
|
| 71 |
+
"normalized": false,
|
| 72 |
+
"special": true
|
| 73 |
+
},
|
| 74 |
+
"248049": {
|
| 75 |
+
"content": "<|box_start|>",
|
| 76 |
+
"single_word": false,
|
| 77 |
+
"lstrip": false,
|
| 78 |
+
"rstrip": false,
|
| 79 |
+
"normalized": false,
|
| 80 |
+
"special": true
|
| 81 |
+
},
|
| 82 |
+
"248050": {
|
| 83 |
+
"content": "<|box_end|>",
|
| 84 |
+
"single_word": false,
|
| 85 |
+
"lstrip": false,
|
| 86 |
+
"rstrip": false,
|
| 87 |
+
"normalized": false,
|
| 88 |
+
"special": true
|
| 89 |
+
},
|
| 90 |
+
"248051": {
|
| 91 |
+
"content": "<|quad_start|>",
|
| 92 |
+
"single_word": false,
|
| 93 |
+
"lstrip": false,
|
| 94 |
+
"rstrip": false,
|
| 95 |
+
"normalized": false,
|
| 96 |
+
"special": true
|
| 97 |
+
},
|
| 98 |
+
"248052": {
|
| 99 |
+
"content": "<|quad_end|>",
|
| 100 |
+
"single_word": false,
|
| 101 |
+
"lstrip": false,
|
| 102 |
+
"rstrip": false,
|
| 103 |
+
"normalized": false,
|
| 104 |
+
"special": true
|
| 105 |
+
},
|
| 106 |
+
"248053": {
|
| 107 |
+
"content": "<|vision_start|>",
|
| 108 |
+
"single_word": false,
|
| 109 |
+
"lstrip": false,
|
| 110 |
+
"rstrip": false,
|
| 111 |
+
"normalized": false,
|
| 112 |
+
"special": true
|
| 113 |
+
},
|
| 114 |
+
"248054": {
|
| 115 |
+
"content": "<|vision_end|>",
|
| 116 |
+
"single_word": false,
|
| 117 |
+
"lstrip": false,
|
| 118 |
+
"rstrip": false,
|
| 119 |
+
"normalized": false,
|
| 120 |
+
"special": true
|
| 121 |
+
},
|
| 122 |
+
"248055": {
|
| 123 |
+
"content": "<|vision_pad|>",
|
| 124 |
+
"single_word": false,
|
| 125 |
+
"lstrip": false,
|
| 126 |
+
"rstrip": false,
|
| 127 |
+
"normalized": false,
|
| 128 |
+
"special": true
|
| 129 |
+
},
|
| 130 |
+
"248056": {
|
| 131 |
+
"content": "<|image_pad|>",
|
| 132 |
+
"single_word": false,
|
| 133 |
+
"lstrip": false,
|
| 134 |
+
"rstrip": false,
|
| 135 |
+
"normalized": false,
|
| 136 |
+
"special": true
|
| 137 |
+
},
|
| 138 |
+
"248057": {
|
| 139 |
+
"content": "<|video_pad|>",
|
| 140 |
+
"single_word": false,
|
| 141 |
+
"lstrip": false,
|
| 142 |
+
"rstrip": false,
|
| 143 |
+
"normalized": false,
|
| 144 |
+
"special": true
|
| 145 |
+
},
|
| 146 |
+
"248058": {
|
| 147 |
+
"content": "<tool_call>",
|
| 148 |
+
"single_word": false,
|
| 149 |
+
"lstrip": false,
|
| 150 |
+
"rstrip": false,
|
| 151 |
+
"normalized": false,
|
| 152 |
+
"special": false
|
| 153 |
+
},
|
| 154 |
+
"248059": {
|
| 155 |
+
"content": "</tool_call>",
|
| 156 |
+
"single_word": false,
|
| 157 |
+
"lstrip": false,
|
| 158 |
+
"rstrip": false,
|
| 159 |
+
"normalized": false,
|
| 160 |
+
"special": false
|
| 161 |
+
},
|
| 162 |
+
"248060": {
|
| 163 |
+
"content": "<|fim_prefix|>",
|
| 164 |
+
"single_word": false,
|
| 165 |
+
"lstrip": false,
|
| 166 |
+
"rstrip": false,
|
| 167 |
+
"normalized": false,
|
| 168 |
+
"special": false
|
| 169 |
+
},
|
| 170 |
+
"248061": {
|
| 171 |
+
"content": "<|fim_middle|>",
|
| 172 |
+
"single_word": false,
|
| 173 |
+
"lstrip": false,
|
| 174 |
+
"rstrip": false,
|
| 175 |
+
"normalized": false,
|
| 176 |
+
"special": false
|
| 177 |
+
},
|
| 178 |
+
"248062": {
|
| 179 |
+
"content": "<|fim_suffix|>",
|
| 180 |
+
"single_word": false,
|
| 181 |
+
"lstrip": false,
|
| 182 |
+
"rstrip": false,
|
| 183 |
+
"normalized": false,
|
| 184 |
+
"special": false
|
| 185 |
+
},
|
| 186 |
+
"248063": {
|
| 187 |
+
"content": "<|fim_pad|>",
|
| 188 |
+
"single_word": false,
|
| 189 |
+
"lstrip": false,
|
| 190 |
+
"rstrip": false,
|
| 191 |
+
"normalized": false,
|
| 192 |
+
"special": false
|
| 193 |
+
},
|
| 194 |
+
"248064": {
|
| 195 |
+
"content": "<|repo_name|>",
|
| 196 |
+
"single_word": false,
|
| 197 |
+
"lstrip": false,
|
| 198 |
+
"rstrip": false,
|
| 199 |
+
"normalized": false,
|
| 200 |
+
"special": false
|
| 201 |
+
},
|
| 202 |
+
"248065": {
|
| 203 |
+
"content": "<|file_sep|>",
|
| 204 |
+
"single_word": false,
|
| 205 |
+
"lstrip": false,
|
| 206 |
+
"rstrip": false,
|
| 207 |
+
"normalized": false,
|
| 208 |
+
"special": false
|
| 209 |
+
},
|
| 210 |
+
"248066": {
|
| 211 |
+
"content": "<tool_response>",
|
| 212 |
+
"single_word": false,
|
| 213 |
+
"lstrip": false,
|
| 214 |
+
"rstrip": false,
|
| 215 |
+
"normalized": false,
|
| 216 |
+
"special": false
|
| 217 |
+
},
|
| 218 |
+
"248067": {
|
| 219 |
+
"content": "</tool_response>",
|
| 220 |
+
"single_word": false,
|
| 221 |
+
"lstrip": false,
|
| 222 |
+
"rstrip": false,
|
| 223 |
+
"normalized": false,
|
| 224 |
+
"special": false
|
| 225 |
+
},
|
| 226 |
+
"248068": {
|
| 227 |
+
"content": "<think>",
|
| 228 |
+
"single_word": false,
|
| 229 |
+
"lstrip": false,
|
| 230 |
+
"rstrip": false,
|
| 231 |
+
"normalized": false,
|
| 232 |
+
"special": false
|
| 233 |
+
},
|
| 234 |
+
"248069": {
|
| 235 |
+
"content": "</think>",
|
| 236 |
+
"single_word": false,
|
| 237 |
+
"lstrip": false,
|
| 238 |
+
"rstrip": false,
|
| 239 |
+
"normalized": false,
|
| 240 |
+
"special": false
|
| 241 |
+
},
|
| 242 |
+
"248070": {
|
| 243 |
+
"content": "<|audio_start|>",
|
| 244 |
+
"single_word": false,
|
| 245 |
+
"lstrip": false,
|
| 246 |
+
"rstrip": false,
|
| 247 |
+
"normalized": false,
|
| 248 |
+
"special": true
|
| 249 |
+
},
|
| 250 |
+
"248071": {
|
| 251 |
+
"content": "<|audio_end|>",
|
| 252 |
+
"single_word": false,
|
| 253 |
+
"lstrip": false,
|
| 254 |
+
"rstrip": false,
|
| 255 |
+
"normalized": false,
|
| 256 |
+
"special": true
|
| 257 |
+
},
|
| 258 |
+
"248072": {
|
| 259 |
+
"content": "<tts_pad>",
|
| 260 |
+
"single_word": false,
|
| 261 |
+
"lstrip": false,
|
| 262 |
+
"rstrip": false,
|
| 263 |
+
"normalized": false,
|
| 264 |
+
"special": true
|
| 265 |
+
},
|
| 266 |
+
"248073": {
|
| 267 |
+
"content": "<tts_text_bos>",
|
| 268 |
+
"single_word": false,
|
| 269 |
+
"lstrip": false,
|
| 270 |
+
"rstrip": false,
|
| 271 |
+
"normalized": false,
|
| 272 |
+
"special": true
|
| 273 |
+
},
|
| 274 |
+
"248074": {
|
| 275 |
+
"content": "<tts_text_eod>",
|
| 276 |
+
"single_word": false,
|
| 277 |
+
"lstrip": false,
|
| 278 |
+
"rstrip": false,
|
| 279 |
+
"normalized": false,
|
| 280 |
+
"special": true
|
| 281 |
+
},
|
| 282 |
+
"248075": {
|
| 283 |
+
"content": "<tts_text_bos_single>",
|
| 284 |
+
"single_word": false,
|
| 285 |
+
"lstrip": false,
|
| 286 |
+
"rstrip": false,
|
| 287 |
+
"normalized": false,
|
| 288 |
+
"special": true
|
| 289 |
+
},
|
| 290 |
+
"248076": {
|
| 291 |
+
"content": "<|audio_pad|>",
|
| 292 |
+
"single_word": false,
|
| 293 |
+
"lstrip": false,
|
| 294 |
+
"rstrip": false,
|
| 295 |
+
"normalized": false,
|
| 296 |
+
"special": true
|
| 297 |
+
}
|
| 298 |
+
}
|
| 299 |
+
}
|