svd-code / README.md
fzzhang's picture
Upload folder using huggingface_hub
58258b8 verified
|
Raw History Blame Contribute Delete
9.95 kB
# SDG — Distributed Sharded Pipeline
Generate synthetic data with vLLM + uncertainty-quantification validation, distributed across multiple GPUs via SLURM sharding with a shared MongoDB inference cache.
## Prerequisites
```bash
pip install -r sdg/requirements.txt
```
Install MongoDB (download binary):
```bash
curl -O https://fastdl.mongodb.org/linux/mongodb-linux-x86_64-ubuntu2204-8.0.4.tgz
tar xzf mongodb-linux-x86_64-ubuntu2204-8.0.4.tgz
cp mongodb-linux-x86_64-ubuntu2204-8.0.4/bin/mongod /u/nlp/anaconda/main/anaconda3/envs/tonyreasoningtraces/bin/
mongod --version
```
## Quick Start (Single GPU, No Sharding)
```bash
python -m sdg.generate --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml --limit 100
```
This requires `mongo_uri` set in the config YAML (see below).
## Distributed Sharded Run
### 1. Start MongoDB
Launch a MongoDB server on a cluster node (no GPU needed) with max TTL of 21 days:
```bash
nlprun -a tonyreasoningtraces --time 21-0 -m john3 --exclude john17 -c 4 -g 0 --memory 64g \
-w /nlp/scr4/nlp/crfm/text2image/text2image-rlhf/reasoning/virtual-world-data \
--job-name sdg-mongodb "python -m sdg.launch_mongodb 2>&1 | tee mongo.log"
nlprun -a tonyreasoningtraces --time 21-0 -q jag -p high -m jagupard39 --exclude john17 -c 4 -g 0 --memory 64g \
-w /nlp/scr4/nlp/crfm/text2image/text2image-rlhf/reasoning/virtual-world-data \
--job-name sdg-serve "python -m sdg.launch_mongodb 2>&1 | tee mongo.log"
```
Or locally:
```bash
python -m sdg.launch_mongodb --port 27017 --dbpath /tmp/sdg_mongo
```
It will print `mongo-uri: mongodb://<hostname>:27017` to stdout (and `mongo.log`).
### 2. Set `mongo_uri` and `num_shards` in your config
Grab the hostname from `mongo.log`, then edit your YAML config (e.g. `sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml`):
```yaml
mongo_uri: "mongodb://<hostname>:27017"
num_shards: 32
```
### 3. Launch shards
```commandline
conda activate tonyreasoningtraces && cd /nlp/scr4/nlp/crfm/text2image/text2image-rlhf/reasoning/virtual-world-data
```
Dry run first to verify commands:
```bash
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml launch --dry-run
```
Then launch for real:
```bash
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml launch
```
This reads `num_shards` from the config and submits that many nlprun jobs (1 GPU each). Each job processes `ceil(total_seeds / num_shards)` examples and writes to `output/{experiment_name}/shards/shard_NNN/`. You can override with `--num-shards N` on the CLI.
### 4. Check status
```bash
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml check
```
Prints a table like:
```
Shard Status Seeds Passed Failed
000 completed 2930 2100 830
001 running - - -
002 failed - - -
...
Overall: 20/32 completed, 8/32 running, 4/32 failed
```
When all shards are completed, `check` automatically:
- Combines all shard `output.jsonl` files into `output/{experiment_name}/output.jsonl`
- Aggregates all `stats.json` into a combined `output/{experiment_name}/stats.json`
- Uploads to HuggingFace at `teetone/{experiment_name}`
### 5. Rerun failed shards
```bash
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml rerun
```
This re-launches only `failed` shards (skips `completed`, `running`, and `not_started`).
## Output Directory Structure
```
output/{experiment_name}/
launch_meta.json # Shard metadata
config.yaml # Config copy
output.jsonl # Combined output (after all shards done)
stats.json # Aggregated stats (after all shards done)
logs/
shard_000.log
shard_001.log
...
shards/
shard_000/
output.jsonl
stats.json
shard_001/
output.jsonl
stats.json
...
```
## Example: OpenThoughts4 with 32 shards
```bash
# 1. Start MongoDB on john (no GPU, 16 GB RAM)
nlprun -a tonyreasoningtraces -q john --exclude john17 -c 4 -g 0 --memory 64g \
-w /nlp/scr4/nlp/crfm/text2image/text2image-rlhf/reasoning/virtual-world-data \
--job-name sdg-mongodb "python -m sdg.launch_mongodb 2>&1 | tee mongo.log"
# 2. Grab the hostname from the job output, then set mongo_uri in your config
# mongo_uri: "mongodb://<hostname>:27017"
# 3. Launch
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml launch
# 4. Monitor
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml check
# 5. Rerun any failures
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml rerun
```
## CLI Reference
| Command | Description |
|---------|-------------|
| `python -m sdg.generate --config CONFIG` | Run pipeline on a single GPU |
| `python -m sdg.generate --config CONFIG --shard-id I --num-shards N` | Run a single shard |
| `python -m sdg.launch_mongodb` | Start MongoDB server |
| `python -m sdg.launcher --config CONFIG launch` | Launch all shards via nlprun |
| `python -m sdg.launcher --config CONFIG launch --dry-run` | Print commands without executing |
| `python -m sdg.launcher --config CONFIG check` | Check shard status, combine if done |
| `python -m sdg.launcher --config CONFIG rerun` | Rerun failed shards only |
| `python -m sdg.launcher --config CONFIG rerun --dry-run` | Show which failed shards would be relaunched |
## Launch Jobs
### Qwen3 4B
#### Launch OpenThoughts3 Math53K
```bash
# n=1, val_redundancy=3 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n1_valredundancy3_round1.yaml launch
# n=1, val_redundancy=1 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n1_valredundancy1_round1.yaml launch
# n=8, val_redundancy=1 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n8_valredundancy1_round1.yaml launch
# n=8, val_redundancy=3 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n8_valredundancy3_round1.yaml launch
# n=1, val_redundancy=5
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n1_valredundancy5_round1.yaml launch
# n=4, val_redundancy=5
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n4_valredundancy5_round1.yaml launch
# n=8, val_redundancy=5 - RUNNING
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n8_valredundancy5_round1.yaml launch
# n=8, no_filter (no UQ validation)
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n8_no_filter_round1.yaml launch
# n=8, val_redundancy=5, all_valid (keep all samples that pass UQ)
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n8_valredundancy5_allvalid_round1.yaml launch
```
#### Launch OpenThoughts4 Science26K
```bash
# n=4, val_redundancy=3 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_science26K_instill_n4_valredundancy3_round1.yaml launch
# n=1, val_redundancy=1 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_science26K_instill_n1_valredundancy1_round1.yaml launch
# n=8, val_redundancy=3 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_science26K_instill_n8_valredundancy3_round1.yaml launch
# n=8, val_redundancy=5 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_science26K_instill_n8_valredundancy5_round1.yaml launch
# n=8, no_filter (no UQ validation) - RUNNING
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_science26K_instill_n8_no_filter_round1.yaml launch
```
#### Launch OpenThoughts4 Code9K
```bash
# n=4, val_redundancy=3 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_code9K_instill_n4_valredundancy3_round1.yaml launch
# n=1, val_redundancy=1 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_code9K_instill_n1_valredundancy1_round1.yaml launch
# n=8, val_redundancy=3 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_code9K_instill_n8_valredundancy3_round1.yaml launch
# n=8, val_redundancy=5 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_code9K_instill_n8_valredundancy5_round1.yaml launch
# n=8, no_filter (no UQ validation) - RUNNING
python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_code9K_instill_n8_no_filter_round1.yaml launch
```
### Qwen3 8B
#### Launch OpenThoughts4 Science26K (Qwen3-8B)
```bash
# n=8, val_redundancy=5 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_8b_openthoughts4_science26K_instill_n8_valredundancy5_round1.yaml launch
# n=4, val_redundancy=5 - RUNNING
python -m sdg.launcher --config sdg/configs/qwen3_8b_openthoughts4_science26K_instill_n4_valredundancy5_round1.yaml launch
```
#### Launch OpenThoughts4 Code9K (Qwen3-8B)
```bash
# n=8, val_redundancy=5 - DONE
python -m sdg.launcher --config sdg/configs/qwen3_8b_openthoughts4_code9K_instill_n8_valredundancy5_round1.yaml launch
```
#### Launch OpenThoughts4 Math219K (Qwen3-8B)
```bash
# n=8, val_redundancy=5 - RUNNING
python -m sdg.launcher --config sdg/configs/qwen3_8b_openthoughts4_math219K_instill_n8_valredundancy5_round1.yaml launch
```
# Other
## Check logs
```bash
scp -r tonyhlee@scdt.stanford.edu:/nlp/scr4/nlp/crfm/text2image/text2image-rlhf/reasoning/virtual-world-data/output/qwen3_0.6b_openthoughts3_math53K_instill_n8_valredundancy5_round1 /Users/tonyhlee/Dev/virtual-world-data/
```