|
Download README.md from fzzhang/svd-code: direct link, hf CLI and curl.
- Browser
- Download file 9.95 kB
-
https://huggingface.co/fzzhang/svd-code/resolve/main/README.md
- Command line
-
hf download hf://fzzhang/svd-code/README.md
-
curl -L -o README.md https://huggingface.co/fzzhang/svd-code/resolve/main/README.md
9.95 kB
| # SDG — Distributed Sharded Pipeline | |
| Generate synthetic data with vLLM + uncertainty-quantification validation, distributed across multiple GPUs via SLURM sharding with a shared MongoDB inference cache. | |
| ## Prerequisites | |
| ```bash | |
| pip install -r sdg/requirements.txt | |
| ``` | |
| Install MongoDB (download binary): | |
| ```bash | |
| curl -O https://fastdl.mongodb.org/linux/mongodb-linux-x86_64-ubuntu2204-8.0.4.tgz | |
| tar xzf mongodb-linux-x86_64-ubuntu2204-8.0.4.tgz | |
| cp mongodb-linux-x86_64-ubuntu2204-8.0.4/bin/mongod /u/nlp/anaconda/main/anaconda3/envs/tonyreasoningtraces/bin/ | |
| mongod --version | |
| ``` | |
| ## Quick Start (Single GPU, No Sharding) | |
| ```bash | |
| python -m sdg.generate --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml --limit 100 | |
| ``` | |
| This requires `mongo_uri` set in the config YAML (see below). | |
| ## Distributed Sharded Run | |
| ### 1. Start MongoDB | |
| Launch a MongoDB server on a cluster node (no GPU needed) with max TTL of 21 days: | |
| ```bash | |
| nlprun -a tonyreasoningtraces --time 21-0 -m john3 --exclude john17 -c 4 -g 0 --memory 64g \ | |
| -w /nlp/scr4/nlp/crfm/text2image/text2image-rlhf/reasoning/virtual-world-data \ | |
| --job-name sdg-mongodb "python -m sdg.launch_mongodb 2>&1 | tee mongo.log" | |
| nlprun -a tonyreasoningtraces --time 21-0 -q jag -p high -m jagupard39 --exclude john17 -c 4 -g 0 --memory 64g \ | |
| -w /nlp/scr4/nlp/crfm/text2image/text2image-rlhf/reasoning/virtual-world-data \ | |
| --job-name sdg-serve "python -m sdg.launch_mongodb 2>&1 | tee mongo.log" | |
| ``` | |
| Or locally: | |
| ```bash | |
| python -m sdg.launch_mongodb --port 27017 --dbpath /tmp/sdg_mongo | |
| ``` | |
| It will print `mongo-uri: mongodb://<hostname>:27017` to stdout (and `mongo.log`). | |
| ### 2. Set `mongo_uri` and `num_shards` in your config | |
| Grab the hostname from `mongo.log`, then edit your YAML config (e.g. `sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml`): | |
| ```yaml | |
| mongo_uri: "mongodb://<hostname>:27017" | |
| num_shards: 32 | |
| ``` | |
| ### 3. Launch shards | |
| ```commandline | |
| conda activate tonyreasoningtraces && cd /nlp/scr4/nlp/crfm/text2image/text2image-rlhf/reasoning/virtual-world-data | |
| ``` | |
| Dry run first to verify commands: | |
| ```bash | |
| python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml launch --dry-run | |
| ``` | |
| Then launch for real: | |
| ```bash | |
| python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml launch | |
| ``` | |
| This reads `num_shards` from the config and submits that many nlprun jobs (1 GPU each). Each job processes `ceil(total_seeds / num_shards)` examples and writes to `output/{experiment_name}/shards/shard_NNN/`. You can override with `--num-shards N` on the CLI. | |
| ### 4. Check status | |
| ```bash | |
| python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml check | |
| ``` | |
| Prints a table like: | |
| ``` | |
| Shard Status Seeds Passed Failed | |
| 000 completed 2930 2100 830 | |
| 001 running - - - | |
| 002 failed - - - | |
| ... | |
| Overall: 20/32 completed, 8/32 running, 4/32 failed | |
| ``` | |
| When all shards are completed, `check` automatically: | |
| - Combines all shard `output.jsonl` files into `output/{experiment_name}/output.jsonl` | |
| - Aggregates all `stats.json` into a combined `output/{experiment_name}/stats.json` | |
| - Uploads to HuggingFace at `teetone/{experiment_name}` | |
| ### 5. Rerun failed shards | |
| ```bash | |
| python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml rerun | |
| ``` | |
| This re-launches only `failed` shards (skips `completed`, `running`, and `not_started`). | |
| ## Output Directory Structure | |
| ``` | |
| output/{experiment_name}/ | |
| launch_meta.json # Shard metadata | |
| config.yaml # Config copy | |
| output.jsonl # Combined output (after all shards done) | |
| stats.json # Aggregated stats (after all shards done) | |
| logs/ | |
| shard_000.log | |
| shard_001.log | |
| ... | |
| shards/ | |
| shard_000/ | |
| output.jsonl | |
| stats.json | |
| shard_001/ | |
| output.jsonl | |
| stats.json | |
| ... | |
| ``` | |
| ## Example: OpenThoughts4 with 32 shards | |
| ```bash | |
| # 1. Start MongoDB on john (no GPU, 16 GB RAM) | |
| nlprun -a tonyreasoningtraces -q john --exclude john17 -c 4 -g 0 --memory 64g \ | |
| -w /nlp/scr4/nlp/crfm/text2image/text2image-rlhf/reasoning/virtual-world-data \ | |
| --job-name sdg-mongodb "python -m sdg.launch_mongodb 2>&1 | tee mongo.log" | |
| # 2. Grab the hostname from the job output, then set mongo_uri in your config | |
| # mongo_uri: "mongodb://<hostname>:27017" | |
| # 3. Launch | |
| python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml launch | |
| # 4. Monitor | |
| python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml check | |
| # 5. Rerun any failures | |
| python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml rerun | |
| ``` | |
| ## CLI Reference | |
| | Command | Description | | |
| |---------|-------------| | |
| | `python -m sdg.generate --config CONFIG` | Run pipeline on a single GPU | | |
| | `python -m sdg.generate --config CONFIG --shard-id I --num-shards N` | Run a single shard | | |
| | `python -m sdg.launch_mongodb` | Start MongoDB server | | |
| | `python -m sdg.launcher --config CONFIG launch` | Launch all shards via nlprun | | |
| | `python -m sdg.launcher --config CONFIG launch --dry-run` | Print commands without executing | | |
| | `python -m sdg.launcher --config CONFIG check` | Check shard status, combine if done | | |
| | `python -m sdg.launcher --config CONFIG rerun` | Rerun failed shards only | | |
| | `python -m sdg.launcher --config CONFIG rerun --dry-run` | Show which failed shards would be relaunched | | |
| ## Launch Jobs | |
| ### Qwen3 4B | |
| #### Launch OpenThoughts3 Math53K | |
| ```bash | |
| # n=1, val_redundancy=3 - DONE | |
| python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n1_valredundancy3_round1.yaml launch | |
| # n=1, val_redundancy=1 - DONE | |
| python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n1_valredundancy1_round1.yaml launch | |
| # n=8, val_redundancy=1 - DONE | |
| python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n8_valredundancy1_round1.yaml launch | |
| # n=8, val_redundancy=3 - DONE | |
| python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n8_valredundancy3_round1.yaml launch | |
| # n=1, val_redundancy=5 | |
| python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n1_valredundancy5_round1.yaml launch | |
| # n=4, val_redundancy=5 | |
| python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n4_valredundancy5_round1.yaml launch | |
| # n=8, val_redundancy=5 - RUNNING | |
| python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n8_valredundancy5_round1.yaml launch | |
| # n=8, no_filter (no UQ validation) | |
| python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n8_no_filter_round1.yaml launch | |
| # n=8, val_redundancy=5, all_valid (keep all samples that pass UQ) | |
| python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n8_valredundancy5_allvalid_round1.yaml launch | |
| ``` | |
| #### Launch OpenThoughts4 Science26K | |
| ```bash | |
| # n=4, val_redundancy=3 - DONE | |
| python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_science26K_instill_n4_valredundancy3_round1.yaml launch | |
| # n=1, val_redundancy=1 - DONE | |
| python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_science26K_instill_n1_valredundancy1_round1.yaml launch | |
| # n=8, val_redundancy=3 - DONE | |
| python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_science26K_instill_n8_valredundancy3_round1.yaml launch | |
| # n=8, val_redundancy=5 - DONE | |
| python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_science26K_instill_n8_valredundancy5_round1.yaml launch | |
| # n=8, no_filter (no UQ validation) - RUNNING | |
| python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_science26K_instill_n8_no_filter_round1.yaml launch | |
| ``` | |
| #### Launch OpenThoughts4 Code9K | |
| ```bash | |
| # n=4, val_redundancy=3 - DONE | |
| python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_code9K_instill_n4_valredundancy3_round1.yaml launch | |
| # n=1, val_redundancy=1 - DONE | |
| python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_code9K_instill_n1_valredundancy1_round1.yaml launch | |
| # n=8, val_redundancy=3 - DONE | |
| python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_code9K_instill_n8_valredundancy3_round1.yaml launch | |
| # n=8, val_redundancy=5 - DONE | |
| python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_code9K_instill_n8_valredundancy5_round1.yaml launch | |
| # n=8, no_filter (no UQ validation) - RUNNING | |
| python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_code9K_instill_n8_no_filter_round1.yaml launch | |
| ``` | |
| ### Qwen3 8B | |
| #### Launch OpenThoughts4 Science26K (Qwen3-8B) | |
| ```bash | |
| # n=8, val_redundancy=5 - DONE | |
| python -m sdg.launcher --config sdg/configs/qwen3_8b_openthoughts4_science26K_instill_n8_valredundancy5_round1.yaml launch | |
| # n=4, val_redundancy=5 - RUNNING | |
| python -m sdg.launcher --config sdg/configs/qwen3_8b_openthoughts4_science26K_instill_n4_valredundancy5_round1.yaml launch | |
| ``` | |
| #### Launch OpenThoughts4 Code9K (Qwen3-8B) | |
| ```bash | |
| # n=8, val_redundancy=5 - DONE | |
| python -m sdg.launcher --config sdg/configs/qwen3_8b_openthoughts4_code9K_instill_n8_valredundancy5_round1.yaml launch | |
| ``` | |
| #### Launch OpenThoughts4 Math219K (Qwen3-8B) | |
| ```bash | |
| # n=8, val_redundancy=5 - RUNNING | |
| python -m sdg.launcher --config sdg/configs/qwen3_8b_openthoughts4_math219K_instill_n8_valredundancy5_round1.yaml launch | |
| ``` | |
| # Other | |
| ## Check logs | |
| ```bash | |
| scp -r tonyhlee@scdt.stanford.edu:/nlp/scr4/nlp/crfm/text2image/text2image-rlhf/reasoning/virtual-world-data/output/qwen3_0.6b_openthoughts3_math53K_instill_n8_valredundancy5_round1 /Users/tonyhlee/Dev/virtual-world-data/ | |
| ``` | |