# SDG — Distributed Sharded Pipeline Generate synthetic data with vLLM + uncertainty-quantification validation, distributed across multiple GPUs via SLURM sharding with a shared MongoDB inference cache. ## Prerequisites ```bash pip install -r sdg/requirements.txt ``` Install MongoDB (download binary): ```bash curl -O https://fastdl.mongodb.org/linux/mongodb-linux-x86_64-ubuntu2204-8.0.4.tgz tar xzf mongodb-linux-x86_64-ubuntu2204-8.0.4.tgz cp mongodb-linux-x86_64-ubuntu2204-8.0.4/bin/mongod /u/nlp/anaconda/main/anaconda3/envs/tonyreasoningtraces/bin/ mongod --version ``` ## Quick Start (Single GPU, No Sharding) ```bash python -m sdg.generate --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml --limit 100 ``` This requires `mongo_uri` set in the config YAML (see below). ## Distributed Sharded Run ### 1. Start MongoDB Launch a MongoDB server on a cluster node (no GPU needed) with max TTL of 21 days: ```bash nlprun -a tonyreasoningtraces --time 21-0 -m john3 --exclude john17 -c 4 -g 0 --memory 64g \ -w /nlp/scr4/nlp/crfm/text2image/text2image-rlhf/reasoning/virtual-world-data \ --job-name sdg-mongodb "python -m sdg.launch_mongodb 2>&1 | tee mongo.log" nlprun -a tonyreasoningtraces --time 21-0 -q jag -p high -m jagupard39 --exclude john17 -c 4 -g 0 --memory 64g \ -w /nlp/scr4/nlp/crfm/text2image/text2image-rlhf/reasoning/virtual-world-data \ --job-name sdg-serve "python -m sdg.launch_mongodb 2>&1 | tee mongo.log" ``` Or locally: ```bash python -m sdg.launch_mongodb --port 27017 --dbpath /tmp/sdg_mongo ``` It will print `mongo-uri: mongodb://:27017` to stdout (and `mongo.log`). ### 2. Set `mongo_uri` and `num_shards` in your config Grab the hostname from `mongo.log`, then edit your YAML config (e.g. `sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml`): ```yaml mongo_uri: "mongodb://:27017" num_shards: 32 ``` ### 3. Launch shards ```commandline conda activate tonyreasoningtraces && cd /nlp/scr4/nlp/crfm/text2image/text2image-rlhf/reasoning/virtual-world-data ``` Dry run first to verify commands: ```bash python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml launch --dry-run ``` Then launch for real: ```bash python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml launch ``` This reads `num_shards` from the config and submits that many nlprun jobs (1 GPU each). Each job processes `ceil(total_seeds / num_shards)` examples and writes to `output/{experiment_name}/shards/shard_NNN/`. You can override with `--num-shards N` on the CLI. ### 4. Check status ```bash python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml check ``` Prints a table like: ``` Shard Status Seeds Passed Failed 000 completed 2930 2100 830 001 running - - - 002 failed - - - ... Overall: 20/32 completed, 8/32 running, 4/32 failed ``` When all shards are completed, `check` automatically: - Combines all shard `output.jsonl` files into `output/{experiment_name}/output.jsonl` - Aggregates all `stats.json` into a combined `output/{experiment_name}/stats.json` - Uploads to HuggingFace at `teetone/{experiment_name}` ### 5. Rerun failed shards ```bash python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml rerun ``` This re-launches only `failed` shards (skips `completed`, `running`, and `not_started`). ## Output Directory Structure ``` output/{experiment_name}/ launch_meta.json # Shard metadata config.yaml # Config copy output.jsonl # Combined output (after all shards done) stats.json # Aggregated stats (after all shards done) logs/ shard_000.log shard_001.log ... shards/ shard_000/ output.jsonl stats.json shard_001/ output.jsonl stats.json ... ``` ## Example: OpenThoughts4 with 32 shards ```bash # 1. Start MongoDB on john (no GPU, 16 GB RAM) nlprun -a tonyreasoningtraces -q john --exclude john17 -c 4 -g 0 --memory 64g \ -w /nlp/scr4/nlp/crfm/text2image/text2image-rlhf/reasoning/virtual-world-data \ --job-name sdg-mongodb "python -m sdg.launch_mongodb 2>&1 | tee mongo.log" # 2. Grab the hostname from the job output, then set mongo_uri in your config # mongo_uri: "mongodb://:27017" # 3. Launch python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml launch # 4. Monitor python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml check # 5. Rerun any failures python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_math30K_instill_n4_valredundancy3_round1.yaml rerun ``` ## CLI Reference | Command | Description | |---------|-------------| | `python -m sdg.generate --config CONFIG` | Run pipeline on a single GPU | | `python -m sdg.generate --config CONFIG --shard-id I --num-shards N` | Run a single shard | | `python -m sdg.launch_mongodb` | Start MongoDB server | | `python -m sdg.launcher --config CONFIG launch` | Launch all shards via nlprun | | `python -m sdg.launcher --config CONFIG launch --dry-run` | Print commands without executing | | `python -m sdg.launcher --config CONFIG check` | Check shard status, combine if done | | `python -m sdg.launcher --config CONFIG rerun` | Rerun failed shards only | | `python -m sdg.launcher --config CONFIG rerun --dry-run` | Show which failed shards would be relaunched | ## Launch Jobs ### Qwen3 4B #### Launch OpenThoughts3 Math53K ```bash # n=1, val_redundancy=3 - DONE python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n1_valredundancy3_round1.yaml launch # n=1, val_redundancy=1 - DONE python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n1_valredundancy1_round1.yaml launch # n=8, val_redundancy=1 - DONE python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n8_valredundancy1_round1.yaml launch # n=8, val_redundancy=3 - DONE python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n8_valredundancy3_round1.yaml launch # n=1, val_redundancy=5 python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n1_valredundancy5_round1.yaml launch # n=4, val_redundancy=5 python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n4_valredundancy5_round1.yaml launch # n=8, val_redundancy=5 - RUNNING python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n8_valredundancy5_round1.yaml launch # n=8, no_filter (no UQ validation) python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n8_no_filter_round1.yaml launch # n=8, val_redundancy=5, all_valid (keep all samples that pass UQ) python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts3_math53K_instill_n8_valredundancy5_allvalid_round1.yaml launch ``` #### Launch OpenThoughts4 Science26K ```bash # n=4, val_redundancy=3 - DONE python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_science26K_instill_n4_valredundancy3_round1.yaml launch # n=1, val_redundancy=1 - DONE python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_science26K_instill_n1_valredundancy1_round1.yaml launch # n=8, val_redundancy=3 - DONE python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_science26K_instill_n8_valredundancy3_round1.yaml launch # n=8, val_redundancy=5 - DONE python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_science26K_instill_n8_valredundancy5_round1.yaml launch # n=8, no_filter (no UQ validation) - RUNNING python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_science26K_instill_n8_no_filter_round1.yaml launch ``` #### Launch OpenThoughts4 Code9K ```bash # n=4, val_redundancy=3 - DONE python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_code9K_instill_n4_valredundancy3_round1.yaml launch # n=1, val_redundancy=1 - DONE python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_code9K_instill_n1_valredundancy1_round1.yaml launch # n=8, val_redundancy=3 - DONE python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_code9K_instill_n8_valredundancy3_round1.yaml launch # n=8, val_redundancy=5 - DONE python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_code9K_instill_n8_valredundancy5_round1.yaml launch # n=8, no_filter (no UQ validation) - RUNNING python -m sdg.launcher --config sdg/configs/qwen3_4b_openthoughts4_code9K_instill_n8_no_filter_round1.yaml launch ``` ### Qwen3 8B #### Launch OpenThoughts4 Science26K (Qwen3-8B) ```bash # n=8, val_redundancy=5 - DONE python -m sdg.launcher --config sdg/configs/qwen3_8b_openthoughts4_science26K_instill_n8_valredundancy5_round1.yaml launch # n=4, val_redundancy=5 - RUNNING python -m sdg.launcher --config sdg/configs/qwen3_8b_openthoughts4_science26K_instill_n4_valredundancy5_round1.yaml launch ``` #### Launch OpenThoughts4 Code9K (Qwen3-8B) ```bash # n=8, val_redundancy=5 - DONE python -m sdg.launcher --config sdg/configs/qwen3_8b_openthoughts4_code9K_instill_n8_valredundancy5_round1.yaml launch ``` #### Launch OpenThoughts4 Math219K (Qwen3-8B) ```bash # n=8, val_redundancy=5 - RUNNING python -m sdg.launcher --config sdg/configs/qwen3_8b_openthoughts4_math219K_instill_n8_valredundancy5_round1.yaml launch ``` # Other ## Check logs ```bash scp -r tonyhlee@scdt.stanford.edu:/nlp/scr4/nlp/crfm/text2image/text2image-rlhf/reasoning/virtual-world-data/output/qwen3_0.6b_openthoughts3_math53K_instill_n8_valredundancy5_round1 /Users/tonyhlee/Dev/virtual-world-data/ ```