|
Download doc/moe_experiments.md from TerryPei/GroundFlow: direct link, hf CLI and curl.
- Browser
- Download file 8.7 kB
-
https://huggingface.co/TerryPei/GroundFlow/resolve/main/doc/moe_experiments.md
- Command line
-
hf download hf://TerryPei/GroundFlow/doc/moe_experiments.md
-
curl -L -o moe_experiments.md https://huggingface.co/TerryPei/GroundFlow/resolve/main/doc/moe_experiments.md
8.7 kB
30B-A3B MoE β Training & Inference File Map
Authoritative reference for every training script, DeepSpeed config, modeling file, and eval path used for Qwen3-VL-30B-A3B-Instruct (MoE). Mirrors the dense (4B/8B) pipeline but uses separate modeling & launcher files so 4B/8B work is never impacted by MoE changes.
Last verified: 2026-04-12.
1. Core code paths
| Role | Path |
|---|---|
| Training entry (shared) | QWENVL-PRIVATE/qwen-vl-finetune/qwenvl/train/train_qwen.py |
| MoE dispatch | train_qwen.py:179-184 β loads Qwen3VLMoeForConditionalGeneration when config.model_type == 'qwen3_vl_moe'. Try/except guarded import at line 48-52 so dense-only env still works. |
| Modeling (MoE) | QWENVL-PRIVATE/qwen_src/qwen3_vl/modeling_qwen3_vl_moe_batch.py |
| IWA-aware attention | modeling_qwen3_vl_moe_batch.py:40 β _iwa_aware_attn_forward, monkey-patches Qwen3VLMoeTextAttention.forward |
| Main class | modeling_qwen3_vl_moe_batch.py:112 β Qwen3VLMoeForConditionalGeneration(_BaseModel) |
| Modeling (dense, for ref) | QWENVL-PRIVATE/qwen_src/qwen3_vl/modeling_qwen3_vl_batch.py β 4B/8B/32B |
| Shared utils (arch-agnostic) | QWENVL-PRIVATE/qwen_src/mm_utils.py, qwen_src/qwen3_vl/iwa.py, configuration_qwen3_vl.py, processing_qwen3_vl.py |
Isolation guarantees (proven 2026-04-12)
- Dense β MoE: dense modeling never imports MoE. Safe.
- MoE β Dense: one controlled import in
modeling_qwen3_vl_moe_batch.py:848pulls three utils (get_singleturn_query_text_hs_mheads, apply_rotary_pos_emb, repeat_kv) from dense. Arch-agnostic β safe. git log 10070ed..5602a58 -- QWENVL-PRIVATE/qwen_src/qwen3_vl/modeling_qwen3_vl_batch.pyreturns empty (the 5 MoE commits did NOT touch the dense file).
2. Training scripts (all in QWENVL-PRIVATE/tools/)
2.1 Pre-training / Stage 1 (twig training β learning to route ROI)
| Script | GPUs | Notes |
|---|---|---|
stage1_30b_8gpu.sh |
1Γ8 | single-node baseline |
stage1_30b_6gpu.sh |
1Γ6 | reduced-GPU variant |
stage1_30b_natural.sh |
β | natural-image pseudo-label variant |
stage1_stage2_30b.sh |
1Γ8 | stage1+stage2 combined (full twig pipeline) |
2.2 Stage 3 β SD-RPN spatial-only (Ξ±=1.0, pure ROI loss, no KL)
| Script | GPUs | Status | Current result |
|---|---|---|---|
stage3_30b_4node_8gpus_spatial_only.sh |
4Γ8=32 | β used | 30B SD-RPN, OCR=826 (tied with vanilla baseline; no gain yet) |
stage3_30b_8node_8gpus_spatial_only.sh |
8Γ8=64 | scale-out | β |
Key flags in spatial_only: --enable_twig True --twig_K 39 --twig_T 3 --distill_weight 1.0 --enable_iwa False.
2.3 Stage 3 β Confluent Distillation (Ξ±=0.667, ROI + KL + IWA)
| Script | GPUs | Version | Status |
|---|---|---|---|
stage3_30b_confluent.sh |
single-node | legacy | superseded |
stage3_30b_4node_8gpus_confluent.sh |
4Γ8=32 | v1 (Ο=2.0) | kept for reference |
stage3_30b_2node_8gpus_confluent_v2.sh |
2Γ8=16 | v2 (Ο=1.5) | CURRENT smoke-test |
stage3_30b_4node_8gpus_confluent_v2.sh |
4Γ8=32 | v2 (Ο=1.5) | CURRENT production |
stage3_30b_8node_8gpus_confluent_v2.sh |
8Γ8=64 | v2 (Ο=1.5) | CURRENT scale-out |
v1 β v2 delta (from diff stage3_30b_4node_8gpus_confluent.sh stage3_30b_4node_8gpus_confluent_v2.sh):
- Output dir tag:
confluent-fixedβconfluent-v2-fixed - Default
TAU: 2.0 β 1.5
Shared confluent v2 flags:
--enable_twig True \
--twig_K 39 \
--twig_T 3 \
--distill_weight 0.667 \
--enable_iwa True \
TAU=1.5 \
ALPHA=0.667
2.4 Stage 3 β Generic / legacy templates
| Script | Purpose |
|---|---|
stage3_30b_4node_8gpus.sh |
generic multi-node (no method tag) |
stage3_30b_8node_8gpus.sh |
generic 8-node |
stage3_30b_multinode.sh |
generic multi-node template |
stage3_30b_fsdp.sh |
FSDP-only template |
stage3_30b_deepspeed.sh |
tiny DeepSpeed wrapper |
These are not currently used for the Confluent pipeline β prefer v2 scripts above.
2.5 DeepSpeed configs
| File | ZeRO Stage | Notes |
|---|---|---|
tools/ds_zero2_30b.json |
2 | smaller memory ceiling |
tools/ds_zero3_30b.json |
3 | used by confluent v2 together with FSDP activation checkpointing. Don't enable HF gradient_checkpointing alongside FSDP's activation_checkpointing: true β double-wrap issues. |
3. Analysis & data-prep scripts
| Path | Purpose |
|---|---|
QWENVL-PRIVATE/make_pseudo_label_natural_30b.py |
natural-image ROI pseudo-labels |
QWENVL-PRIVATE/make_pseudo_label_textual_30b.py |
text-image ROI pseudo-labels |
QWENVL-PRIVATE/analysis/qwen3vl_debug_30b.py |
ROI pseudo-label map generator (placeholder heads) |
QWENVL-PRIVATE/analysis/qwen3vl_identify_heads_30b.py |
per-layer + per-head attention visualization (picks grounding heads) |
QWENVL-PRIVATE/analysis/gen_roi_map_30b.py |
ROI map gen helper |
QWENVL-PRIVATE/tools/identify_heads_30b.sh |
head-identification runner wrapper |
QWENVL-PRIVATE/tools/identify_heads_30b_generic.sh |
head-ID generic wrapper |
4. Inference / eval file map
CLI --model |
Source file | Used for |
|---|---|---|
qwen3_vl |
lmms-eval/lmms_eval/models/simple/qwen3_vl.py |
30B MoE (auto-dispatches via config.model_type == 'qwen3_vl_moe') β also handles 4B/8B fallback |
qwen3_vl_hybrid |
.../simple/qwen3_vl_hybrid.py |
4B/8B dense (run7b-golden, matches commit 6360f83 byte-for-byte) |
qwen3_vl_hybrid_30b_moe |
.../simple/qwen3_vl_hybrid_30b_moe.py |
30B MoE hybrid (NEW 2026-04-12, commits 76697057e + a560bd626) β mirror of qwen3_vl_hybrid.py using MoE class |
Canonical eval commands
30B MoE via simple (auto-dispatch):
export PYTHONPATH=/opt/tiger/thothvl_pretrain/QWENVL-PRIVATE:$PYTHONPATH
export HF_HOME=/mnt/bn/leonworkspace/HF_HOME
export HF_TOKEN=hf_REDACTED
CKPT=/mnt/bn/leonworkspace/terry/model/qwen3vl-30b-roi-K39T3-185k-confluent-v2-fixed-<RUN_ID>
python3 -m lmms_eval --model qwen3_vl --model_args \
"pretrained=$CKPT,device_map=auto,two_stage_roi=True,roi_baseline=True,roi_conf_thresh=0.15,high_res_thresh=0.1,attn_implementation=flash_attention_2" \
--tasks ocrbench --batch_size 1 --output_path ./results
30B MoE via hybrid (isolated file):
python3 -m lmms_eval --model qwen3_vl_hybrid_30b_moe --model_args "<same args>" \
--tasks ocrbench --batch_size 1 --output_path ./results
5. Known issues & fixes (2026-04-12 saga)
| Commit | Fix |
|---|---|
10070ed9c |
OOM: KL distillation recomputed lm_head three times; reuse twig_logits |
e7428e791 |
IWA was created but never wired in MoE training forward β added compute_importance + per-layer _iwa_bias |
67063ecb7 |
Added _compute_roi_maps_moe helper |
506d1b666 |
Critical: DeepStack visual features were being dropped (silent tuple [0] truncation). Added get_image_features / get_video_features path |
5602a581f |
Implemented two-stage inference in MoE forward (shared prefix β twig β ROI map β sub-image reroute β backbone) |
b8724e4f1 |
Skip IWA in two-stage reroute path (matches dense) |
b914a711c |
Eval dispatch (MoE via hybrid) β later deleted |
dd18961ba |
Deleted qwen3_vl_hybrid.py β regression (4B/8B 852β845) |
4d9e06434 + da60dfc29 |
Recovered 4B/8B by porting 14 SD-RPN attrs into qwen3_vl.py simple |
Current smoke-test status (as of 2026-04-12)
- 30B Confluent v2 training: not yet launched end-to-end after the 5-commit MoE alignment. Needs a run with
stage3_30b_2node_8gpus_confluent_v2.shto verify (a) no OOM, (b) DeepStack flows, (c) IWA gamma in optimizer, (d) loss decreases. - 30B SD-RPN spatial_only: current result OCR=826 (= vanilla baseline) β may be missing DeepStack (see fix
506d1b666). Retrain required. - 30B MoE hybrid eval file (
qwen3_vl_hybrid_30b_moe.py): registered, import-tested, not yet run against a real MoE checkpoint.
6. GPU / node requirements
| Scale | Total GPUs | GPU RAM needed | Script |
|---|---|---|---|
| Smoke-test | 16 | each ~60GB during training | stage3_30b_2node_8gpus_confluent_v2.sh |
| Production | 32 | each ~45GB | stage3_30b_4node_8gpus_confluent_v2.sh |
| Scale-out | 64 | each ~30GB | stage3_30b_8node_8gpus_confluent_v2.sh |
ETA per FSDP pass of ~185k samples:
- 16 GPUs: ~44 h
- 32 GPUs: ~22 h
- 64 GPUs: ~11 h