# 30B-A3B MoE — Training & Inference File Map Authoritative reference for every training script, DeepSpeed config, modeling file, and eval path used for **Qwen3-VL-30B-A3B-Instruct** (MoE). Mirrors the dense (4B/8B) pipeline but uses separate modeling & launcher files so 4B/8B work is never impacted by MoE changes. Last verified: 2026-04-12. --- ## 1. Core code paths | Role | Path | |------|------| | Training entry (shared) | `QWENVL-PRIVATE/qwen-vl-finetune/qwenvl/train/train_qwen.py` | | MoE dispatch | `train_qwen.py:179-184` — loads `Qwen3VLMoeForConditionalGeneration` when `config.model_type == 'qwen3_vl_moe'`. Try/except guarded import at line 48-52 so dense-only env still works. | | **Modeling (MoE)** | **`QWENVL-PRIVATE/qwen_src/qwen3_vl/modeling_qwen3_vl_moe_batch.py`** | | IWA-aware attention | `modeling_qwen3_vl_moe_batch.py:40` — `_iwa_aware_attn_forward`, monkey-patches `Qwen3VLMoeTextAttention.forward` | | Main class | `modeling_qwen3_vl_moe_batch.py:112` — `Qwen3VLMoeForConditionalGeneration(_BaseModel)` | | Modeling (dense, for ref) | `QWENVL-PRIVATE/qwen_src/qwen3_vl/modeling_qwen3_vl_batch.py` — 4B/8B/32B | | Shared utils (arch-agnostic) | `QWENVL-PRIVATE/qwen_src/mm_utils.py`, `qwen_src/qwen3_vl/iwa.py`, `configuration_qwen3_vl.py`, `processing_qwen3_vl.py` | ### Isolation guarantees (proven 2026-04-12) - **Dense → MoE**: dense modeling never imports MoE. Safe. - **MoE → Dense**: one controlled import in `modeling_qwen3_vl_moe_batch.py:848` pulls three utils (`get_singleturn_query_text_hs_mheads, apply_rotary_pos_emb, repeat_kv`) from dense. Arch-agnostic — safe. - `git log 10070ed..5602a58 -- QWENVL-PRIVATE/qwen_src/qwen3_vl/modeling_qwen3_vl_batch.py` returns **empty** (the 5 MoE commits did NOT touch the dense file). --- ## 2. Training scripts (all in `QWENVL-PRIVATE/tools/`) ### 2.1 Pre-training / Stage 1 (twig training — learning to route ROI) | Script | GPUs | Notes | |--------|------|-------| | `stage1_30b_8gpu.sh` | 1×8 | single-node baseline | | `stage1_30b_6gpu.sh` | 1×6 | reduced-GPU variant | | `stage1_30b_natural.sh` | — | natural-image pseudo-label variant | | `stage1_stage2_30b.sh` | 1×8 | stage1+stage2 combined (full twig pipeline) | ### 2.2 Stage 3 — SD-RPN spatial-only (α=1.0, pure ROI loss, no KL) | Script | GPUs | Status | Current result | |--------|------|--------|----------------| | `stage3_30b_4node_8gpus_spatial_only.sh` | 4×8=32 | ✓ used | 30B SD-RPN, OCR=826 (tied with vanilla baseline; no gain yet) | | `stage3_30b_8node_8gpus_spatial_only.sh` | 8×8=64 | scale-out | — | Key flags in spatial_only: `--enable_twig True --twig_K 39 --twig_T 3 --distill_weight 1.0 --enable_iwa False`. ### 2.3 Stage 3 — Confluent Distillation (α=0.667, ROI + KL + IWA) | Script | GPUs | Version | Status | |--------|------|---------|--------| | `stage3_30b_confluent.sh` | single-node | legacy | superseded | | `stage3_30b_4node_8gpus_confluent.sh` | 4×8=32 | **v1** (τ=2.0) | kept for reference | | **`stage3_30b_2node_8gpus_confluent_v2.sh`** | 2×8=**16** | **v2** (τ=1.5) | **CURRENT smoke-test** | | **`stage3_30b_4node_8gpus_confluent_v2.sh`** | 4×8=**32** | **v2** (τ=1.5) | **CURRENT production** | | **`stage3_30b_8node_8gpus_confluent_v2.sh`** | 8×8=**64** | **v2** (τ=1.5) | **CURRENT scale-out** | **v1 → v2 delta** (from `diff stage3_30b_4node_8gpus_confluent.sh stage3_30b_4node_8gpus_confluent_v2.sh`): - Output dir tag: `confluent-fixed` → `confluent-v2-fixed` - Default `TAU`: 2.0 → 1.5 **Shared confluent v2 flags:** ``` --enable_twig True \ --twig_K 39 \ --twig_T 3 \ --distill_weight 0.667 \ --enable_iwa True \ TAU=1.5 \ ALPHA=0.667 ``` ### 2.4 Stage 3 — Generic / legacy templates | Script | Purpose | |--------|---------| | `stage3_30b_4node_8gpus.sh` | generic multi-node (no method tag) | | `stage3_30b_8node_8gpus.sh` | generic 8-node | | `stage3_30b_multinode.sh` | generic multi-node template | | `stage3_30b_fsdp.sh` | FSDP-only template | | `stage3_30b_deepspeed.sh` | tiny DeepSpeed wrapper | These are **not** currently used for the Confluent pipeline — prefer v2 scripts above. ### 2.5 DeepSpeed configs | File | ZeRO Stage | Notes | |------|-----------|-------| | `tools/ds_zero2_30b.json` | 2 | smaller memory ceiling | | `tools/ds_zero3_30b.json` | 3 | used by confluent v2 together with FSDP activation checkpointing. **Don't enable HF `gradient_checkpointing` alongside FSDP's `activation_checkpointing: true` — double-wrap issues.** | --- ## 3. Analysis & data-prep scripts | Path | Purpose | |------|---------| | `QWENVL-PRIVATE/make_pseudo_label_natural_30b.py` | natural-image ROI pseudo-labels | | `QWENVL-PRIVATE/make_pseudo_label_textual_30b.py` | text-image ROI pseudo-labels | | `QWENVL-PRIVATE/analysis/qwen3vl_debug_30b.py` | ROI pseudo-label map generator (placeholder heads) | | `QWENVL-PRIVATE/analysis/qwen3vl_identify_heads_30b.py` | per-layer + per-head attention visualization (picks grounding heads) | | `QWENVL-PRIVATE/analysis/gen_roi_map_30b.py` | ROI map gen helper | | `QWENVL-PRIVATE/tools/identify_heads_30b.sh` | head-identification runner wrapper | | `QWENVL-PRIVATE/tools/identify_heads_30b_generic.sh` | head-ID generic wrapper | --- ## 4. Inference / eval file map | CLI `--model` | Source file | Used for | |---------------|-------------|----------| | `qwen3_vl` | `lmms-eval/lmms_eval/models/simple/qwen3_vl.py` | 30B MoE (auto-dispatches via `config.model_type == 'qwen3_vl_moe'`) — also handles 4B/8B fallback | | `qwen3_vl_hybrid` | `.../simple/qwen3_vl_hybrid.py` | 4B/8B dense (run7b-golden, matches commit `6360f83` byte-for-byte) | | **`qwen3_vl_hybrid_30b_moe`** | `.../simple/qwen3_vl_hybrid_30b_moe.py` | **30B MoE hybrid (NEW 2026-04-12, commits 76697057e + a560bd626)** — mirror of `qwen3_vl_hybrid.py` using MoE class | ### Canonical eval commands **30B MoE via simple (auto-dispatch):** ```bash export PYTHONPATH=/opt/tiger/thothvl_pretrain/QWENVL-PRIVATE:$PYTHONPATH export HF_HOME=/mnt/bn/leonworkspace/HF_HOME export HF_TOKEN=hf_REDACTED CKPT=/mnt/bn/leonworkspace/terry/model/qwen3vl-30b-roi-K39T3-185k-confluent-v2-fixed- python3 -m lmms_eval --model qwen3_vl --model_args \ "pretrained=$CKPT,device_map=auto,two_stage_roi=True,roi_baseline=True,roi_conf_thresh=0.15,high_res_thresh=0.1,attn_implementation=flash_attention_2" \ --tasks ocrbench --batch_size 1 --output_path ./results ``` **30B MoE via hybrid (isolated file):** ```bash python3 -m lmms_eval --model qwen3_vl_hybrid_30b_moe --model_args "" \ --tasks ocrbench --batch_size 1 --output_path ./results ``` --- ## 5. Known issues & fixes (2026-04-12 saga) | Commit | Fix | |--------|-----| | `10070ed9c` | OOM: KL distillation recomputed `lm_head` three times; reuse `twig_logits` | | `e7428e791` | IWA was created but never wired in MoE training forward — added `compute_importance` + per-layer `_iwa_bias` | | `67063ecb7` | Added `_compute_roi_maps_moe` helper | | `506d1b666` | Critical: DeepStack visual features were being dropped (silent tuple `[0]` truncation). Added `get_image_features` / `get_video_features` path | | `5602a581f` | Implemented two-stage inference in MoE forward (shared prefix → twig → ROI map → sub-image reroute → backbone) | | `b8724e4f1` | Skip IWA in two-stage reroute path (matches dense) | | `b914a711c` | Eval dispatch (MoE via hybrid) — later deleted | | `dd18961ba` | Deleted `qwen3_vl_hybrid.py` — regression (4B/8B 852→845) | | `4d9e06434` + `da60dfc29` | Recovered 4B/8B by porting 14 SD-RPN attrs into `qwen3_vl.py` simple | ### Current smoke-test status (as of 2026-04-12) - 30B Confluent v2 training: not yet launched end-to-end after the 5-commit MoE alignment. Needs a run with `stage3_30b_2node_8gpus_confluent_v2.sh` to verify (a) no OOM, (b) DeepStack flows, (c) IWA gamma in optimizer, (d) loss decreases. - 30B SD-RPN spatial_only: current result OCR=826 (= vanilla baseline) — may be missing DeepStack (see fix `506d1b666`). Retrain required. - 30B MoE hybrid eval file (`qwen3_vl_hybrid_30b_moe.py`): registered, import-tested, not yet run against a real MoE checkpoint. --- ## 6. GPU / node requirements | Scale | Total GPUs | GPU RAM needed | Script | |-------|------------|----------------|--------| | Smoke-test | 16 | each ~60GB during training | `stage3_30b_2node_8gpus_confluent_v2.sh` | | Production | 32 | each ~45GB | `stage3_30b_4node_8gpus_confluent_v2.sh` | | Scale-out | 64 | each ~30GB | `stage3_30b_8node_8gpus_confluent_v2.sh` | ETA per FSDP pass of ~185k samples: - 16 GPUs: ~44 h - 32 GPUs: ~22 h - 64 GPUs: ~11 h