GroundFlow / doc /moe_experiments.md
TerryPei's picture
sync doc from /opt/tiger/thothvl_pretrain (HF tokens redacted)
849b4fe verified
|
Raw History Blame Contribute Delete
8.7 kB

30B-A3B MoE β€” Training & Inference File Map

Authoritative reference for every training script, DeepSpeed config, modeling file, and eval path used for Qwen3-VL-30B-A3B-Instruct (MoE). Mirrors the dense (4B/8B) pipeline but uses separate modeling & launcher files so 4B/8B work is never impacted by MoE changes.

Last verified: 2026-04-12.


1. Core code paths

Role Path
Training entry (shared) QWENVL-PRIVATE/qwen-vl-finetune/qwenvl/train/train_qwen.py
MoE dispatch train_qwen.py:179-184 β€” loads Qwen3VLMoeForConditionalGeneration when config.model_type == 'qwen3_vl_moe'. Try/except guarded import at line 48-52 so dense-only env still works.
Modeling (MoE) QWENVL-PRIVATE/qwen_src/qwen3_vl/modeling_qwen3_vl_moe_batch.py
IWA-aware attention modeling_qwen3_vl_moe_batch.py:40 β€” _iwa_aware_attn_forward, monkey-patches Qwen3VLMoeTextAttention.forward
Main class modeling_qwen3_vl_moe_batch.py:112 β€” Qwen3VLMoeForConditionalGeneration(_BaseModel)
Modeling (dense, for ref) QWENVL-PRIVATE/qwen_src/qwen3_vl/modeling_qwen3_vl_batch.py β€” 4B/8B/32B
Shared utils (arch-agnostic) QWENVL-PRIVATE/qwen_src/mm_utils.py, qwen_src/qwen3_vl/iwa.py, configuration_qwen3_vl.py, processing_qwen3_vl.py

Isolation guarantees (proven 2026-04-12)

  • Dense β†’ MoE: dense modeling never imports MoE. Safe.
  • MoE β†’ Dense: one controlled import in modeling_qwen3_vl_moe_batch.py:848 pulls three utils (get_singleturn_query_text_hs_mheads, apply_rotary_pos_emb, repeat_kv) from dense. Arch-agnostic β€” safe.
  • git log 10070ed..5602a58 -- QWENVL-PRIVATE/qwen_src/qwen3_vl/modeling_qwen3_vl_batch.py returns empty (the 5 MoE commits did NOT touch the dense file).

2. Training scripts (all in QWENVL-PRIVATE/tools/)

2.1 Pre-training / Stage 1 (twig training β€” learning to route ROI)

Script GPUs Notes
stage1_30b_8gpu.sh 1Γ—8 single-node baseline
stage1_30b_6gpu.sh 1Γ—6 reduced-GPU variant
stage1_30b_natural.sh β€” natural-image pseudo-label variant
stage1_stage2_30b.sh 1Γ—8 stage1+stage2 combined (full twig pipeline)

2.2 Stage 3 β€” SD-RPN spatial-only (Ξ±=1.0, pure ROI loss, no KL)

Script GPUs Status Current result
stage3_30b_4node_8gpus_spatial_only.sh 4Γ—8=32 βœ“ used 30B SD-RPN, OCR=826 (tied with vanilla baseline; no gain yet)
stage3_30b_8node_8gpus_spatial_only.sh 8Γ—8=64 scale-out β€”

Key flags in spatial_only: --enable_twig True --twig_K 39 --twig_T 3 --distill_weight 1.0 --enable_iwa False.

2.3 Stage 3 β€” Confluent Distillation (Ξ±=0.667, ROI + KL + IWA)

Script GPUs Version Status
stage3_30b_confluent.sh single-node legacy superseded
stage3_30b_4node_8gpus_confluent.sh 4Γ—8=32 v1 (Ο„=2.0) kept for reference
stage3_30b_2node_8gpus_confluent_v2.sh 2Γ—8=16 v2 (Ο„=1.5) CURRENT smoke-test
stage3_30b_4node_8gpus_confluent_v2.sh 4Γ—8=32 v2 (Ο„=1.5) CURRENT production
stage3_30b_8node_8gpus_confluent_v2.sh 8Γ—8=64 v2 (Ο„=1.5) CURRENT scale-out

v1 β†’ v2 delta (from diff stage3_30b_4node_8gpus_confluent.sh stage3_30b_4node_8gpus_confluent_v2.sh):

  • Output dir tag: confluent-fixed β†’ confluent-v2-fixed
  • Default TAU: 2.0 β†’ 1.5

Shared confluent v2 flags:

--enable_twig True     \
--twig_K 39            \
--twig_T 3             \
--distill_weight 0.667 \
--enable_iwa True      \
TAU=1.5                \
ALPHA=0.667

2.4 Stage 3 β€” Generic / legacy templates

Script Purpose
stage3_30b_4node_8gpus.sh generic multi-node (no method tag)
stage3_30b_8node_8gpus.sh generic 8-node
stage3_30b_multinode.sh generic multi-node template
stage3_30b_fsdp.sh FSDP-only template
stage3_30b_deepspeed.sh tiny DeepSpeed wrapper

These are not currently used for the Confluent pipeline β€” prefer v2 scripts above.

2.5 DeepSpeed configs

File ZeRO Stage Notes
tools/ds_zero2_30b.json 2 smaller memory ceiling
tools/ds_zero3_30b.json 3 used by confluent v2 together with FSDP activation checkpointing. Don't enable HF gradient_checkpointing alongside FSDP's activation_checkpointing: true β€” double-wrap issues.

3. Analysis & data-prep scripts

Path Purpose
QWENVL-PRIVATE/make_pseudo_label_natural_30b.py natural-image ROI pseudo-labels
QWENVL-PRIVATE/make_pseudo_label_textual_30b.py text-image ROI pseudo-labels
QWENVL-PRIVATE/analysis/qwen3vl_debug_30b.py ROI pseudo-label map generator (placeholder heads)
QWENVL-PRIVATE/analysis/qwen3vl_identify_heads_30b.py per-layer + per-head attention visualization (picks grounding heads)
QWENVL-PRIVATE/analysis/gen_roi_map_30b.py ROI map gen helper
QWENVL-PRIVATE/tools/identify_heads_30b.sh head-identification runner wrapper
QWENVL-PRIVATE/tools/identify_heads_30b_generic.sh head-ID generic wrapper

4. Inference / eval file map

CLI --model Source file Used for
qwen3_vl lmms-eval/lmms_eval/models/simple/qwen3_vl.py 30B MoE (auto-dispatches via config.model_type == 'qwen3_vl_moe') β€” also handles 4B/8B fallback
qwen3_vl_hybrid .../simple/qwen3_vl_hybrid.py 4B/8B dense (run7b-golden, matches commit 6360f83 byte-for-byte)
qwen3_vl_hybrid_30b_moe .../simple/qwen3_vl_hybrid_30b_moe.py 30B MoE hybrid (NEW 2026-04-12, commits 76697057e + a560bd626) β€” mirror of qwen3_vl_hybrid.py using MoE class

Canonical eval commands

30B MoE via simple (auto-dispatch):

export PYTHONPATH=/opt/tiger/thothvl_pretrain/QWENVL-PRIVATE:$PYTHONPATH
export HF_HOME=/mnt/bn/leonworkspace/HF_HOME
export HF_TOKEN=hf_REDACTED

CKPT=/mnt/bn/leonworkspace/terry/model/qwen3vl-30b-roi-K39T3-185k-confluent-v2-fixed-<RUN_ID>
python3 -m lmms_eval --model qwen3_vl --model_args \
    "pretrained=$CKPT,device_map=auto,two_stage_roi=True,roi_baseline=True,roi_conf_thresh=0.15,high_res_thresh=0.1,attn_implementation=flash_attention_2" \
    --tasks ocrbench --batch_size 1 --output_path ./results

30B MoE via hybrid (isolated file):

python3 -m lmms_eval --model qwen3_vl_hybrid_30b_moe --model_args "<same args>" \
    --tasks ocrbench --batch_size 1 --output_path ./results

5. Known issues & fixes (2026-04-12 saga)

Commit Fix
10070ed9c OOM: KL distillation recomputed lm_head three times; reuse twig_logits
e7428e791 IWA was created but never wired in MoE training forward β€” added compute_importance + per-layer _iwa_bias
67063ecb7 Added _compute_roi_maps_moe helper
506d1b666 Critical: DeepStack visual features were being dropped (silent tuple [0] truncation). Added get_image_features / get_video_features path
5602a581f Implemented two-stage inference in MoE forward (shared prefix β†’ twig β†’ ROI map β†’ sub-image reroute β†’ backbone)
b8724e4f1 Skip IWA in two-stage reroute path (matches dense)
b914a711c Eval dispatch (MoE via hybrid) β€” later deleted
dd18961ba Deleted qwen3_vl_hybrid.py β€” regression (4B/8B 852β†’845)
4d9e06434 + da60dfc29 Recovered 4B/8B by porting 14 SD-RPN attrs into qwen3_vl.py simple

Current smoke-test status (as of 2026-04-12)

  • 30B Confluent v2 training: not yet launched end-to-end after the 5-commit MoE alignment. Needs a run with stage3_30b_2node_8gpus_confluent_v2.sh to verify (a) no OOM, (b) DeepStack flows, (c) IWA gamma in optimizer, (d) loss decreases.
  • 30B SD-RPN spatial_only: current result OCR=826 (= vanilla baseline) β€” may be missing DeepStack (see fix 506d1b666). Retrain required.
  • 30B MoE hybrid eval file (qwen3_vl_hybrid_30b_moe.py): registered, import-tested, not yet run against a real MoE checkpoint.

6. GPU / node requirements

Scale Total GPUs GPU RAM needed Script
Smoke-test 16 each ~60GB during training stage3_30b_2node_8gpus_confluent_v2.sh
Production 32 each ~45GB stage3_30b_4node_8gpus_confluent_v2.sh
Scale-out 64 each ~30GB stage3_30b_8node_8gpus_confluent_v2.sh

ETA per FSDP pass of ~185k samples:

  • 16 GPUs: ~44 h
  • 32 GPUs: ~22 h
  • 64 GPUs: ~11 h