colon-videogen: mask track conditioned colonoscopy video generation
Research artifacts from an early stage project on temporally coherent colonoscopy video synthesis with exact lesion labels, evaluated by downstream polyp detection. Code and the full research record are in the companion Git repository (bundle included under code/).
This is preliminary research output, not a product and not a clinical tool.
What is here
models/
Video diffusion checkpoints (57M parameter latent video DiT, mask track channel concatenated in latent space, SD f8 VAE frozen, adaLN-zero blocks, full space time attention). Each file holds model, ema, step, args.
colon-exp024c-f1-256: the main conditional model at 256px. ckpt_0020000.pt is the best checkpoint (adherence Dice 0.486, judge found rate 0.89, area ratio 1.33). Quality degrades after this point: 40k gives 0.462, 60k gives 0.366.colon-exp024-f1-pilot: 128px conditional, 200k steps. Kept as a documented negative result: at 128px the conditioning does not survive the 16x16 latent grid (adherence 0.071).colon-exp024b-uncond-pilot: 128px unconditional null control (adherence 0.013).colon-exp028b-uncond256: 256px unconditional run that diverged to NaN between step 20k and 40k. Only ckpt_0020000.pt is valid; the 40k and 60k files are all NaN and are kept only as a record of the failure.colon-exp028c-uncond256-s1: replacement unconditional control, different seed.colon-exp029-256-lw12: lesion loss weight 12 instead of 4. Raises found rate to 0.80 but does not fix small lesion adherence (0.221 versus 0.233).colon-exp030-256-long: second seed of the 256px conditional run, stopped early at 80k once the degradation trajectory was confirmed.colon-exp023-f1-overfit: 8 clip overfit sanity check.kvasir_segformer_b1_judge: SegFormer-b1 fine tuned on Kvasir-SEG (val Dice 0.897), used as an independent adherence judge. It is an evaluation utility, not a benchmark segmentation model.
data/
Only artifacts generated or annotated by this project.
synth_video_v2_from20k: 2,740 synthetic 256px frames with exact YOLO boxes derived from the conditioning masks, generated from the best checkpoint. Alsolabels_pseudo/from the SegFormer judge for the label provenance experiment.synth_video_v1_from60k: the same pipeline run from the degraded 60k checkpoint, kept because the comparison between the two is the point of the experiment.synth_b2_image_insertions: 2,740 image level polyp insertions produced with the released Polyp-Gen SD2 inpainting weights, used as the image level control arm.ldpv_pseudo_mask_tracks: 804 SAM2 propagated pseudo mask tracks over LDPolypVideo clips (16 frames each), with a per clip quality score. Masks and manifests only, no source frames. These are pseudo labels used as training scaffolding, never as evaluation ground truth.
results/
Evaluation outputs, metric json files, contact sheets, and sample videos for every experiment referenced in the research notes.
Provenance and licensing, please read
The generative models were trained on LDPolypVideo clips, and the judge was fine tuned on Kvasir-SEG. Neither source dataset is redistributed here, and neither is any encoded form of it: the VAE latent caches were deliberately excluded because they can be decoded back to approximations of the original frames.
Source datasets carry their own terms and must be obtained from their original providers:
- LDPolypVideo: github.com/dashishi/LDPolypVideo-Benchmark, no explicit licence, cite the MICCAI 2021 paper.
- Kvasir-SEG and Hyper-Kvasir: datasets.simula.no, research and education use, cite the source papers.
- REAL-Colon: figshare, CC BY 4.0.
Third party model weights used during the work (EndoGen, Polyp-Gen, SAM2, LlamaGen, Stable Diffusion VAE) are not included; obtain them from their original authors under their own licences.
The synthetic frames here are model outputs, but the models were trained on real de-identified patient colonoscopy video. Memorisation of training frames is unlikely but has not been formally audited. Treat this repository as research material.
Known limitations
- Small lesions are the main failure mode: adherence Dice is about 0.50 on the larger half of lesions and about 0.30 on the smaller half.
- Trained on 714 clips, so the model overfits after roughly 20k steps.
- 256px was necessary: at 128px, the resolution most of this literature reports at, lesion level spatial control does not work at all.
- All downstream numbers are single dataset (LDPolypVideo detection) and use three seeds at most.