Audio-Visual Scene-Aware Dialog, DSTC7

Answer a question about a video in words, given the dialog so far. Six modalities are attended jointly -- the question, four VGG19 conv grids of the video, and the audio -- and an LSTM decoder writes the answer. Built on the attention layer of Factor Graph Attention (CVPR 2019); a reimplementation of A Simple Baseline for Audio-Visual Scene-Aware Dialog (CVPR 2019).

Code: github.com/idansc/fga.

Results

DSTC7 test set, scored against six references with the organisers' dstc7avsd_eval metrics:

this model organisers' I3D baseline
BLEU-1 0.651 0.621
BLEU-4 0.348 0.305
METEOR 0.233 0.217
ROUGE_L 0.510 0.481
CIDEr 0.887 0.733

These are not comparable to Table 1 of the paper, which reports CIDEr against a single reference on the validation set. Six references and one reference put CIDEr on different scales; the baseline column above is the anchor for this one.

Seeds, and which checkpoint this is

The published weights widen the decoder's output projection to 1024 from the released configuration's 128; everything else follows the original. Seeds:

projection CIDEr per seed
128 (released config, 8.36M params) 0.812, 0.813, 0.814
512 0.836, 0.843, 0.847, 0.884
1024 (this repo) 0.839, 0.866, 0.887

The published checkpoint is the best seed at 1024. Read the spread before the best number: seed-to-seed variation at these widths runs to ~0.05 CIDEr.

The projection width is not a detail

At the released width of 128, zeroing the video features costs nothing measurable -- the model answers from text alone and no training metric notices. At 512, zeroing the video costs about 0.12 CIDEr, and a text-only model becomes unstable across seeds (0.809 vs 0.659). The narrow projection after the attended representation was the reason the audio-visual signal could not express itself; widening the layer where the reduction happens is what let the video matter.

Faithful configuration

For the released configuration exactly -- projection 128, 8,359,107 parameters, matching the paper's reported count -- the numbers are CIDEr 0.813 +/- 0.001 over three seeds, still above the challenge baseline on every metric. Word embeddings of 128 with a separate table per consumer, a single-layer question LSTM of 256, a hierarchical history encoder reading each question-answer pair before reading the turns in order, a 256-wide decoder. Adam at 1e-3, batch 64, dropout 0.5, He initialization except recurrent weights which get Xavier, beam width 3.

Video features are four equally-spaced frames per clip through VGG19's last conv layer, 7x7 = 49 regions of 512 -- what the original release calls i3d_rgb_vgg19_4. Audio is VGGish, 32 steps sampled evenly across the clip.

Files

pred_q1024_s3.json holds the generated answer for all 1,710 test questions, keyed {video}_{turn}.

Citation

@inproceedings{schwartz2019simple,
  title={A Simple Baseline for Audio-Visual Scene-Aware Dialog},
  author={Schwartz, Idan and Schwing, Alexander G and Hazan, Tamir},
  booktitle={Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition},
  year={2019}
}

@inproceedings{schwartz2019factor,
  title={Factor graph attention},
  author={Schwartz, Idan and Yu, Seunghak and Hazan, Tamir and Schwing, Alexander G},
  booktitle={Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition},
  pages={2039--2048},
  year={2019}
}

@article{alamri2019audio,
  title={Audio Visual Scene-Aware Dialog},
  author={Alamri, Huda and Cartillier, Vincent and Das, Abhishek and Wang, Jue and
          Cherian, Anoop and Essa, Irfan and Batra, Dhruv and Marks, Tim K and
          Hori, Chiori and Anderson, Peter and Lee, Stefan and Parikh, Devi},
  journal={CVPR},
  year={2019}
}
Downloads last month
-
Safetensors
Model size
13.6M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Papers for Idan/fga-avsd