|
Download paper_notes.md from emilyusmith/text-image-retrieval-dev: direct link, hf CLI and curl.
- Browser
- Download file 3.12 kB
-
https://huggingface.co/emilyusmith/text-image-retrieval-dev/resolve/main/paper_notes.md
- Command line
-
hf download hf://emilyusmith/text-image-retrieval-dev/paper_notes.md
-
curl -L -o paper_notes.md https://huggingface.co/emilyusmith/text-image-retrieval-dev/resolve/main/paper_notes.md
3.12 kB
| # Text Image Retrieval: Research Notes | |
| ## Status | |
| Working note / experiment plan. No completed benchmark results are claimed here. | |
| ## 1. Scope and motivation | |
| This is a working plan for studying alignment quality between image and text representations; it is not a results paper. The central question is whether the proposed change improves the target behavior under a matched training and evaluation budget. The note deliberately separates hypotheses from observations so that future results can be added without rewriting the rationale. | |
| ## 2. Context | |
| Research on text image retrieval often mixes improvements from architecture, data scale, preprocessing, and compute. A useful comparison therefore needs controlled baselines and explicit reporting of resource use. For this topic, the main confound is that near-duplicate captions and dataset overlap can overstate generalization. | |
| ## 3. Working hypothesis | |
| A focused change to the representation or interaction mechanism may improve Recall@1 without increasing deployment cost disproportionately. The hypothesis should be rejected if gains disappear after matching parameter count, data exposure, or tuning budget. | |
| ## 4. Proposed approach | |
| The first implementation should keep modality-specific preprocessing simple, project inputs into a shared representation space, and isolate the new component behind a small interface. Baselines should include a comparable model without the component and a stronger off-the-shelf reference. Any optimization should be applied to all systems, not only the proposed one. | |
| ## 5. Evaluation plan | |
| | Dataset | Role | Primary measure | | |
| |---|---|---| | |
| | Flickr30k | primary evaluation | Recall@1 | | |
| | MS COCO Captions | transfer / robustness | Recall@5 | | |
| | Winoground | transfer / robustness | median rank | | |
| Planned comparisons include a matched-capacity baseline, an ablation that removes the proposed component, and an out-of-domain transfer check. Default training values for the first controlled run are learning rate `0.0003`, batch size `24`, and `5` independent seeds. These are planning values, not claims about a finished experiment. | |
| ## 6. Reproducibility checklist | |
| - Separate model selection from final evaluation. | |
| - Run at least one out-of-domain test. | |
| - Track failed runs as well as successful runs. | |
| - Document every exclusion rule. | |
| ## 7. Failure modes and responsible use | |
| The analysis should report subgroup and category-level failures instead of relying only on a single aggregate score. Particular attention is needed because near-duplicate captions and dataset overlap can overstate generalization. No production use is recommended without task-specific validation, data review, and an assessment of privacy and bias. | |
| ## 8. Open questions | |
| - Where does the method fail on compositional or out-of-domain examples? | |
| - Can a simpler baseline recover the same gain with more careful tuning? | |
| - Which gain survives when the compute budget is matched? | |
| ## References | |
| [1] Radford et al., CLIP, 2021. | |
| [2] Li et al., BLIP, 2022. | |
| [3] Thrush et al., Winoground, 2022. | |