- webshart format, indexed with aspect bucketing metadata - four captions per sample - simple VLM output - verbose caption - a concise caption produced by passing the verbose caption through the model again to remove speculation and flowery language - an ideogram4 style JSON schema