Extract and chunk text/images from PDFs for dataset creation
Live pretraining progress for OpenEuroLLM 9B