Institutional Books - Enriched Text: A customizable multilingual open-source pipeline for denoising, deduplicating, and annotating OCR text at scale Paper • 2608.19026 • Published Aug 19 • 3
Institutional Books - Visual Elements: An open-source pipeline for extracting, classifying, deduplicating, and captioning visual elements from digital book collections Paper • 2608.18957 • Published Aug 19 • 1
Institutional Books Collection A growing corpus of public domain books from library collections, seeded by Harvard Library. • 12 items • Updated Aug 21 • 10
Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers Paper • 2608.18972 • Published Aug 19 • 13
Institutional Newspapers Collection A growing corpus of newspapers, parsed and optimized for computational access. • 6 items • Updated 20 days ago • 10