Institutional Books - Enriched Text: A customizable multilingual open-source pipeline for denoising, deduplicating, and annotating OCR text at scale Paper • 2608.19026 • Published Aug 19 • 3
Institutional Books - Visual Elements: An open-source pipeline for extracting, classifying, deduplicating, and captioning visual elements from digital book collections Paper • 2608.18957 • Published Aug 19 • 1
Institutional Books Collection A growing corpus of public domain books from library collections, seeded by Harvard Library. • 12 items • Updated Aug 21 • 10
Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers Paper • 2608.18972 • Published Aug 19 • 13
Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers Paper • 2608.18972 • Published Aug 19 • 13
Institutional Newspapers Collection A growing corpus of newspapers, parsed and optimized for computational access. • 6 items • Updated 17 days ago • 9
institutional/institutional-newspapers-crop-classifier-text-model2vec Text Classification • 32.4M • Updated 29 days ago • 35 • 2
Institutional Books - Visual Elements: An open-source pipeline for extracting, classifying, deduplicating, and captioning visual elements from digital book collections Paper • 2608.18957 • Published Aug 19 • 1
Institutional Books - Enriched Text: A customizable multilingual open-source pipeline for denoising, deduplicating, and annotating OCR text at scale Paper • 2608.19026 • Published Aug 19 • 3
Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers Paper • 2608.18972 • Published Aug 19 • 13
Institutional Books Collection A growing corpus of public domain books from library collections, seeded by Harvard Library. • 12 items • Updated Aug 21 • 10
Institutional Newspapers Collection A growing corpus of newspapers, parsed and optimized for computational access. • 6 items • Updated 17 days ago • 9