File size: 1,630 Bytes
594866b
 
43ac739
 
 
594866b
 
43ac739
594866b
 
43ac739
28fe309
43ac739
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
---
title: README
emoji: πŸ“š
colorFrom: blue
colorTo: indigo
sdk: static
pinned: false
short_description: Phrase-level segmentation and alignment for medieval texts.
---

# ProMeTEXT

**ProMeTEXT** β€” the **Centre for PROcessing MEdieval TEXTs** β€” develops datasets, models and tools for the computational study of medieval and historical texts.

Our work focuses on **phrase-level segmentation**, **multilingual alignment**, and the processing of medieval textual traditions across Romance languages, Latin, and Middle English.

## Resources

- **Aquilign** β€” a multilingual aligner for historical and philological corpora.
- **Aquilign Multilingual Segmenter** β€” a Hugging Face model for phrase-level segmentation of historical texts.
- **Aquilign Explorer** β€” a demo app for demonstrating multilingual alignment workflows.
- **Multilingual Segmentation Dataset** β€” gold-standard segmentation data for medieval prose.
- **Parallel Alignment Corpora** β€” multilingual aligned corpora used for fine-tuning LaBSE and evaluating multilingual alignment across historical textual traditions.
   
## Links

- [GitHub organization](https://github.com/ProMeText)
- [Alignment tool: Aquilign](https://github.com/ProMeText/Aquilign)
- [Demo app: Aquilign Explorer](https://huggingface.co/spaces/ProMeText/aquilign-explorer)
- [Segmentation model: Aquilign Multilingual Segmenter](https://huggingface.co/ProMeText/aquilign-multilingual-segmenter)
- [Segmentation dataset](https://github.com/ProMeText/multilingual-segmentation-dataset)
- [Parallel corpora](https://github.com/ProMeText/parallelium-scriptures-alignment-dataset)