310 GB
67,609 files
Updated 22 days ago
Name
Size
2013-20
2013-48
2014-10
2014-15
2014-23
2014-35
2014-41
2014-42
2014-49
2014-52
2015-06
2015-11
2015-14
2015-18
2015-22
2015-27
2015-32
2015-35
2015-40
2015-48
2016-07
2016-18
2016-22
2016-26
2016-30
2016-36
2016-40
2016-44
2016-50
2017-04
2017-09
2017-13
2017-17
2017-22
2017-26
2017-30
2017-34
2017-39
2017-43
2017-51
2018-05
2018-09
2018-13
2018-17
2018-22
2018-26
2018-30
2018-34
2018-39
2018-43
2018-47
2018-51
2019-04
2019-09
2019-13
2019-18
2019-22
2019-26
2019-30
2019-35
2019-39
2019-43
2019-47
2019-51
2020-05
2020-10
2020-16
2020-24
2020-29
2020-34
2020-40
2020-45
2020-50
2021-04
2021-10
2021-17
2021-21
2021-25
2021-31
2021-39
2021-43
2021-49
2022-05
2022-21
2022-27
2022-33
2022-40
2022-49
2023-06
2023-14
2023-23
2023-40
2023-50
2024-10
2024-18
.gitattributes2.31 kB
xet
LICENSE.md19.9 kB
xet
README.md2.17 kB
xet
README.md

❄️FuLG

The FuLG dataset is a comprehensive Romanian language corpus comprising 150 billion tokens, carefully extracted from Common Crawl. This extensive dataset is the result of rigorous filtering and deduplication processes applied to 95 Common Crawl snapshots. The compressed dataset has 289 GB.

For more details, check the arXiv preprint.

How do I download this?

Using 🤗 Datasets
from datasets import load_dataset

# Full dataset
dataset = load_dataset("faur-ai/fulg")

# To load the data from a specific CC snapshot
dataset = load_dataset("faur-ai/fulg", data_dir='2018-05')
Using Git
git clone https://huggingface.co/datasets/faur-ai/fulg

Data Fields

The data have several fields:

  • url: url of the source as a string
  • date_download: date of crawl
  • digest: hash of content
  • length: length of content
  • nlines: number of lines
  • source_domain: domain of document
  • title: title of document
  • raw_content: text content as a string
  • cc_segment: source CommonCrawl segment
  • original_nlines: original number of lines before processing
  • original_length: original length before processing
  • language: language (ro)
  • language_score: score for language

Licensing Information

We are releasing this dataset under the terms of ODC-BY. By using this dataset, you are also bound any license agreements and terms of use of the original data sources.

Bibtex

If you use our dataset, please cite us at:

@misc{fulg150bromaniancorpus,
      title={FuLG: 150B Romanian Corpus for Language Model Pretraining}, 
      author={Vlad-Andrei Bădoiu and Mihai-Valentin Dumitru and Alexandru M. Gherghescu and Alexandru Agache and Costin Raiciu},
      year={2024},
      eprint={2407.13657},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2407.13657}, 
}
Total size
310 GB
Files
67,609
Last updated
Sep 13
Pre-warmed CDN
US EU US EU

Contributors