762 GB
898 files
Updated 9 days ago
Name
Size
2013_20
2013_48
2014_10
2014_15
2014_23
2014_35
2014_41
2014_42
2014_49
2014_52
2015_06
2015_11
2015_14
2015_18
2015_22
2015_27
2015_32
2015_35
2015_40
2015_48
2016_07
2016_18
2016_22
2016_26
2016_30
2016_36
2016_40
2016_44
2016_50
2017_04
2017_09
2017_13
2017_17
2017_22
2017_26
2017_30
2017_34
2017_39
2017_43
2017_47
2017_51
2018_05
2018_09
2018_13
2018_17
2018_22
2018_26
2018_30
2018_34
2018_39
2018_43
2018_47
2018_51
2019_04
2019_09
2019_13
2019_18
2019_22
2019_26
2019_30
2019_35
2019_43
2019_47
2019_51
2020_05
2020_10
2020_16
2020_24
2020_29
2020_34
2020_40
2020_45
2020_50
2021_04
2021_10
2021_17
2021_21
2021_25
2021_31
2021_39
2021_43
2021_49
2022_05
2022_21
2022_27
2022_33
2022_40
2022_49
2023_06
2023_14
2023_23
2023_40
2023_50
2024_10
2024_18
2024_22
2024_26
2024_30
2024_33
.gitattributes2.95 kB
xet
README.md

Traditional Chinese C4

Dataset Summary

Data obtained from 2013~2025 Common Crawl. Original source from here.

Downloaded and processed using code based on another project attempting to recreate the C4 dataset.

The resultant dataset contains both simplified and traditional Chinese, which could be found here. It was then filtered using a modified list of simplified Chinese characters to obtain this traditional Chinese dataset.

Unfortunately, I don't have enough funding to run a deduplication across all dumps at the moment. I will do it in the future (hopefully).

Acknowoledgement

Special thanks to Eons Data Communications Limited and Votee AI for the compute resources to download and process the Common Crawl Dumps.

Deduplication carried out using compute resource offered under the category of General Projects by Research Insituite for Information Technology, Kyushu University. This work was also partly achieved through the use of SQUID at D3 Center, The University of Osaka.

Total size
762 GB
Files
898
Last updated
Jul 28
Pre-warmed CDN
US EU US EU

Contributors