KorByte-128K / DATA_SOURCES.md
DongHyeok-Seo
Release KorByte-128K v2 tokenizer
5a98e33
|
Raw History Blame Contribute Delete
1.33 kB

Training data sources

The tokenizer was trained only on the deterministic public-corpus slices below. Raw training text is not redistributed in this repository.

Source Config Revision License noted by source Accepted characters
HuggingFaceFW/fineweb-2 kor_Hang af9c13333eb981300149d5ca60a8e9d659b276b9 ODC-By-1.0 500,000,045
wikimedia/wikipedia 20231101.ko b04c8d1ceb2f5cd4588862100d08de323dccfbaa CC-BY-SA-3.0-and-GFDL 150,000,192
wikimedia/wikipedia 20231101.en b04c8d1ceb2f5cd4588862100d08de323dccfbaa CC-BY-SA-3.0-and-GFDL 50,000,499
  • Shuffle seed: 20260804
  • Corpus SHA-256: 85c80c8fe99cfe9aab3c21b7cf189cdcbf78834c714aa1c00fb2b7406ed4c0ca
  • Accepted characters: 700,000,736
  • Accepted lines: 6,064,614
  • Filtering: control-character removal, obvious email/URL/long-number redaction, script-ratio filtering, and exact-line deduplication

Users remain responsible for reviewing the original dataset cards and terms. The Apache-2.0 license in this repository applies to the released tokenizer artifact and project code; it does not relicense source datasets.