|
Download DATA_SOURCES.md from dawncr0w/KorByte-128K: direct link, hf CLI and curl.
- Browser
- Download file 1.33 kB
-
https://huggingface.co/dawncr0w/KorByte-128K/resolve/main/DATA_SOURCES.md
- Command line
-
hf download hf://dawncr0w/KorByte-128K/DATA_SOURCES.md
-
curl -L -o DATA_SOURCES.md https://huggingface.co/dawncr0w/KorByte-128K/resolve/main/DATA_SOURCES.md
1.33 kB
Training data sources
The tokenizer was trained only on the deterministic public-corpus slices below. Raw training text is not redistributed in this repository.
| Source | Config | Revision | License noted by source | Accepted characters |
|---|---|---|---|---|
| HuggingFaceFW/fineweb-2 | kor_Hang |
af9c13333eb981300149d5ca60a8e9d659b276b9 |
ODC-By-1.0 |
500,000,045 |
| wikimedia/wikipedia | 20231101.ko |
b04c8d1ceb2f5cd4588862100d08de323dccfbaa |
CC-BY-SA-3.0-and-GFDL |
150,000,192 |
| wikimedia/wikipedia | 20231101.en |
b04c8d1ceb2f5cd4588862100d08de323dccfbaa |
CC-BY-SA-3.0-and-GFDL |
50,000,499 |
- Shuffle seed:
20260804 - Corpus SHA-256:
85c80c8fe99cfe9aab3c21b7cf189cdcbf78834c714aa1c00fb2b7406ed4c0ca - Accepted characters:
700,000,736 - Accepted lines:
6,064,614 - Filtering: control-character removal, obvious email/URL/long-number redaction, script-ratio filtering, and exact-line deduplication
Users remain responsible for reviewing the original dataset cards and terms. The Apache-2.0 license in this repository applies to the released tokenizer artifact and project code; it does not relicense source datasets.