CoolFace
Datasetpublic

zh-tw-llm-dv/zh-tw-pythia-ta8000-v1-it1-sg-002

zh-tw-pythia-ta8000-v1-it1-sg-002 This dataset is a part of the zh-tw-llm project. Tokenizer: zh-tw-pythia-tokenizer-a8000-v1 Built with: sharegpt Rows: train 8054, test 83 Max length: 2048 Full config:{"build_with": ["sharegpt"], "preview_length": 512, "sharegpt_settings": {"source_dataset": "zetavg/ShareGPT-Processed", "train_on_inputs": false, "languages": [{"en": 0.3}, {"zh": 0.2}, "zh_Hant"], "rows_limit": 10000, "test_size": 0.01, "test_split_seed": 42, "test_rows_limit":… See the full description on the dataset page: https://huggingface.co/datasets/zh-tw-llm-dv/zh-tw-pythia-ta8000-v1-it1-sg-002.

sourceHugging Faceupdated 3y agoView on Hugging Face
0likes21downloads

zh-tw-llm-dv/zh-tw-pythia-ta8000-v1-it1-sg-002 · main · files are served by the source, never re-hosted here