Lelonthecodeur/every-language-dataset-v3
Every Language Dataset V3 Next-generation synthetic multilingual dataset. Size Total: 25,000,000 Train: 24,000,000 Validation: 500,000 Test: 500,000 Diversity Human-language catalog: 171 language codes. Programming languages: 50. Task families: conversation question answering reasoning logic arithmetic translation summarization explanation code generation code explanation debugging Format Parquet + ZSTD. Generation… See the full description on the dataset page: https://huggingface.co/datasets/Lelonthecodeur/every-language-dataset-v3.
This repository belongs to Lelonthecodeur on Hugging Face.
CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.
