datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TowerOfBabel
About
A synthetic multilingual translation dataset generated with GPT-5.5 XHigh.
The goal is to teach AI models accurate translation from English to target languages — ranging from well-documented languages like Spanish and French, to lower-resource languages like Swahili and Yoruba.
Size
273,980 examples across 61 languages.
Quality
Each example uses complex sentence structures including subordinate clauses and conditionals, making the data richer… See the full description on the dataset page: https://huggingface.co/datasets/8BitStudio/TowerOfBabel.Rocket3B_8bit_r32_alph32_batch4_ResponsesRocket3B_8bit_r32_alph32_eval8BitFinetune
