CoolFace
Datasetpublic

8BitStudio/TowerOfBabel

About A synthetic multilingual translation dataset generated with GPT-5.5 XHigh. The goal is to teach AI models accurate translation from English to target languages — ranging from well-documented languages like Spanish and French, to lower-resource languages like Swahili and Yoruba. Size 273,980 examples across 61 languages. Quality Each example uses complex sentence structures including subordinate clauses and conditionals, making the data richer… See the full description on the dataset page: https://huggingface.co/datasets/8BitStudio/TowerOfBabel.

sourceHugging Facecc0-1.0updated 2mo agoView on Hugging Face
0likes102downloads
Dataset Card

About

A synthetic multilingual translation dataset generated with GPT-5.5 XHigh. The goal is to teach AI models accurate translation from English to target languages — ranging from well-documented languages like Spanish and French, to lower-resource languages like Swahili and Yoruba.

Size

273,980 examples across 61 languages.

Quality

Each example uses complex sentence structures including subordinate clauses and conditionals, making the data richer and more generalizable than simple phrase datasets. All translations were verified using multi-agent pipelines to ensure quality across all language pairs. Generation Synthetically generated and verified using GPT-5.5 XHigh sub-agents. some examples may be incorrect due to this being such a large synthetic dataset