CoolFace
Datasetpublic

agentlans/fineweb2hq-vs-c4

This dataset includes 5000 rows per language from each of two sources: the higher-quality epfml/FineWeb2-HQ and the lower-quality allenai/c4. The data is split 80/20 into training and test sets. Languages were carefully chosen to ensure balanced representation across both splits: Arabic, Chinese, Czech, Danish, Dutch, French, German, Greek, Hungarian, Indonesian, Italian, Japanese, Persian, Polish, Portuguese, Russian, Spanish, Swedish, Turkish, and Vietnamese.

sourceHugging Faceodc-byupdated 1y agoView on Hugging Face
0likes23downloads
Dataset Card

This dataset includes 5000 rows per language from each of two sources: the higher-quality epfml/FineWeb2-HQ and the lower-quality allenai/c4. The data is split 80/20 into training and test sets.

Languages were carefully chosen to ensure balanced representation across both splits: Arabic, Chinese, Czech, Danish, Dutch, French, German, Greek, Hungarian, Indonesian, Italian, Japanese, Persian, Polish, Portuguese, Russian, Spanish, Swedish, Turkish, and Vietnamese.