agentlans/fineweb2hq-vs-c4
This dataset includes 5000 rows per language from each of two sources: the higher-quality epfml/FineWeb2-HQ and the lower-quality allenai/c4. The data is split 80/20 into training and test sets. Languages were carefully chosen to ensure balanced representation across both splits: Arabic, Chinese, Czech, Danish, Dutch, French, German, Greek, Hungarian, Indonesian, Italian, Japanese, Persian, Polish, Portuguese, Russian, Spanish, Swedish, Turkish, and Vietnamese.
This dataset includes 5000 rows per language from each of two sources: the higher-quality epfml/FineWeb2-HQ and the lower-quality allenai/c4. The data is split 80/20 into training and test sets.
Languages were carefully chosen to ensure balanced representation across both splits: Arabic, Chinese, Czech, Danish, Dutch, French, German, Greek, Hungarian, Indonesian, Italian, Japanese, Persian, Polish, Portuguese, Russian, Spanish, Swedish, Turkish, and Vietnamese.
