datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
flores_101One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the
lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource
languages, consider only restricted domains, or are low quality because they are constructed using
semi-automatic procedures. In this work, we introduce the FLORES evaluation benchmark, consisting of 3001
sentences extracted from English Wikipedia and covering a variety of different topics and domains.
These sentences have been translated in 101 languages by professional translators through a carefully
controlled process. The resulting dataset enables better assessment of model quality on the long tail of
low-resource languages, including the evaluation of many-to-many multilingual translation systems, as all
translations are multilingually aligned. By publicly releasing such a high-quality and high-coverage dataset,
we hope to foster progress in the machine translation community and beyond.eureka-rebus
Dataset Card for EurekaRebus
Last data update: February 14th, 2026. Refer to the changelog for a list of revisions that can be loaded with the revision parameter in load_dataset.
Dataset Summary
This dataset contains the original collection of over 200k first passes and solution for Italian rebuses published in various Italian magazines dating back to 1869. The original data are hosted in the Eureka5 platform of the Associazione Culturale "Biblioteca Enigmistica Italiana… See the full description on the dataset page: https://huggingface.co/datasets/gsarti/eureka-rebus.tw-gsat-chat
tw-gsat — 台灣學測 SFT 資料集(國文 + 社會科,110–115 學年度)
本資料集為合成 SFT 訓練資料,涵蓋台灣學測國文與社會科選擇題。
子集
Subset
筆數
說明
chinese
152
學測國文(110–115)
society
237
學測社會(110–115)
default (merged)
389
合併版
chinese_v2
152
國文 v2——結構化 think + 豐富 output + \boxed{X}
society_v2
237
社會 v2——同上
merged_v2
389
v2 合併版
Schema
與 lianghsun/secret-chat 相同格式:
unique_id, messages, turn, question, think, answer, tools,
system_prompt, lang_question, lang_answer, lang_think,
tags… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-gsat-chat.
