CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gsarti /flores_101One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource languages, consider only restricted domains, or are low quality because they are constructed using semi-automatic procedures. In this work, we introduce the FLORES evaluation benchmark, consisting of 3001 sentences extracted from English Wikipedia and covering a variety of different topics and domains. These sentences have been translated in 101 languages by professional translators through a carefully controlled process. The resulting dataset enables better assessment of model quality on the long tail of low-resource languages, including the evaluation of many-to-many multilingual translation systems, as all translations are multilingually aligned. By publicly releasing such a high-quality and high-coverage dataset, we hope to foster progress in the machine translation community and beyond.tabulartext-generation100K<n<1M33 likes27k downloads4y agoHugging Face02gsarti /clean_mc4_itA thoroughly cleaned version of the Italian portion of the multilingual colossal, cleaned version of Common Crawl's web crawl corpus (mC4) by AllenAI. Based on Common Crawl dataset: "https://commoncrawl.org". This is the processed version of Google's mC4 dataset by AllenAI, with further cleaning detailed in the repository README file.text-generation100M<n<1B18 likes2.1k downloads2y agoHugging Face03gsarti /change_itThe CHANGE-IT dataset contains approximately 152,000 article-headline pairs, collected from two Italian newspapers situated at opposite ends of the political spectrum, namely la Repubblica (left) and Il Giornale (right), with the two newspapers equally represented. The dataset has been used in the context of the CHANGE-IT task (https://sites.google.com/view/change-it) during the Evalita 2020 evaluation campaign (http://www.evalita.it/2020). CHANGE-IT is a generation task for Italian – more specifically, a style transfer task for headlines of Italian newspapers. Given a (collection of) headlines from one newspaper, namely Il Giornale (G) or La Repubblica (R), it challenges automatic systems to change all G-headlines to headlines in style R, and all R-headlines to headlines in style G. Although the task only concerns headline change, the dataset comprehends both the headlines as well as their respective full articles.summarization1 likes126 downloads4y agoHugging Face04gsarti /wmt_vatThe Variance-Aware Machine Translation corpus contains 70 small and discriminative test sets for machine translation (MT) evaluation called variance-aware test sets (VAT), covering 35 translation directions from WMT16 to WMT20 competitions. VAT is automatically created by a novel variance-aware filtering method that filters the indiscriminative test instances of the current MT benchmark without any human labor. Experimental results show that VAT outperforms the original WMT benchmark in terms of the correlation with human judgment across mainstream language pairs and test sets. Further analysis on the properties of VAT reveals the challenging linguistic features (e.g., translation of low-frequency words and proper nouns) for the competitive MT systems, providing guidance for constructing future MT test sets.text-generation7 likes81 downloads4y agoHugging Face05GSAI-ML /ReFusion ReFusion Dataset Summary This dataset is the training corpus used for ReFusion, as described in our paper. It comprises approximately 3.7 million high-quality instruction tuning samples consolidated from several state-of-the-art open-source datasets. The data covers diverse domains including mathematics, coding, and general instruction following. Composition & Sources The dataset is constructed from the following sources: MAmmoTH OpenMathInstruct-2 (1M… See the full description on the dataset page: https://huggingface.co/datasets/GSAI-ML/ReFusion.texttext-generation1M<n<10M4 likes62 downloads9mo agoHugging Face06gsarti /eureka-rebusgated Dataset Card for EurekaRebus Last data update: February 14th, 2026. Refer to the changelog for a list of revisions that can be loaded with the revision parameter in load_dataset. Dataset Summary This dataset contains the original collection of over 200k first passes and solution for Italian rebuses published in various Italian magazines dating back to 1869. The original data are hosted in the Eureka5 platform of the Associazione Culturale "Biblioteca Enigmistica Italiana… See the full description on the dataset page: https://huggingface.co/datasets/gsarti/eureka-rebus.tabulartext-generation100K<n<1M2 likes22 downloads7mo agoHugging Face07lianghsun /tw-gsat-chatgated tw-gsat — 台灣學測 SFT 資料集(國文 + 社會科,110–115 學年度) 本資料集為合成 SFT 訓練資料,涵蓋台灣學測國文與社會科選擇題。 子集 Subset 筆數 說明 chinese 152 學測國文(110–115) society 237 學測社會(110–115) default (merged) 389 合併版 chinese_v2 152 國文 v2——結構化 think + 豐富 output + \boxed{X} society_v2 237 社會 v2——同上 merged_v2 389 v2 合併版 Schema 與 lianghsun/secret-chat 相同格式: unique_id, messages, turn, question, think, answer, tools, system_prompt, lang_question, lang_answer, lang_think, tags… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-gsat-chat.tabulartext-generation1K<n<10K0 likes8 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.