CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gsarti /flores_101One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource languages, consider only restricted domains, or are low quality because they are constructed using semi-automatic procedures. In this work, we introduce the FLORES evaluation benchmark, consisting of 3001 sentences extracted from English Wikipedia and covering a variety of different topics and domains. These sentences have been translated in 101 languages by professional translators through a carefully controlled process. The resulting dataset enables better assessment of model quality on the long tail of low-resource languages, including the evaluation of many-to-many multilingual translation systems, as all translations are multilingually aligned. By publicly releasing such a high-quality and high-coverage dataset, we hope to foster progress in the machine translation community and beyond.tabulartext-generation100K<n<1M33 likes27k downloads4y agoHugging Face02gsarti /clean_mc4_itA thoroughly cleaned version of the Italian portion of the multilingual colossal, cleaned version of Common Crawl's web crawl corpus (mC4) by AllenAI. Based on Common Crawl dataset: "https://commoncrawl.org". This is the processed version of Google's mC4 dataset by AllenAI, with further cleaning detailed in the repository README file.text-generation100M<n<1B18 likes2.1k downloads2y agoHugging Face03gsarti /qe4pe Quality Estimation for Post-Editing (QE4PE) For more details on QE4PE, see our paper and our Github repository Gabriele Sarti • Vilém Zouhar • Grzegorz Chrupała • Ana Guerberof Arenas • Malvina Nissim • Arianna Bisazza Word-level quality estimation (QE) detects erroneous spans in machine translations, which can direct and facilitate human post-editing. While the accuracy of word-level QE systems has been assessed extensively, their usability and downstream influence on the… See the full description on the dataset page: https://huggingface.co/datasets/gsarti/qe4pe.tabulartranslation10K<n<100K4 likes455 downloads1y agoHugging Face04gsaltintas /controlled-datatext10M<n<100M1 likes328 downloads8mo agoHugging Face05TCabbage /gsat-vocab-sentences-tts GSAT Vocabulary TTS Audio Text-to-speech audio files for GSAT (General Scholastic Ability Test) English vocabulary. Structure audio/ - MP3 audio files organized by hash prefix (e.g., audio/ab/abcd1234....mp3) index.jsonl - Index file mapping hashes to text and TTS engine used Engines Kokoro (af_heart voice) - Used for lemmas (single words/phrases) Supertonic (M1 voice) - Used for example sentences Audio Format Format: MP3 Sample rate: 24kHz… See the full description on the dataset page: https://huggingface.co/datasets/TCabbage/gsat-vocab-sentences-tts.audiotext-to-speech10K<n<100K0 likes323 downloads9mo agoHugging Face06gist-sparse-attention /GSA-PT-Qwen2-7B-Instruct-chunk4-chunk4-data GSA-PT-Qwen2-7B-Instruct-chunk4-chunk4-data This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk4-chunk4. Each sample is tokenized and formatted with GSA gist tokens for continue pretraining. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Model yuzhenm/GSA-PT-Qwen2-7B-Instruct-chunk4-chunk4 — model trained on this dataset tabular10K<n<100K0 likes280 downloads6mo agoHugging Face07gist-sparse-attention /GSA-PT-Llama-3.2-1B-chunk8-data GSA-PT-Llama-3.2-1B-chunk8-data This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models with chunk size chunk8. Each sample is tokenized and formatted with GSA gist tokens for continued pretraining. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Models yuzhenm/GSA-PT-Llama-3.2-1B-chunk8 — model trained on this dataset 0 likes253 downloads6mo agoHugging Face08gsarch /vigorl_datasets ViGoRL Datasets This repository contains the official datasets associated with the paper "Grounded Reinforcement Learning for Visual Reasoning (ViGoRL)", by Gabriel Sarch, Snigdha Saha, Naitik Khandelwal, Ayush Jain, Michael J. Tarr, Aviral Kumar, and Katerina Fragkiadaki. Dataset Overview These datasets are designed for training and evaluating visually grounded vision-language models (VLMs). Datasets are organized by the visual reasoning tasks described in the ViGoRL… See the full description on the dataset page: https://huggingface.co/datasets/gsarch/vigorl_datasets.imagevisual-question-answering100K<n<1M1 likes203 downloads1y agoHugging Face09OALL /details_GSAI-ML__LLaDA-8B-Instruct_v2 Dataset Card for Evaluation run of GSAI-ML/LLaDA-8B-Instruct Dataset automatically created during the evaluation run of model GSAI-ML/LLaDA-8B-Instruct. The dataset is composed of 116 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_GSAI-ML__LLaDA-8B-Instruct_v2.text100K<n<1M0 likes185 downloads2y agoHugging Face10gsarti /mt_genevalThe MT-GenEval benchmark evaluates gender translation accuracy on English -> {Arabic, French, German, Hindi, Italian, Portuguese, Russian, Spanish}. The dataset contains individual sentences with annotations on the gendered target words, and contrastive original-invertend translations with additional preceding context.texttranslation10K<n<100K7 likes174 downloads4y agoHugging Face11gist-sparse-attention /GSA-PT-Qwen2-7B-Instruct-chunk16-data GSA-PT-Qwen2-7B-Instruct-chunk16-data This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk16. Each sample is tokenized and formatted with GSA gist tokens for continue pretraining. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Model yuzhenm/GSA-PT-Qwen2-7B-Instruct-chunk16 — model trained on this dataset tabular10K<n<100K0 likes164 downloads6mo agoHugging Face12gist-sparse-attention /GSA-FT-Qwen2-7B-Instruct-chunk4-chunk4-data GSA-FT-Qwen2-7B-Instruct-chunk4-chunk4-data This is the supervised fine-tuning dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk4-chunk4. Each sample is tokenized and formatted with GSA gist tokens for supervised fine-tuning. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Model yuzhenm/GSA-FT-Qwen2-7B-Instruct-chunk4-chunk4 — model trained on this dataset 0 likes147 downloads6mo agoHugging Face13gsarti /iwslt2017_contextThe IWSLT 2017 Multilingual Task addresses text translation, including zero-shot translation, with a single MT system across all directions including English, German, Dutch, Italian and Romanian. As unofficial task, conventional bilingual text translation is offered between English and Arabic, French, Japanese, Chinese, German and Korean.tabulartranslation1M<n<10M1 likes141 downloads3y agoHugging Face14gist-sparse-attention /GSA-PT-Llama-3.2-1B-chunk16-data GSA-PT-Llama-3.2-1B-chunk16-data This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models with chunk size chunk16. Each sample is tokenized and formatted with GSA gist tokens for continued pretraining. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Models yuzhenm/GSA-PT-Llama-3.2-1B-chunk16 — model trained on this dataset 0 likes139 downloads6mo agoHugging Face15gsarti /change_itThe CHANGE-IT dataset contains approximately 152,000 article-headline pairs, collected from two Italian newspapers situated at opposite ends of the political spectrum, namely la Repubblica (left) and Il Giornale (right), with the two newspapers equally represented. The dataset has been used in the context of the CHANGE-IT task (https://sites.google.com/view/change-it) during the Evalita 2020 evaluation campaign (http://www.evalita.it/2020). CHANGE-IT is a generation task for Italian – more specifically, a style transfer task for headlines of Italian newspapers. Given a (collection of) headlines from one newspaper, namely Il Giornale (G) or La Repubblica (R), it challenges automatic systems to change all G-headlines to headlines in style R, and all R-headlines to headlines in style G. Although the task only concerns headline change, the dataset comprehends both the headlines as well as their respective full articles.summarization1 likes126 downloads4y agoHugging Face16gist-sparse-attention /GSA-FT-Qwen2-7B-Instruct-chunk8-chunk4-data GSA-FT-Qwen2-7B-Instruct-chunk8-chunk4-data This is the supervised fine-tuning dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk8-chunk4. Each sample is tokenized and formatted with GSA gist tokens for supervised fine-tuning. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Model yuzhenm/GSA-FT-Qwen2-7B-Instruct-chunk8-chunk4 — model trained on this dataset 0 likes125 downloads6mo agoHugging Face17gist-sparse-attention /GSA-PT-Qwen2-7B-Instruct-chunk32-data GSA-PT-Qwen2-7B-Instruct-chunk32-data This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk32. Each sample is tokenized and formatted with GSA gist tokens for continue pretraining. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Model yuzhenm/GSA-PT-Qwen2-7B-Instruct-chunk32 — model trained on this dataset 0 likes115 downloads6mo agoHugging Face18gsarch /countqa_lite gsarch/countqa_lite A deterministic lite evaluation subset of Jayant-Sravan/CountQA. Source revision: f92cc6fe46542c61e2916e3d2ae9a911e2216b1a Source split: test Sampling seed: 43 Output rows: 500 Schema: unchanged from the upstream dataset CountQA is sampled at the QA-pair level. Each output row retains the original schema and contains one-element questions and answers lists, so lmms-eval's existing countqa_process_docs produces exactly 500 prompts. Generated by… See the full description on the dataset page: https://huggingface.co/datasets/gsarch/countqa_lite.imagevisual-question-answeringn<1K0 likes111 downloads2mo agoHugging Face19Gsatlwf /TCGA_Cancer0 likes98 downloads3y agoHugging Face20gsarti /magpieThe MAGPIE corpus is a large sense-annotated corpus of potentially idiomatic expressions (PIEs), based on the British National Corpus (BNC). Potentially idiomatic expressions are like idiomatic expressions, but the term also covers literal uses of idiomatic expressions, such as 'I leave work at the end of the day.' for the idiom 'at the end of the day'. This version of the dataset reflects the filtered subset used by Dankers et al. (2022) in their investigation on how PIEs are represented by NMT models. Authors use 37k samples annotated as fully figurative or literal, for 1482 idioms that contain nouns, numerals or adjectives that are colours (which they refer to as keywords). Because idioms show syntactic and morphological variability, the focus is mostly put on nouns. PIEs and their context are separated using the original corpus’s word-level annotations.texttext-classification10K<n<100K7 likes85 downloads4y agoHugging Face21gsarti /wmt_vatThe Variance-Aware Machine Translation corpus contains 70 small and discriminative test sets for machine translation (MT) evaluation called variance-aware test sets (VAT), covering 35 translation directions from WMT16 to WMT20 competitions. VAT is automatically created by a novel variance-aware filtering method that filters the indiscriminative test instances of the current MT benchmark without any human labor. Experimental results show that VAT outperforms the original WMT benchmark in terms of the correlation with human judgment across mainstream language pairs and test sets. Further analysis on the properties of VAT reveals the challenging linguistic features (e.g., translation of low-frequency words and proper nouns) for the competitive MT systems, providing guidance for constructing future MT test sets.text-generation7 likes81 downloads4y agoHugging Face22gsarch /ScreenSpot-Pro-Lite ScreenSpot-Pro-Lite A fixed 500-example representative/challenging subset of the 1,581-example ScreenSpot-Pro benchmark. Selection Sampling is proportional over the cross-product of platform, application, and UI type, so all 26 applications remain represented and the original icon/text mix is retained. Within every stratum, 80% is deterministic seeded sampling and 20% is a hard tier. Hardness combines failure rate and disagreement across five full-run anchor… See the full description on the dataset page: https://huggingface.co/datasets/gsarch/ScreenSpot-Pro-Lite.imageimage-text-to-textn<1K0 likes78 downloads1mo agoHugging Face23gist-sparse-attention /GSA-PT-Qwen2-7B-Instruct-chunk8-chunk4-data GSA-PT-Qwen2-7B-Instruct-chunk8-chunk4-data This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk8-chunk4. Each sample is tokenized and formatted with GSA gist tokens for continue pretraining. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Model yuzhenm/GSA-PT-Qwen2-7B-Instruct-chunk8-chunk4 — model trained on this dataset tabular10K<n<100K0 likes76 downloads6mo agoHugging Face24gist-sparse-attention /GSA-FT-Qwen2-7B-Instruct-chunk32-data GSA-FT-Qwen2-7B-Instruct-chunk32-data This is the supervised fine-tuning dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk32. Each sample is tokenized and formatted with GSA gist tokens for supervised fine-tuning. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Model yuzhenm/GSA-FT-Qwen2-7B-Instruct-chunk32 — model trained on this dataset tabular10K<n<100K0 likes75 downloads6mo agoHugging Face25gsaltintas /turkihs Dataset Card for Tokenization Robustness TokSuite Benchmark (Turkish Collection) Dataset Description This dataset is part of TokSuite, a comprehensive benchmark designed to measure how different tokenization strategies affect language model performance and robustness. This specific subset contains Turkish language multiple-choice text completion questions with various real-world perturbations that test tokenizer robustness. Curated by: R3 Research Team… See the full description on the dataset page: https://huggingface.co/datasets/gsaltintas/turkihs.multiple-choicen<1K0 likes70 downloads8mo agoHugging Face26gsarti /grote-logs Dataset Card for Dataset Name Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More Information Needed] Paper [optional]: [More Information Needed] Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/gsarti/grote-logs.textn<1K0 likes65 downloads1y agoHugging Face27gist-sparse-attention /GSA-PT-Qwen2-7B-Instruct-chunk8-data GSA-PT-Qwen2-7B-Instruct-chunk8-data This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk8. Each sample is tokenized and formatted with GSA gist tokens for continue pretraining. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Model yuzhenm/GSA-PT-Qwen2-7B-Instruct-chunk8 — model trained on this dataset tabular10K<n<100K0 likes65 downloads6mo agoHugging Face28GSAI-ML /ReFusion ReFusion Dataset Summary This dataset is the training corpus used for ReFusion, as described in our paper. It comprises approximately 3.7 million high-quality instruction tuning samples consolidated from several state-of-the-art open-source datasets. The data covers diverse domains including mathematics, coding, and general instruction following. Composition & Sources The dataset is constructed from the following sources: MAmmoTH OpenMathInstruct-2 (1M… See the full description on the dataset page: https://huggingface.co/datasets/GSAI-ML/ReFusion.texttext-generation1M<n<10M4 likes62 downloads9mo agoHugging Face29gsaltintas /temporal_expressions Dataset Card for Tokenization Robustness A comprehensive evaluation dataset for testing robustness of different tokenization strategies. Dataset Details Dataset Description This dataset evaluates how robust language models are to different tokenization strategies and edge cases. It includes questions with multiple choice answers designed to test various aspects of tokenization handling. Curated by: R3 Funded by [optional]: [More Information Needed] Shared… See the full description on the dataset page: https://huggingface.co/datasets/gsaltintas/temporal_expressions.tabularmultiple-choicen<1K1 likes61 downloads1y agoHugging Face30gist-sparse-attention /GSA-FT-Qwen2-7B-Instruct-chunk16-data GSA-FT-Qwen2-7B-Instruct-chunk16-data This is the supervised fine-tuning dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk16. Each sample is tokenized and formatted with GSA gist tokens for supervised fine-tuning. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Model yuzhenm/GSA-FT-Qwen2-7B-Instruct-chunk16 — model trained on this dataset 0 likes61 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.