CoolFace
20 results

gsa

gsarti /flores_101One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource languages, consider only restricted domains, or are low quality because they are constructed using semi-automatic procedures. In this work, we introduce the FLORES evaluation benchmark, consisting of 3001 sentences extracted from English Wikipedia and covering a variety of different topics and domains. These sentences have been translated in 101 languages by professional translators through a carefully controlled process. The resulting dataset enables better assessment of model quality on the long tail of low-resource languages, including the evaluation of many-to-many multilingual translation systems, as all translations are multilingually aligned. By publicly releasing such a high-quality and high-coverage dataset, we hope to foster progress in the machine translation community and beyond.tabulartext-generation100K<n<1M33 likes26k downloads4y agoHugging Facegsarti /clean_mc4_itA thoroughly cleaned version of the Italian portion of the multilingual colossal, cleaned version of Common Crawl's web crawl corpus (mC4) by AllenAI. Based on Common Crawl dataset: "https://commoncrawl.org". This is the processed version of Google's mC4 dataset by AllenAI, with further cleaning detailed in the repository README file.text-generation100M<n<1B18 likes2.1k downloads2y agoHugging Facegsarti /qe4pe Quality Estimation for Post-Editing (QE4PE) For more details on QE4PE, see our paper and our Github repository Gabriele Sarti • Vilém Zouhar • Grzegorz Chrupała • Ana Guerberof Arenas • Malvina Nissim • Arianna Bisazza Word-level quality estimation (QE) detects erroneous spans in machine translations, which can direct and facilitate human post-editing. While the accuracy of word-level QE systems has been assessed extensively, their usability and downstream influence on the… See the full description on the dataset page: https://huggingface.co/datasets/gsarti/qe4pe.tabulartranslation10K<n<100K4 likes461 downloads1y agoHugging Facegist-sparse-attention /GSA-PT-Qwen2-7B-Instruct-chunk4-chunk4-data GSA-PT-Qwen2-7B-Instruct-chunk4-chunk4-data This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk4-chunk4. Each sample is tokenized and formatted with GSA gist tokens for continue pretraining. Paper GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding Related Model yuzhenm/GSA-PT-Qwen2-7B-Instruct-chunk4-chunk4 — model trained on this dataset tabular10K<n<100K0 likes357 downloads6mo agoHugging Facegsaltintas /controlled-datatext10M<n<100M1 likes344 downloads8mo agoHugging FaceTCabbage /gsat-vocab-sentences-tts GSAT Vocabulary TTS Audio Text-to-speech audio files for GSAT (General Scholastic Ability Test) English vocabulary. Structure audio/ - MP3 audio files organized by hash prefix (e.g., audio/ab/abcd1234....mp3) index.jsonl - Index file mapping hashes to text and TTS engine used Engines Kokoro (af_heart voice) - Used for lemmas (single words/phrases) Supertonic (M1 voice) - Used for example sentences Audio Format Format: MP3 Sample rate: 24kHz… See the full description on the dataset page: https://huggingface.co/datasets/TCabbage/gsat-vocab-sentences-tts.audiotext-to-speech10K<n<100K0 likes322 downloads9mo agoHugging Face