gsa
Datasets
All datasets matching “gsa”flores_101One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the
lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource
languages, consider only restricted domains, or are low quality because they are constructed using
semi-automatic procedures. In this work, we introduce the FLORES evaluation benchmark, consisting of 3001
sentences extracted from English Wikipedia and covering a variety of different topics and domains.
These sentences have been translated in 101 languages by professional translators through a carefully
controlled process. The resulting dataset enables better assessment of model quality on the long tail of
low-resource languages, including the evaluation of many-to-many multilingual translation systems, as all
translations are multilingually aligned. By publicly releasing such a high-quality and high-coverage dataset,
we hope to foster progress in the machine translation community and beyond.clean_mc4_itA thoroughly cleaned version of the Italian portion of the multilingual
colossal, cleaned version of Common Crawl's web crawl corpus (mC4) by AllenAI.
Based on Common Crawl dataset: "https://commoncrawl.org".
This is the processed version of Google's mC4 dataset by AllenAI, with further cleaning
detailed in the repository README file.qe4pe
Quality Estimation for Post-Editing (QE4PE)
For more details on QE4PE, see our paper and our Github repository
Gabriele Sarti • Vilém Zouhar • Grzegorz Chrupała • Ana Guerberof Arenas • Malvina Nissim • Arianna Bisazza
Word-level quality estimation (QE) detects erroneous spans in machine translations, which can direct and facilitate human post-editing. While the accuracy of word-level QE systems has been assessed extensively, their usability and downstream influence on the… See the full description on the dataset page: https://huggingface.co/datasets/gsarti/qe4pe.GSA-PT-Qwen2-7B-Instruct-chunk4-chunk4-data
GSA-PT-Qwen2-7B-Instruct-chunk4-chunk4-data
This is the continue pretraining dataset used for training GSA (Gist Sparse Attention) models based on Qwen2-7B-Instruct with chunk size chunk4-chunk4.
Each sample is tokenized and formatted with GSA gist tokens for continue pretraining.
Paper
GSA: Gist Sparse Attention via Learnable Compression and Selective Unfolding
Related Model
yuzhenm/GSA-PT-Qwen2-7B-Instruct-chunk4-chunk4 — model trained on this dataset
controlled-datagsat-vocab-sentences-tts
GSAT Vocabulary TTS Audio
Text-to-speech audio files for GSAT (General Scholastic Ability Test) English vocabulary.
Structure
audio/ - MP3 audio files organized by hash prefix (e.g., audio/ab/abcd1234....mp3)
index.jsonl - Index file mapping hashes to text and TTS engine used
Engines
Kokoro (af_heart voice) - Used for lemmas (single words/phrases)
Supertonic (M1 voice) - Used for example sentences
Audio Format
Format: MP3
Sample rate: 24kHz… See the full description on the dataset page: https://huggingface.co/datasets/TCabbage/gsat-vocab-sentences-tts.
