datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
flores_101One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the
lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource
languages, consider only restricted domains, or are low quality because they are constructed using
semi-automatic procedures. In this work, we introduce the FLORES evaluation benchmark, consisting of 3001
sentences extracted from English Wikipedia and covering a variety of different topics and domains.
These sentences have been translated in 101 languages by professional translators through a carefully
controlled process. The resulting dataset enables better assessment of model quality on the long tail of
low-resource languages, including the evaluation of many-to-many multilingual translation systems, as all
translations are multilingually aligned. By publicly releasing such a high-quality and high-coverage dataset,
we hope to foster progress in the machine translation community and beyond.qe4pe
Quality Estimation for Post-Editing (QE4PE)
For more details on QE4PE, see our paper and our Github repository
Gabriele Sarti • Vilém Zouhar • Grzegorz Chrupała • Ana Guerberof Arenas • Malvina Nissim • Arianna Bisazza
Word-level quality estimation (QE) detects erroneous spans in machine translations, which can direct and facilitate human post-editing. While the accuracy of word-level QE systems has been assessed extensively, their usability and downstream influence on the… See the full description on the dataset page: https://huggingface.co/datasets/gsarti/qe4pe.controlled-datavigorl_datasets
ViGoRL Datasets
This repository contains the official datasets associated with the paper "Grounded Reinforcement Learning for Visual Reasoning (ViGoRL)", by Gabriel Sarch, Snigdha Saha, Naitik Khandelwal, Ayush Jain, Michael J. Tarr, Aviral Kumar, and Katerina Fragkiadaki.
Dataset Overview
These datasets are designed for training and evaluating visually grounded vision-language models (VLMs).
Datasets are organized by the visual reasoning tasks described in the ViGoRL… See the full description on the dataset page: https://huggingface.co/datasets/gsarch/vigorl_datasets.details_GSAI-ML__LLaDA-8B-Instruct_v2
Dataset Card for Evaluation run of GSAI-ML/LLaDA-8B-Instruct
Dataset automatically created during the evaluation run of model GSAI-ML/LLaDA-8B-Instruct.
The dataset is composed of 116 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_GSAI-ML__LLaDA-8B-Instruct_v2.mt_genevalThe MT-GenEval benchmark evaluates gender translation accuracy on English -> {Arabic, French, German, Hindi, Italian,
Portuguese, Russian, Spanish}. The dataset contains individual sentences with annotations on the gendered target words,
and contrastive original-invertend translations with additional preceding context.iwslt2017_contextThe IWSLT 2017 Multilingual Task addresses text translation, including zero-shot translation, with a single MT system across all directions including English, German, Dutch, Italian and Romanian. As unofficial task, conventional bilingual text translation is offered between English and Arabic, French, Japanese, Chinese, German and Korean.countqa_lite
gsarch/countqa_lite
A deterministic lite evaluation subset of Jayant-Sravan/CountQA.
Source revision: f92cc6fe46542c61e2916e3d2ae9a911e2216b1a
Source split: test
Sampling seed: 43
Output rows: 500
Schema: unchanged from the upstream dataset
CountQA is sampled at the QA-pair level. Each output row retains the original schema and contains one-element questions and answers lists, so lmms-eval's existing countqa_process_docs produces exactly 500 prompts.
Generated by… See the full description on the dataset page: https://huggingface.co/datasets/gsarch/countqa_lite.magpieThe MAGPIE corpus is a large sense-annotated corpus of potentially idiomatic expressions (PIEs), based on the British National Corpus (BNC). Potentially idiomatic expressions are like idiomatic expressions, but the term also covers literal uses of idiomatic expressions, such as 'I leave work at the end of the day.' for the idiom 'at the end of the day'. This version of the dataset reflects the filtered subset used by Dankers et al. (2022) in their investigation on how PIEs are represented by NMT models. Authors use 37k samples annotated as fully figurative or literal, for 1482 idioms that contain nouns, numerals or adjectives that are colours (which they refer to as keywords). Because idioms show syntactic and morphological variability, the focus is mostly put on nouns. PIEs and their context are separated using the original corpus’s word-level annotations.ScreenSpot-Pro-Lite
ScreenSpot-Pro-Lite
A fixed 500-example representative/challenging subset of the 1,581-example
ScreenSpot-Pro benchmark.
Selection
Sampling is proportional over the cross-product of platform, application, and UI type,
so all 26 applications remain represented and the original icon/text mix is retained.
Within every stratum, 80% is deterministic seeded sampling and 20% is a hard tier.
Hardness combines failure rate and disagreement across five full-run anchor… See the full description on the dataset page: https://huggingface.co/datasets/gsarch/ScreenSpot-Pro-Lite.grote-logs
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/gsarti/grote-logs.ReFusion
ReFusion
Dataset Summary
This dataset is the training corpus used for ReFusion, as described in our paper. It comprises approximately 3.7 million high-quality instruction tuning samples consolidated from several state-of-the-art open-source datasets. The data covers diverse domains including mathematics, coding, and general instruction following.
Composition & Sources
The dataset is constructed from the following sources:
MAmmoTH
OpenMathInstruct-2 (1M… See the full description on the dataset page: https://huggingface.co/datasets/GSAI-ML/ReFusion.temporal_expressions
Dataset Card for Tokenization Robustness
A comprehensive evaluation dataset for testing robustness of different tokenization strategies.
Dataset Details
Dataset Description
This dataset evaluates how robust language models are to different tokenization strategies and edge cases. It includes questions with multiple choice answers designed to test various aspects of tokenization handling.
Curated by: R3
Funded by [optional]: [More Information Needed]
Shared… See the full description on the dataset page: https://huggingface.co/datasets/gsaltintas/temporal_expressions.mathematical_scientific_notation
Dataset Card for Tokenization Robustness
A comprehensive evaluation dataset for testing robustness of different tokenization strategies.
Dataset Details
Dataset Description
This dataset evaluates how robust language models are to different tokenization strategies and edge cases. It includes questions with multiple choice answers designed to test various aspects of tokenization handling.
Curated by: R3
Funded by [optional]: [More Information Needed]
Shared… See the full description on the dataset page: https://huggingface.co/datasets/gsaltintas/mathematical_scientific_notation.us-gsa-surplus-auctions
U.S. Government (GSA) Surplus Auction Dataset
This dataset lists completed U.S. federal surplus auction lots sold through GSA Auctions (https://gsaauctions.gov), one row per lot. It is compiled and published by GovAuctions.app (https://govauctions.app) and is the lot-level companion to the GovAuctions.app Surplus Price Index (https://govauctions.app/research/surplus-price-index).
Canonical page: https://govauctions.app/research/open-dataset
Source repository, updated monthly:… See the full description on the dataset page: https://huggingface.co/datasets/govauctions/us-gsa-surplus-auctions.seq_level_training_dataHStarBench-Lite
HStarBench-Lite
A fixed 500-example representative/challenging subset of the 4,000-example
HStarBench mixed test set,
packaged for the Vero/lmms-eval panorama-strip evaluation.
Selection
The subset preserves the full benchmark's HOS/HPS and difficulty-level proportions:
Split
Level
Count
HOS
0
77
HOS
1
22
HOS
2
201
HPS
0
63
HPS
1
57
HPS
2
53
HPS
3
27
Within every split/level stratum, 80% is deterministic seeded sampling and 20% is a… See the full description on the dataset page: https://huggingface.co/datasets/gsarch/HStarBench-Lite.social_media_informal_text
Dataset Card for Tokenization Robustness
A comprehensive evaluation dataset for testing robustness of different tokenization strategies.
Dataset Details
Dataset Description
This dataset evaluates how robust language models are to different tokenization strategies and edge cases. It includes questions with multiple choice answers designed to test various aspects of tokenization handling.
Curated by: R3
Funded by [optional]: [More Information Needed]
Shared… See the full description on the dataset page: https://huggingface.co/datasets/gsaltintas/social_media_informal_text.GS_Mortgage_Securities_Trust_2020_GSA2_1833522
GS Mortgage Securities Trust 2020-GSA2
SEC ABS-EE asset-level filings for CIK 1833522 (GS Mortgage Securities Trust 2020-GSA2).
Filings: 45
Parquet files: 175
Total size: 68.2 MB
Reporting period start: 2020-12-07
Reporting period end: 2024-07-08
Parquet files are loan-level / asset-level data extracted from XML exhibits, organised as {accession_nodash}/{exhibit_name}.parquet. Reporting-period dates are derived from the asset-level XML (reportingPeriodEndingDate).… See the full description on the dataset page: https://huggingface.co/datasets/DenyTranDFW/GS_Mortgage_Securities_Trust_2020_GSA2_1833522.cruciverbit_augmentedGS_Mortgage_Securities_Trust_2021_GSA3_1895116
GS Mortgage Securities Trust 2021-GSA3
SEC ABS-EE asset-level filings for CIK 1895116 (GS Mortgage Securities Trust 2021-GSA3).
Filings: 39
Parquet files: 104
Total size: 4.5 MB
Reporting period start: 2021-12-13
Reporting period end: 2026-02-11
Parquet files are loan-level / asset-level data extracted from XML exhibits, organised as {accession_nodash}/{exhibit_name}.parquet. Reporting-period dates are derived from the asset-level XML (reportingPeriodEndingDate).… See the full description on the dataset page: https://huggingface.co/datasets/DenyTranDFW/GS_Mortgage_Securities_Trust_2021_GSA3_1895116.gsat-115Machine-gradable exam benchmarks produced by any-to-bench. Each subset is one
exam: the viewer table shows one row per answerable question (figures embedded);
the raw, byte-faithful bundle lives under <subset>/bundle/ — exam.json
(structured paper), answer_schema.json (strict JSON Schema an answer sheet must
satisfy), grading.json (deterministic rules + judge rubrics), manifest.json
(provenance), and assets/ (figures).
Usage
Benchmark any model against an exam:
a2b download… See the full description on the dataset page: https://huggingface.co/datasets/JacobLinCool/gsat-115.SimpleVQA-ENrebus-reasoningSYSTEM_PROMPT = """# Come risolvere un rebus
Sei un esperto risolutore di giochi enigmistici. Il seguente gioco contiene una frase cifrata (**Rebus**) nella quale alcune parole sono state sostituite da delle **Definizioni** di cruciverba fornite tra parentesi quadre. Tutte le parole e le frasi sono esclusivamente in lingua italiana. Lo scopo del gioco è quello di identificare le **Risposte** corrette e sostituirle alle definizioni nel Rebus, producendo una **Prima Lettura** che verrà poi… See the full description on the dataset page: https://huggingface.co/datasets/gsarti/rebus-reasoning.eng_Latn
Dataset Card for Tokenization Robustness
A comprehensive evaluation dataset for testing robustness of different tokenization strategies.
Dataset Details
Dataset Description
This dataset evaluates how robust language models are to different tokenization strategies and edge cases. It includes questions with multiple choice answers designed to test various aspects of tokenization handling.
Curated by: R3
Funded by [optional]: [More Information Needed]
Shared… See the full description on the dataset page: https://huggingface.co/datasets/gsaltintas/eng_Latn.faa-balloon-flying-handbook
FAA Balloon Flying Handbook Dataset
This dataset was created by processing the official FAA Balloon Flying Handbook (FAA-H-8083-11B).
If you're interested in understanding how this dataset was created, check out this blog post
or explore the details directly in the GitHub repository.
Usage:
from datasets import load_dataset
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("gsantopaolo/faa-balloon-flying-handbook")
print(dataset)
# Print the first 5 rows… See the full description on the dataset page: https://huggingface.co/datasets/gsantopaolo/faa-balloon-flying-handbook.EvoChart-QA
gsarch/EvoChart-QA
This dataset packages the EvoChart QA annotations with embedded chart images.
Each row contains the keys image, question, answer, attribute, is_clear, and chart_type.
Total rows: 1250
Game-QA-Litescript__orthography
Dataset Card for Tokenization Robustness
A comprehensive evaluation dataset for testing robustness of different tokenization strategies.
Dataset Details
Dataset Description
This dataset evaluates how robust language models are to different tokenization strategies and edge cases. It includes questions with multiple choice answers designed to test various aspects of tokenization handling.
Curated by: R3
Funded by [optional]: [More Information Needed]
Shared… See the full description on the dataset page: https://huggingface.co/datasets/gsaltintas/script__orthography.reflection
