datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
qe4pe
Quality Estimation for Post-Editing (QE4PE)
For more details on QE4PE, see our paper and our Github repository
Gabriele Sarti • Vilém Zouhar • Grzegorz Chrupała • Ana Guerberof Arenas • Malvina Nissim • Arianna Bisazza
Word-level quality estimation (QE) detects erroneous spans in machine translations, which can direct and facilitate human post-editing. While the accuracy of word-level QE systems has been assessed extensively, their usability and downstream influence on the… See the full description on the dataset page: https://huggingface.co/datasets/gsarti/qe4pe.grote-logs
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/gsarti/grote-logs.us-gsa-surplus-auctions
U.S. Government (GSA) Surplus Auction Dataset
This dataset lists completed U.S. federal surplus auction lots sold through GSA Auctions (https://gsaauctions.gov), one row per lot. It is compiled and published by GovAuctions.app (https://govauctions.app) and is the lot-level companion to the GovAuctions.app Surplus Price Index (https://govauctions.app/research/surplus-price-index).
Canonical page: https://govauctions.app/research/open-dataset
Source repository, updated monthly:… See the full description on the dataset page: https://huggingface.co/datasets/govauctions/us-gsa-surplus-auctions.seq_level_training_datarebus-reasoningSYSTEM_PROMPT = """# Come risolvere un rebus
Sei un esperto risolutore di giochi enigmistici. Il seguente gioco contiene una frase cifrata (**Rebus**) nella quale alcune parole sono state sostituite da delle **Definizioni** di cruciverba fornite tra parentesi quadre. Tutte le parole e le frasi sono esclusivamente in lingua italiana. Lo scopo del gioco è quello di identificare le **Risposte** corrette e sostituirle alle definizioni nel Rebus, producendo una **Prima Lettura** che verrà poi… See the full description on the dataset page: https://huggingface.co/datasets/gsarti/rebus-reasoning.motorola-reviewseureka-rebus-calamita-2024
Dataset Card for EurekaRebus CALAMITA 2024
Dataset Summary
This dataset contains a frozen version of the EurekaRebus dataset (Sarti et al. 2024) used for LLM evaluation in the 2024 CALAMITA evaluation campaign.
Refer to the original dataset for an overview of the contents.
Data Splits
Train: 80158 examples, ignored for the purpose of the CALAMITA evaluation campaign unless fine-tuning is involved.
Test: 3167 examples, used for the CALAMITA evaluation… See the full description on the dataset page: https://huggingface.co/datasets/gsarti/eureka-rebus-calamita-2024.crowdsourced-sentiment
Dataset Card for Dataset Name
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper [optional]: [More Information Needed]
Demo [optional]: [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/gsar78/crowdsourced-sentiment.gold_dataset_8kgold_dataset_4k_balancedgold_dataset_4k
