datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1645559101lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1646052073
GEM Submission
Submission name: Hugging Face test T5-base.outputs.json 36bf2a59
lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1646049601t5_large_supervised_proportional_1MThis data set is created by randomly sampling 1M documents from the large supervised proportional mixture from the T5 repository.
The code to produce this sampled dataset can be found here.
lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1645800191lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1646049378lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1646050898lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1646049876lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1645558682lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1646049424lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1646051364triviaqa_gar-t5_expansions
Dataset Summary
The repo provides answer,title and sentence expansions for the Trivia QA corpus with gar-T5.
Dataset Structure
There are dev and test folds
An example data entry of the dev split looks as follows:
{
"id": "1",
"predicted_answers": ["Bz"], "predicted_titles": ["Vehicle registration plates of Belize *** Vehicle registration plate"], "predicted_sentences": ["The international code for Belize is \"\"BZ\"\"."]
}
An example data entry of the test split… See the full description on the dataset page: https://huggingface.co/datasets/castorini/triviaqa_gar-t5_expansions.nq_gar-t5_expansions
Dataset Summary
The repo provides answer, title and sentence expansions for the Natural Questions corpus with gar-T5.
Dataset Structure
There are dev and test folds
An example data entry of the dev split looks as follows:
{
"id": "1",
"predicted_answers": ["312"], "predicted_titles": ["Invisible Man"], "predicted_sentences": ["The Invisible Man First edition Author Ralph Ellison Cover artist M."]
}
An example data entry of the test split looks as follows:
{… See the full description on the dataset page: https://huggingface.co/datasets/castorini/nq_gar-t5_expansions.tis-quantile-datasets-gtr-t5-base
Targeted Instruction Selection: Quantile Datasets (EMBED)
This repository contains distance quantile subsets computed using the EMBED data representation method, as presented in the paper A Critical Look at Targeted Instruction Selection: Disentangling What Matters (and What Doesn't).
Project Resources
Paper: arXiv:2602.14696
GitHub: dcml-lab/targeted-instruction-selection
Dataset Description
Instruction fine-tuning of large language models (LLMs) often… See the full description on the dataset page: https://huggingface.co/datasets/Harvard-DCML/tis-quantile-datasets-gtr-t5-base.google__flan-t5-large-details
Dataset Card for Evaluation run of google/flan-t5-large
Dataset automatically created during the evaluation run of model google/flan-t5-large
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/google__flan-t5-large-details.google__flan-t5-small-details
Dataset Card for Evaluation run of google/flan-t5-small
Dataset automatically created during the evaluation run of model google/flan-t5-small
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/google__flan-t5-small-details.google__flan-t5-base-details
Dataset Card for Evaluation run of google/flan-t5-base
Dataset automatically created during the evaluation run of model google/flan-t5-base
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/google__flan-t5-base-details.google__flan-t5-xl-details
Dataset Card for Evaluation run of google/flan-t5-xl
Dataset automatically created during the evaluation run of model google/flan-t5-xl
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/google__flan-t5-xl-details.SO-Python_basics_QA-filtered-2023-T5_paraphrased-tanh_scoregoogle__flan-t5-xxl-details
Dataset Card for Evaluation run of google/flan-t5-xxl
Dataset automatically created during the evaluation run of model google/flan-t5-xxl
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/google__flan-t5-xxl-details.flan_t5_qnaflan_t5_qna
Biology-T5-QA-Datat5_codePenalDatasetglossa-bidirectional-t5-bartinlegal-laysum-for-t5-with-similarity-scoreduckdb-docbench
DocBench: A Synthetic DuckDB Text-to-SQL Benchmark
DocBench is a synthetic Text-to-SQL benchmark dataset consisting of 2430 question/sql pairs derived from the DuckDB documentation, specifically designed to probe language models for knowledge of DuckDB-specific SQL functionality.
The dataset covers functions, aggregates, operators, statements, keywords, and multi-keyword expressions available in DuckDB 1.1.3 and its default extensions.
Dataset Structure
Each example… See the full description on the dataset page: https://huggingface.co/datasets/t5r/duckdb-docbench.T5_german_summaries_filtered_convos
Synthetic Call Center Summaries Dataset T5 German
Overview
This dataset contains synthetic summaries of call center conversations generated by different prompt configurations.
Each record (in JSON Lines format) includes:
The original dialogue metadata.
A generated summary tailored to provide quick insights for call center service agents.
Evaluation metrics
Information on model
T-Systems-onsite/mt5-small-sum-de-en-v2
source_prefix: "summarize: "… See the full description on the dataset page: https://huggingface.co/datasets/marccgrau/T5_german_summaries_filtered_convos.inlegal-laysum-flan-t5test-t5-fine-tunet5-fine-tune-mixture
