CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01GEM-submissions /lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1645559101textn<1K0 likes141 downloads5y agoHugging Face02GEM-submissions /lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1646052073 GEM Submission Submission name: Hugging Face test T5-base.outputs.json 36bf2a59 textn<1K0 likes141 downloads5y agoHugging Face03GEM-submissions /lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1646049601textn<1K0 likes140 downloads5y agoHugging Face04jchenyu /t5_large_supervised_proportional_1MThis data set is created by randomly sampling 1M documents from the large supervised proportional mixture from the T5 repository. The code to produce this sampled dataset can be found here. tabular1M<n<10M0 likes140 downloads4y agoHugging Face05GEM-submissions /lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1645800191textn<1K0 likes130 downloads5y agoHugging Face06GEM-submissions /lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1646049378textn<1K0 likes130 downloads5y agoHugging Face07GEM-submissions /lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1646050898textn<1K0 likes128 downloads5y agoHugging Face08GEM-submissions /lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1646049876textn<1K0 likes127 downloads5y agoHugging Face09GEM-submissions /lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1645558682textn<1K0 likes126 downloads5y agoHugging Face10GEM-submissions /lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1646049424textn<1K0 likes125 downloads5y agoHugging Face11GEM-submissions /lewtun__hugging-face-test-t5-base.outputs.json-36bf2a59__1646051364textn<1K0 likes125 downloads5y agoHugging Face12castorini /triviaqa_gar-t5_expansions Dataset Summary The repo provides answer,title and sentence expansions for the Trivia QA corpus with gar-T5. Dataset Structure There are dev and test folds An example data entry of the dev split looks as follows: { "id": "1", "predicted_answers": ["Bz"], "predicted_titles": ["Vehicle registration plates of Belize *** Vehicle registration plate"], "predicted_sentences": ["The international code for Belize is \"\"BZ\"\"."] } An example data entry of the test split… See the full description on the dataset page: https://huggingface.co/datasets/castorini/triviaqa_gar-t5_expansions.text10K<n<100K0 likes77 downloads5y agoHugging Face13castorini /nq_gar-t5_expansions Dataset Summary The repo provides answer, title and sentence expansions for the Natural Questions corpus with gar-T5. Dataset Structure There are dev and test folds An example data entry of the dev split looks as follows: { "id": "1", "predicted_answers": ["312"], "predicted_titles": ["Invisible Man"], "predicted_sentences": ["The Invisible Man First edition Author Ralph Ellison Cover artist M."] } An example data entry of the test split looks as follows: {… See the full description on the dataset page: https://huggingface.co/datasets/castorini/nq_gar-t5_expansions.text10K<n<100K1 likes74 downloads3y agoHugging Face14Harvard-DCML /tis-quantile-datasets-gtr-t5-base Targeted Instruction Selection: Quantile Datasets (EMBED) This repository contains distance quantile subsets computed using the EMBED data representation method, as presented in the paper A Critical Look at Targeted Instruction Selection: Disentangling What Matters (and What Doesn't). Project Resources Paper: arXiv:2602.14696 GitHub: dcml-lab/targeted-instruction-selection Dataset Description Instruction fine-tuning of large language models (LLMs) often… See the full description on the dataset page: https://huggingface.co/datasets/Harvard-DCML/tis-quantile-datasets-gtr-t5-base.texttext-generation10K<n<100K0 likes36 downloads7mo agoHugging Face15open-llm-leaderboard /google__flan-t5-large-detailsgated Dataset Card for Evaluation run of google/flan-t5-large Dataset automatically created during the evaluation run of model google/flan-t5-large The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/google__flan-t5-large-details.tabular10K<n<100K0 likes25 downloads2y agoHugging Face16open-llm-leaderboard /google__flan-t5-small-detailsgated Dataset Card for Evaluation run of google/flan-t5-small Dataset automatically created during the evaluation run of model google/flan-t5-small The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/google__flan-t5-small-details.tabular10K<n<100K0 likes22 downloads2y agoHugging Face17open-llm-leaderboard /google__flan-t5-base-detailsgated Dataset Card for Evaluation run of google/flan-t5-base Dataset automatically created during the evaluation run of model google/flan-t5-base The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/google__flan-t5-base-details.tabular10K<n<100K0 likes22 downloads2y agoHugging Face18open-llm-leaderboard /google__flan-t5-xl-detailsgated Dataset Card for Evaluation run of google/flan-t5-xl Dataset automatically created during the evaluation run of model google/flan-t5-xl The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/google__flan-t5-xl-details.tabular10K<n<100K0 likes21 downloads2y agoHugging Face19Myashka /SO-Python_basics_QA-filtered-2023-T5_paraphrased-tanh_scoretabular100K<n<1M0 likes20 downloads3y agoHugging Face20open-llm-leaderboard /google__flan-t5-xxl-detailsgated Dataset Card for Evaluation run of google/flan-t5-xxl Dataset automatically created during the evaluation run of model google/flan-t5-xxl The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results. An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/google__flan-t5-xxl-details.tabular10K<n<100K0 likes20 downloads2y agoHugging Face21dongwoojung /flan_t5_qnaflan_t5_qna text100K<n<1M0 likes19 downloads3y agoHugging Face22QasTurk /Biology-T5-QA-Datatext1K<n<10K1 likes16 downloads2y agoHugging Face23AssemienDev /t5_codePenalDatasettextn<1K0 likes13 downloads2y agoHugging Face24rrrr66254 /glossa-bidirectional-t5-barttext10K<n<100K0 likes12 downloads1y agoHugging Face25mohitskaushal /inlegal-laysum-for-t5-with-similarity-scoretext10K<n<100K0 likes11 downloads6mo agoHugging Face26t5r /duckdb-docbenchgated DocBench: A Synthetic DuckDB Text-to-SQL Benchmark DocBench is a synthetic Text-to-SQL benchmark dataset consisting of 2430 question/sql pairs derived from the DuckDB documentation, specifically designed to probe language models for knowledge of DuckDB-specific SQL functionality. The dataset covers functions, aggregates, operators, statements, keywords, and multi-keyword expressions available in DuckDB 1.1.3 and its default extensions. Dataset Structure Each example… See the full description on the dataset page: https://huggingface.co/datasets/t5r/duckdb-docbench.text1K<n<10K1 likes11 downloads5mo agoHugging Face27marccgrau /T5_german_summaries_filtered_convos Synthetic Call Center Summaries Dataset T5 German Overview This dataset contains synthetic summaries of call center conversations generated by different prompt configurations. Each record (in JSON Lines format) includes: The original dialogue metadata. A generated summary tailored to provide quick insights for call center service agents. Evaluation metrics Information on model T-Systems-onsite/mt5-small-sum-de-en-v2 source_prefix: "summarize: "… See the full description on the dataset page: https://huggingface.co/datasets/marccgrau/T5_german_summaries_filtered_convos.textsummarizationn<1K0 likes9 downloads1y agoHugging Face28mohitskaushal /inlegal-laysum-flan-t5text10K<n<100K0 likes9 downloads7mo agoHugging Face29amydeng2000 /test-t5-fine-tunetext1K<n<10K0 likes6 downloads4y agoHugging Face30amydeng2000 /t5-fine-tune-mixturetext1M<n<10M0 likes6 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.