CoolFace
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gplsi /cocoterosCOCOTEROS Dataset V1.1 Dataset Summary: The COCOTEROS dataset is designed for constrained text generation tasks with the added feature of providing contextual information to assist models in generating text. The dataset is structured to allow models to generate coherent phrases based on a set of keywords and a linguistic context which serves as the co-text of the keywords provided. This makes COCOTEROS suitable for tasks where the generated text needs to be related both to a set of specific… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/cocoteros.texttext-generation1K<n<10K0 likes859 downloads11mo agoHugging Face02yuchenlin /G-PlanET Dataset Card for Dataset Name Dataset Summary This G-PlanET dataset is built on AI2 ALFRED. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation Rationale [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/yuchenlin/G-PlanET.texttext-generation10K<n<100K5 likes156 downloads3y agoHugging Face03gplsi /alia_dogv 📘 ALIA_DOGV Dataset The ALIA_DOGV dataset is a multilingual resource designed for text generation. The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries. Each entry includes information about the text's language, format, text, and metadata. 🧾 Column Descriptions Field Type Description format string Indicates the text format. All entries use "md" (Markdown). language string Language of the… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_dogv.texttext-generation100K<n<1M1 likes128 downloads5mo agoHugging Face04gplsi /alia_les_corts 📘 ALIA_LES_CORTS Dataset The ALIA_LES_CORTS dataset is a multilingual resource designed for text generation. The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries. Each entry includes information about the text's language, format, text, and metadata. 🧾 Column Descriptions Field Type Description format string Indicates the text format. All entries use "md" (Markdown). language string Language of… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_les_corts.texttext-generation1K<n<10K0 likes55 downloads9mo agoHugging Face05Xiaofeng77 /gp-l-only-10k General Points (gp-l-only-10k) Dataset for Debunking SFT Generalization This dataset is part of the research presented in the paper "Debunk the Myth of SFT Generalization". The paper challenges the conventional belief that Supervised Fine-Tuning (SFT) primarily memorizes training data and lacks generalization capabilities, while Reinforcement Learning (RL) achieves broader robustness. Through systematic evaluation on decision-making benchmarks like Sokoban and General Points, the… See the full description on the dataset page: https://huggingface.co/datasets/Xiaofeng77/gp-l-only-10k.texttext-generation10K<n<100K0 likes54 downloads1y agoHugging Face06gplsi /MULTICOMMULTICOM V1.1 This repository hosts the MULTICOM dataset, a novel benchmark for evaluating the multilingual commonsense generation abilities of Large Language Models (LLMs), as presented in the paper Do LLMs exhibit the same commonsense capabilities across languages?. The dataset extends the COCOTEROS dataset to four languages: English, Spanish, Dutch, and Valencian. The task involves generating a commonsensical sentence that includes a given triplet of words. texttext-generation10K<n<100K0 likes54 downloads11mo agoHugging Face07gplsi /alia_amic 📘 ALIA_AMIC Dataset The ALIA_AMIC dataset is a monolingual resource designed for text generation. The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries. Each entry includes information about the text's language, format, text, and metadata. 🧾 Column Descriptions Field Type Description format string Indicates the text format. All entries use "md" (Markdown). language string Language of the… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_amic.texttext-generation100K<n<1M0 likes53 downloads9mo agoHugging Face08gplsi /cocoteros_vagatedCOCOTEROS_VA Dataset Dataset Summary: The COCOTEROS_VA dataset is a translation of the COCOTEROS dataset, carried out by a linguist specialized in Valencian. It is designed for constrained text generation tasks with the added feature of providing contextual information to assist models in generating text. The dataset is structured to allow models to generate coherent phrases based on a set of keywords and a linguistic context, which serves as the co-text of the keywords provided. This makes… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/cocoteros_va.texttext-generationn<1K0 likes48 downloads11mo agoHugging Face09gplsi /truthfulqa_vagated TRUTHFULQA_VA Dataset Dataset Summary TruthfulQA_va is the Valencian version of the TruthfulQA dataset. This dataset is used to measure the truthfulness of a language model when generating answers to questions. It includes questions from different categories that some humans would answer wrongly due to false beliefs or misconceptions. Note that this version includes only the generation split. Dataset Structure Each row in the dataset includes the following… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/truthfulqa_va.texttext-generationn<1K0 likes42 downloads11mo agoHugging Face10chungimungi /gpl_queries_pubmedtexttext-classification1M<n<10M0 likes40 downloads2y agoHugging Face11gplsi /alia_multilingual_parallel_sentences MULTILINGUAL PARALLEL SENTENCES Dataset The dataset is built from parallel corpora for translation tasks and is intended to be used for continual pretraining of language models. It provides aligned sentences in multiple languages to facilitate multilingual learning. Dataset Structure The dataset is stored in a single file: a JSON Lines file where each line contains sentences in multiple languages. Each sentence is prefixed with the full name of the language. The following… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_multilingual_parallel_sentences.texttext-generation1M<n<10M0 likes36 downloads8mo agoHugging Face12gplsi /alia_gva_communications 📘 ALIA_GVA_Communications Dataset The ALIA_GVA_Communications dataset is a multilingual resource designed for text generation. The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries. Each entry includes information about the text's language, format, text, source, and metadata. 🧾 Column Descriptions Field Type Description format string Indicates the text format. All entries use "md" (Markdown).… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_gva_communications.texttext-generation10K<n<100K0 likes27 downloads4mo agoHugging Face13gplsi /alia_uji 📘 ALIA_UJI Dataset The ALIA_UJI dataset is a multilingual resource designed for text generation, with documents sourced from the Universitat Jaume I. The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries. Each entry includes information about the text's language, format, text, source, and metadata. 🧾 Column Descriptions Field Type Description format string Indicates the text format. All entries… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_uji.texttext-generation10K<n<100K0 likes27 downloads4mo agoHugging Face14gplsi /alia_uv 📘 ALIA_UV Dataset The ALIA_UV dataset is a multilingual resource designed for text generation, with documents sourced from the Universitat de València. The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries. Each entry includes information about the text's language, format, text, source, and metadata. 🧾 Column Descriptions Field Type Description format string Indicates the text format. All entries… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_uv.texttext-generation1K<n<10K0 likes27 downloads4mo agoHugging Face15gplsi /alia_gva_grammar 📘 ALIA_GVA_Grammar Dataset The ALIA_GVA_Grammar dataset is a resource designed for text generation. The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries. Each entry includes information about the text's language, format, text, source, and metadata. 🧾 Column Descriptions Field Type Description format string Indicates the text format. All entries use "md" (Markdown). language string Language of… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_gva_grammar.texttext-generation1K<n<10K0 likes14 downloads4mo agoHugging Face16gplsi /cieaCOVA cieaCOVA Dataset cieaCOVA is a Valencian-language (va) evaluation dataset designed to benchmark large language models (LLMs) on structured reasoning and generative tasks. The dataset contains 1,982 curated examples and is specifically developed for evaluation purposes — not for model training. The dataset is organized into two task-oriented directories, each containing a train and test split: multiple_choice/ (train.parquet, test.parquet) — multiple-choice question answering… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/cieaCOVA.texttext-generation1K<n<10K0 likes12 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.