datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
cocoterosCOCOTEROS Dataset V1.1
Dataset Summary: The COCOTEROS dataset is designed for constrained text generation tasks with the added feature of providing contextual information to assist models in generating text. The dataset is structured to allow models to generate coherent phrases based on a set of keywords and a linguistic context which serves as the co-text of the keywords provided. This makes COCOTEROS suitable for tasks where the generated text needs to be related both to a set of specific… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/cocoteros.G-PlanET
Dataset Card for Dataset Name
Dataset Summary
This G-PlanET dataset is built on AI2 ALFRED.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/yuchenlin/G-PlanET.alia_dogv
📘 ALIA_DOGV Dataset
The ALIA_DOGV dataset is a multilingual resource designed for text generation.
The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries.
Each entry includes information about the text's language, format, text, and metadata.
🧾 Column Descriptions
Field
Type
Description
format
string
Indicates the text format. All entries use "md" (Markdown).
language
string
Language of the… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_dogv.alia_les_corts
📘 ALIA_LES_CORTS Dataset
The ALIA_LES_CORTS dataset is a multilingual resource designed for text generation.
The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries.
Each entry includes information about the text's language, format, text, and metadata.
🧾 Column Descriptions
Field
Type
Description
format
string
Indicates the text format. All entries use "md" (Markdown).
language
string
Language of… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_les_corts.gp-l-only-10k
General Points (gp-l-only-10k) Dataset for Debunking SFT Generalization
This dataset is part of the research presented in the paper "Debunk the Myth of SFT Generalization".
The paper challenges the conventional belief that Supervised Fine-Tuning (SFT) primarily memorizes training data and lacks generalization capabilities, while Reinforcement Learning (RL) achieves broader robustness. Through systematic evaluation on decision-making benchmarks like Sokoban and General Points, the… See the full description on the dataset page: https://huggingface.co/datasets/Xiaofeng77/gp-l-only-10k.MULTICOMMULTICOM V1.1
This repository hosts the MULTICOM dataset, a novel benchmark for evaluating the multilingual commonsense generation abilities of Large Language Models (LLMs), as presented in the paper Do LLMs exhibit the same commonsense capabilities across languages?.
The dataset extends the COCOTEROS dataset to four languages: English, Spanish, Dutch, and Valencian. The task involves generating a commonsensical sentence that includes a given triplet of words.
alia_amic
📘 ALIA_AMIC Dataset
The ALIA_AMIC dataset is a monolingual resource designed for text generation.
The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries.
Each entry includes information about the text's language, format, text, and metadata.
🧾 Column Descriptions
Field
Type
Description
format
string
Indicates the text format. All entries use "md" (Markdown).
language
string
Language of the… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_amic.cocoteros_vaCOCOTEROS_VA Dataset
Dataset Summary:
The COCOTEROS_VA dataset is a translation of the COCOTEROS dataset, carried out by a linguist specialized in Valencian. It is designed for constrained text generation tasks with the added feature of providing contextual information to assist models in generating text. The dataset is structured to allow models to generate coherent phrases based on a set of keywords and a linguistic context, which serves as the co-text of the keywords provided. This makes… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/cocoteros_va.truthfulqa_va
TRUTHFULQA_VA Dataset
Dataset Summary
TruthfulQA_va is the Valencian version of the TruthfulQA dataset. This dataset is used to measure the truthfulness of a language model when generating answers to questions. It includes questions from different categories that some humans would answer wrongly due to false beliefs or misconceptions. Note that this version includes only the generation split.
Dataset Structure
Each row in the dataset includes the following… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/truthfulqa_va.gpl_queries_pubmedalia_multilingual_parallel_sentences
MULTILINGUAL PARALLEL SENTENCES Dataset
The dataset is built from parallel corpora for translation tasks and is intended to be used for continual pretraining of language models.
It provides aligned sentences in multiple languages to facilitate multilingual learning.
Dataset Structure
The dataset is stored in a single file: a JSON Lines file where each line contains sentences in multiple languages. Each sentence is prefixed with the full name of the language.
The following… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_multilingual_parallel_sentences.alia_gva_communications
📘 ALIA_GVA_Communications Dataset
The ALIA_GVA_Communications dataset is a multilingual resource designed for text generation.
The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries.
Each entry includes information about the text's language, format, text, source, and metadata.
🧾 Column Descriptions
Field
Type
Description
format
string
Indicates the text format. All entries use "md" (Markdown).… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_gva_communications.alia_uji
📘 ALIA_UJI Dataset
The ALIA_UJI dataset is a multilingual resource designed for text generation, with documents sourced from the Universitat Jaume I.
The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries.
Each entry includes information about the text's language, format, text, source, and metadata.
🧾 Column Descriptions
Field
Type
Description
format
string
Indicates the text format. All entries… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_uji.alia_uv
📘 ALIA_UV Dataset
The ALIA_UV dataset is a multilingual resource designed for text generation, with documents sourced from the Universitat de València.
The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries.
Each entry includes information about the text's language, format, text, source, and metadata.
🧾 Column Descriptions
Field
Type
Description
format
string
Indicates the text format. All entries… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_uv.alia_gva_grammar
📘 ALIA_GVA_Grammar Dataset
The ALIA_GVA_Grammar dataset is a resource designed for text generation.
The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries.
Each entry includes information about the text's language, format, text, source, and metadata.
🧾 Column Descriptions
Field
Type
Description
format
string
Indicates the text format. All entries use "md" (Markdown).
language
string
Language of… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_gva_grammar.cieaCOVA
cieaCOVA Dataset
cieaCOVA is a Valencian-language (va) evaluation dataset designed to benchmark large language models (LLMs) on structured reasoning and generative tasks. The dataset contains 1,982 curated examples and is specifically developed for evaluation purposes — not for model training.
The dataset is organized into two task-oriented directories, each containing a train and test split:
multiple_choice/ (train.parquet, test.parquet) — multiple-choice question answering… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/cieaCOVA.
