datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fake_job_postings_balanced_en
🧠 BALANCED_FAKE_JOB_POSTINGS_EN Dataset
📘 Overview
This dataset is a balanced English version of the original Fake Job Postings dataset from Kaggle:
Real or Fake? Fake Job Posting Prediction.
It contains 1,730 job postings, equally divided between fraudulent (fake) and non-fraudulent (real) listings.
All text fields remain in English, preserving the semantic meaning and structure of the original dataset. Only balancing was performed — no translation or additional… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/fake_job_postings_balanced_en.cocoterosCOCOTEROS Dataset V1.1
Dataset Summary: The COCOTEROS dataset is designed for constrained text generation tasks with the added feature of providing contextual information to assist models in generating text. The dataset is structured to allow models to generate coherent phrases based on a set of keywords and a linguistic context which serves as the co-text of the keywords provided. This makes COCOTEROS suitable for tasks where the generated text needs to be related both to a set of specific… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/cocoteros.fake_job_postings_balanced_va
🧠 BALANCED_FAKE_JOB_POSTINGS_VA Dataset
📘 Overview
This dataset is a manually translated and balanced Valencian version of the original Fake Job Postings dataset from Kaggle:Real or Fake? Fake Job Posting Prediction.
It contains 1,730 job postings, equally divided between fraudulent (fake) and non-fraudulent (real) listings.All text fields (e.g., job title, company profile, description, requirements) have been manually translated into Valencian, preserving the semantic… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/fake_job_postings_balanced_va.details_lilloukas__GPlatty-30B
Dataset Card for Evaluation run of lilloukas/GPlatty-30B
Dataset Summary
Dataset automatically created during the evaluation run of model lilloukas/GPlatty-30B on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_lilloukas__GPlatty-30B.G-PlanET
Dataset Card for Dataset Name
Dataset Summary
This G-PlanET dataset is built on AI2 ALFRED.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/yuchenlin/G-PlanET.ES-VA_translation_test
Subtask (ES-VA_translation) of Phrases adaptability task
This dataset was built from 200,000 sentences extracted from the Common Voice tool, an open resource that collects text contributions in various languages. These sentences were subjected to a rigorous filtering process, selecting only those with the greatest linguistic richness to ensure their usefulness in applications requiring language diversity and complexity.
Subsequently, the selected sentences were translated from… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/ES-VA_translation_test.alia_dogv
📘 ALIA_DOGV Dataset
The ALIA_DOGV dataset is a multilingual resource designed for text generation.
The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries.
Each entry includes information about the text's language, format, text, and metadata.
🧾 Column Descriptions
Field
Type
Description
format
string
Indicates the text format. All entries use "md" (Markdown).
language
string
Language of the… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_dogv.OpenGameArt-GPL-3.0
Dataset Card for OpenGameArt-GPL-3.0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the GNU General Public License version 3.0 (GPL-3.0). The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, and textures along with their associated metadata.
Languages
The dataset is primarily monolingual:
English (en): All asset descriptions… See the full description on the dataset page: https://huggingface.co/datasets/irfankabir02/OpenGameArt-GPL-3.0.DOM-Dataset-v1
DOM-Dataset-v1
This dataset contains a collection of raw HTML files (DOM snapshots) organized under the data/ directory.
Total HTML files: 52323
File layout: data/**/.html
Access
Use snapshot_download to retrieve the full dataset locally:
from huggingface_hub import snapshot_download
local_dir = snapshot_download(repo_id="gplsi/DOM-Dataset-v1", repo_type="dataset")
print(local_dir)
HTML Corpus
This dataset contains a collection of raw HTML files uploaded… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/DOM-Dataset-v1.OpenGameArt-GPL-2.0
Dataset Card for OpenGameArt-GPL-2.0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the GNU General Public License version 2.0 (GPL-2.0). The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata.
Languages
The dataset is primarily monolingual:
English (en): All asset… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-GPL-2.0.warabi-dictionary-extended-gpl
IME Dictionary Extended GPL
Japanese IME dictionary pack. The SQLite database and TSV files contain the
same conversion and prediction records. See NOTICE.md before redistribution.
Compilation license: GPL-3.0-only
SQLite tables: entries, entry_sources, predictions, prediction_sources
TSV exports: entries.tsv, predictions.tsv
Kaomoji flag
The SQLite kaomoji column (0/1) marks symbol-structured emoticons; filter
with WHERE NOT kaomoji for a face-free dictionary.… See the full description on the dataset page: https://huggingface.co/datasets/fa0311/warabi-dictionary-extended-gpl.dogv_parallel
DOGV_PARALLEL Dataset
Dataset Summary
DOGV_PARALLEL is a parallel dataset for Valencian (VA) to Spanish (ES) translation. It consists of sentence pairs in Valencian and Spanish, along with the source file from which the data was extracted. This dataset is designed to support machine translation tasks and linguistic research.
Dataset Structure
Each row in the dataset includes the following columns:
VA: A sentence in Valencian.
ES: The corresponding translation… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/dogv_parallel.alia_tourism
📘 ALIA_TOURISM Dataset
The ALIA_TOURISM dataset is a multilingual resource designed for text generation within the tourism domain.
The source field indicates the data origin.
Documents in the train partition are restricted to sources with LLM-permissive licenses or explicit donations to the ALIA project,
whereas the test partition contains no such restrictions.
The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries.
Each… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_tourism.kjv_gpl_en
King James Version (GPL Edition)
Description
The King James Version (KJV), also known as the Authorized Version, is the most influential English Bible translation ever produced. First published in 1611 by a commission of 47 scholars appointed by King James I, it was translated from the original Hebrew and Greek texts. This edition is released under the GNU General Public License (GPL) by the getBible project. While the KJV text itself is public domain, this… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/kjv_gpl_en.gplates-tectonic-intelligence
GPlates Tectonic Intelligence Dataset
Overview
GPlates Tectonic Intelligence Dataset is a standardized geospatial dataset maintained by NORA Research Lab, converting the Müller et al. (2019) global plate reconstruction model into AI-ready spatial features.
Every cell in a 1.0° global grid (360 × 180 = 64,800 points, present day / 0 Ma) is tagged with its reconstruction plate ID, seafloor age, and distance to the nearest mid-ocean ridge, subduction trench… See the full description on the dataset page: https://huggingface.co/datasets/NoraResearchLab/gplates-tectonic-intelligence.alia_les_corts
📘 ALIA_LES_CORTS Dataset
The ALIA_LES_CORTS dataset is a multilingual resource designed for text generation.
The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries.
Each entry includes information about the text's language, format, text, and metadata.
🧾 Column Descriptions
Field
Type
Description
format
string
Indicates the text format. All entries use "md" (Markdown).
language
string
Language of… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_les_corts.CA-VA_alignment_test
Subtask (CA-VA_Alignment) of Phrases adaptability task
This dataset was built from 200,000 sentences extracted from the Common Voice tool, an open resource that collects text contributions in various languages. These sentences were subjected to a rigorous filtering process, selecting only those with the greatest linguistic richness to ensure their usefulness in applications requiring language diversity and complexity.
Subsequently, the selected sentences were translated from Spanish… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/CA-VA_alignment_test.boua_parallel
BOUA_PARALLEL Dataset
Dataset Summary
BOUA_PARALLEL is a parallel dataset for Valencian (VA) to Spanish (ES) translation. It consists of sentence pairs in Valencian and Spanish, along with the source file from which the data was extracted. This dataset is designed to support machine translation tasks and linguistic research.
Dataset Structure
Each row in the dataset includes the following columns:
VA: A sentence in Valencian.
ES: The corresponding translation… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/boua_parallel.OpenGameArt-GPL-3.0
Dataset Card for OpenGameArt-GPL-3.0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the GNU General Public License version 3.0 (GPL-3.0). The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, and textures along with their associated metadata.
Languages
The dataset is primarily monolingual:
English (en): All asset descriptions… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-GPL-3.0.alia_boua
📘 ALIA_BOUA Dataset
The ALIA_BOUA dataset is a multilingual resource designed for text generation.
The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries.
Each entry includes information about the text's language, format, text, and metadata.
🧾 Column Descriptions
Field
Type
Description
format
string
Indicates the text format. All entries use "md" (Markdown).
language
string
Language of the… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_boua.alia_intellectual_property
📘 ALIA_INTELLECTUAL_PROPERTY Dataset
The ALIA_INTELLECTUAL_PROPERTY dataset is a multilingual resource designed for text generation tasks within the intellectual property (IP) domain, including topics such as copyrights, patents, trademarks, and related legal and institutional information.
The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries.
Each entry includes information about the text's source, language, format, text… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_intellectual_property.xnli_va
Dataset Summary
This dataset is a professional translation into Valencian of the Cross-lingual Natural Language Inference XNLI dataset.
XNLI-va is a collection of 5.010 sentence pairs annotated with textual entailment.
The original dataset was restricted to only non-commercial research purposes under the Creative Commons Attribution Non-commercial 4.0 International Public License.
Dataset Structure
premise: a string feature.
hypothesis: a string feature.
label: a… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/xnli_va.SocialTOX
📚 Dataset: Comentarios anotados con toxicidad y constructividad
El corpus desarrollado en esta investigación es una extensión del NECOS-TOX corpus. A diferencia del original, nuestra anotación incluye comentarios de 16 medios de noticias en español, mientras que el corpus NECOS-TOX contenía 1,419 comentarios de 10 noticias del periódico El Mundo.
📝 Guía de anotación de toxicidad y constructividad
La ampliación del corpus se realizó siguiendo la guía de anotación de… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/SocialTOX.alia_amic
📘 ALIA_AMIC Dataset
The ALIA_AMIC dataset is a monolingual resource designed for text generation.
The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries.
Each entry includes information about the text's language, format, text, and metadata.
🧾 Column Descriptions
Field
Type
Description
format
string
Indicates the text format. All entries use "md" (Markdown).
language
string
Language of the… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_amic.MULTICOMMULTICOM V1.1
This repository hosts the MULTICOM dataset, a novel benchmark for evaluating the multilingual commonsense generation abilities of Large Language Models (LLMs), as presented in the paper Do LLMs exhibit the same commonsense capabilities across languages?.
The dataset extends the COCOTEROS dataset to four languages: English, Spanish, Dutch, and Valencian. The task involves generating a commonsensical sentence that includes a given triplet of words.
alia_multilingual_parallel_sentences
MULTILINGUAL PARALLEL SENTENCES Dataset
The dataset is built from parallel corpora for translation tasks and is intended to be used for continual pretraining of language models.
It provides aligned sentences in multiple languages to facilitate multilingual learning.
Dataset Structure
The dataset is stored in a single file: a JSON Lines file where each line contains sentences in multiple languages. Each sentence is prefixed with the full name of the language.
The following… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_multilingual_parallel_sentences.amic_parallel
AMIC_PARALLEL Dataset
Dataset Summary
AMIC_PARALLEL is a parallel dataset for Valencian (VA) to Spanish (ES) translation. It consists of sentence pairs in Valencian and Spanish, along with the source file from which the data was extracted. This dataset is designed to support machine translation tasks and linguistic research.
Dataset Structure
Each row in the dataset includes the following columns:
VA: A sentence in Valencian.
ES: The corresponding translation… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/amic_parallel.cocoteros_vaCOCOTEROS_VA Dataset
Dataset Summary:
The COCOTEROS_VA dataset is a translation of the COCOTEROS dataset, carried out by a linguist specialized in Valencian. It is designed for constrained text generation tasks with the added feature of providing contextual information to assist models in generating text. The dataset is structured to allow models to generate coherent phrases based on a set of keywords and a linguistic context, which serves as the co-text of the keywords provided. This makes… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/cocoteros_va.truthfulqa_va
TRUTHFULQA_VA Dataset
Dataset Summary
TruthfulQA_va is the Valencian version of the TruthfulQA dataset. This dataset is used to measure the truthfulness of a language model when generating answers to questions. It includes questions from different categories that some humans would answer wrongly due to false beliefs or misconceptions. Note that this version includes only the generation split.
Dataset Structure
Each row in the dataset includes the following… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/truthfulqa_va.es_vaca
ES_VACA Dataset
Dataset Summary
ES_VACA is a Spanish–Valencian–Catalan lexical correspondence dataset designed to facilitate research on lexical variation and equivalence between Spanish (ES), exclusive Valencian (VA), exclusive Catalan (CA), and shared Valencian–Catalan (VA-CA) lexical items.
The dataset is indexed by Spanish words (ES) and provides the corresponding terms in the three other categories, highlighting cases where the Valencian and Catalan varieties… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/es_vaca.
