datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OpenGameArt-GPL-2.0
Dataset Card for OpenGameArt-GPL-2.0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the GNU General Public License version 2.0 (GPL-2.0). The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata.
Languages
The dataset is primarily monolingual:
English (en): All asset… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-GPL-2.0.G-PlanET
Dataset Card for Dataset Name
Dataset Summary
This G-PlanET dataset is built on AI2 ALFRED.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/yuchenlin/G-PlanET.OpenGameArt-GPL-3.0
Dataset Card for OpenGameArt-GPL-3.0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the GNU General Public License version 3.0 (GPL-3.0). The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, and textures along with their associated metadata.
Languages
The dataset is primarily monolingual:
English (en): All asset descriptions… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-GPL-3.0.OpenGameArt-GPL-3.0
Dataset Card for OpenGameArt-GPL-3.0
Dataset Summary
This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the GNU General Public License version 3.0 (GPL-3.0). The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, and textures along with their associated metadata.
Languages
The dataset is primarily monolingual:
English (en): All asset descriptions… See the full description on the dataset page: https://huggingface.co/datasets/irfankabir02/OpenGameArt-GPL-3.0.alia_dogv
📘 ALIA_DOGV Dataset
The ALIA_DOGV dataset is a multilingual resource designed for text generation.
The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries.
Each entry includes information about the text's language, format, text, and metadata.
🧾 Column Descriptions
Field
Type
Description
format
string
Indicates the text format. All entries use "md" (Markdown).
language
string
Language of the… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_dogv.dogv_parallel
DOGV_PARALLEL Dataset
Dataset Summary
DOGV_PARALLEL is a parallel dataset for Valencian (VA) to Spanish (ES) translation. It consists of sentence pairs in Valencian and Spanish, along with the source file from which the data was extracted. This dataset is designed to support machine translation tasks and linguistic research.
Dataset Structure
Each row in the dataset includes the following columns:
VA: A sentence in Valencian.
ES: The corresponding translation… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/dogv_parallel.alia_les_corts
📘 ALIA_LES_CORTS Dataset
The ALIA_LES_CORTS dataset is a multilingual resource designed for text generation.
The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries.
Each entry includes information about the text's language, format, text, and metadata.
🧾 Column Descriptions
Field
Type
Description
format
string
Indicates the text format. All entries use "md" (Markdown).
language
string
Language of… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_les_corts.boua_parallel
BOUA_PARALLEL Dataset
Dataset Summary
BOUA_PARALLEL is a parallel dataset for Valencian (VA) to Spanish (ES) translation. It consists of sentence pairs in Valencian and Spanish, along with the source file from which the data was extracted. This dataset is designed to support machine translation tasks and linguistic research.
Dataset Structure
Each row in the dataset includes the following columns:
VA: A sentence in Valencian.
ES: The corresponding translation… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/boua_parallel.alia_amic
📘 ALIA_AMIC Dataset
The ALIA_AMIC dataset is a monolingual resource designed for text generation.
The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries.
Each entry includes information about the text's language, format, text, and metadata.
🧾 Column Descriptions
Field
Type
Description
format
string
Indicates the text format. All entries use "md" (Markdown).
language
string
Language of the… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_amic.alia_multilingual_parallel_sentences
MULTILINGUAL PARALLEL SENTENCES Dataset
The dataset is built from parallel corpora for translation tasks and is intended to be used for continual pretraining of language models.
It provides aligned sentences in multiple languages to facilitate multilingual learning.
Dataset Structure
The dataset is stored in a single file: a JSON Lines file where each line contains sentences in multiple languages. Each sentence is prefixed with the full name of the language.
The following… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_multilingual_parallel_sentences.amic_parallel
AMIC_PARALLEL Dataset
Dataset Summary
AMIC_PARALLEL is a parallel dataset for Valencian (VA) to Spanish (ES) translation. It consists of sentence pairs in Valencian and Spanish, along with the source file from which the data was extracted. This dataset is designed to support machine translation tasks and linguistic research.
Dataset Structure
Each row in the dataset includes the following columns:
VA: A sentence in Valencian.
ES: The corresponding translation… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/amic_parallel.uv_parallel_va_en
UV_PARALLEL_VA_EN Dataset
Dataset Summary
UV_PARALLEL_VA_EN is a parallel dataset for machine translation between Valencian (VA) and English (EN).It consists of aligned sentence pairs along with the source file from which each pair was extracted.The dataset is intended for research in machine translation, cross-lingual NLP, and linguistic analysis.
Dataset Structure
Each row in the dataset includes the following fields:
VA: A sentence in Valencian.
EN: The… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/uv_parallel_va_en.uv_parallel_va_es
UV_PARALLEL_VA_ES Dataset
Dataset Summary
UV_PARALLEL_VA_ES is a parallel dataset for machine translation between Valencian (VA) and Spanish (ES).It consists of aligned sentence pairs along with the source file from which each pair was extracted.The dataset is intended for research in machine translation, cross-lingual NLP, and linguistic analysis.
Dataset Structure
Each row in the dataset includes the following fields:
VA: A sentence in Valencian.
ES: The… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/uv_parallel_va_es.uji_parallel_va_es
UJI_PARALLEL_VA_ES Dataset
Dataset Summary
UJI_PARALLEL_VA_ES is a parallel dataset for machine translation between Valencian (VA) and Spanish (ES).It consists of aligned sentence pairs along with the source file from which each pair was extracted.The dataset is intended for research in machine translation, cross-lingual NLP, and linguistic analysis.
Dataset Structure
Each row in the dataset includes the following fields:
VA: A sentence in Valencian.
ES: The… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/uji_parallel_va_es.alia_gva_communications
📘 ALIA_GVA_Communications Dataset
The ALIA_GVA_Communications dataset is a multilingual resource designed for text generation.
The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries.
Each entry includes information about the text's language, format, text, source, and metadata.
🧾 Column Descriptions
Field
Type
Description
format
string
Indicates the text format. All entries use "md" (Markdown).… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_gva_communications.gpl-nfcorpusgpl-trec-covidalia_uji
📘 ALIA_UJI Dataset
The ALIA_UJI dataset is a multilingual resource designed for text generation, with documents sourced from the Universitat Jaume I.
The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries.
Each entry includes information about the text's language, format, text, source, and metadata.
🧾 Column Descriptions
Field
Type
Description
format
string
Indicates the text format. All entries… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_uji.alia_uv
📘 ALIA_UV Dataset
The ALIA_UV dataset is a multilingual resource designed for text generation, with documents sourced from the Universitat de València.
The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries.
Each entry includes information about the text's language, format, text, source, and metadata.
🧾 Column Descriptions
Field
Type
Description
format
string
Indicates the text format. All entries… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_uv.uji_parallel_va_en
UJI_PARALLEL_VA_EN Dataset
Dataset Summary
UJI_PARALLEL_VA_EN is a parallel dataset for machine translation between Valencian (VA) and English (EN).It consists of aligned sentence pairs along with the source file from which each pair was extracted.The dataset is intended for research in machine translation, cross-lingual NLP, and linguistic analysis.
Dataset Structure
Each row in the dataset includes the following fields:
VA: A sentence in Valencian.
EN: The… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/uji_parallel_va_en.discriminative_counterfeit_es
🔍 DISCRIMINATIVE COUNTERFEIT_ES Dataset
The DISCRIMINATIVE COUNTERFEIT_ES dataset is a discriminative corpus for counterfeit and validity checking of trademark-related claims.The task consists of classifying claims into one of two categories:
fake
not-fake
Claims follow a standardized textual pattern in Spanish:
"La marca {name} se dedica a {description}."
This dataset enables training and evaluating models for fine-grained legal-status classification, supporting research in… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/discriminative_counterfeit_es.gpl-fiqagpl-scifactalia_gva_grammar
📘 ALIA_GVA_Grammar Dataset
The ALIA_GVA_Grammar dataset is a resource designed for text generation.
The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries.
Each entry includes information about the text's language, format, text, source, and metadata.
🧾 Column Descriptions
Field
Type
Description
format
string
Indicates the text format. All entries use "md" (Markdown).
language
string
Language of… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_gva_grammar.discriminative_counterfeit_en
🔍 DISCRIMINATIVE COUNTERFEIT_EN Dataset
The DISCRIMINATIVE COUNTERFEIT_EN dataset is a discriminative corpus for counterfeit and validity checking of trademark-related claims.The task consists of classifying claims into one of two categories:
fake
not-fake
Claims follow a standardized textual pattern in English:
"The trademark {name} is involved in {description}."
This dataset enables training and evaluating models for fine-grained legal-status classification, supporting… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/discriminative_counterfeit_en.gva_translation
GVA_TRANSLATION Dataset
Dataset Summary
GVA_TRANSLATION Dataset is a parallel dataset for machine translation between Valencian (VA) and Spanish (ES).The data in this dataset was shared by the "Dirección General de Política Lingüística de la Generalitat Valenciana" exclusively for use as training data for the ALIA family of models.
The dataset is intended for research in machine translation, cross-lingual NLP, and linguistic analysis.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/gva_translation.gpl-fevergpl-quoragpl-all-mix-450kgpl-nq
