CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gplsi /fake_job_postings_balanced_en 🧠 BALANCED_FAKE_JOB_POSTINGS_EN Dataset 📘 Overview This dataset is a balanced English version of the original Fake Job Postings dataset from Kaggle: Real or Fake? Fake Job Posting Prediction. It contains 1,730 job postings, equally divided between fraudulent (fake) and non-fraudulent (real) listings. All text fields remain in English, preserving the semantic meaning and structure of the original dataset. Only balancing was performed — no translation or additional… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/fake_job_postings_balanced_en.tabulartext-classification1K<n<10K0 likes1.5k downloads9mo agoHugging Face02gplsi /cocoterosCOCOTEROS Dataset V1.1 Dataset Summary: The COCOTEROS dataset is designed for constrained text generation tasks with the added feature of providing contextual information to assist models in generating text. The dataset is structured to allow models to generate coherent phrases based on a set of keywords and a linguistic context which serves as the co-text of the keywords provided. This makes COCOTEROS suitable for tasks where the generated text needs to be related both to a set of specific… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/cocoteros.texttext-generation1K<n<10K0 likes857 downloads11mo agoHugging Face03gplsi /fake_job_postings_balanced_va 🧠 BALANCED_FAKE_JOB_POSTINGS_VA Dataset 📘 Overview This dataset is a manually translated and balanced Valencian version of the original Fake Job Postings dataset from Kaggle:Real or Fake? Fake Job Posting Prediction. It contains 1,730 job postings, equally divided between fraudulent (fake) and non-fraudulent (real) listings.All text fields (e.g., job title, company profile, description, requirements) have been manually translated into Valencian, preserving the semantic… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/fake_job_postings_balanced_va.text-classification1K<n<10K0 likes612 downloads9mo agoHugging Face04open-llm-leaderboard-old /details_lilloukas__GPlatty-30B Dataset Card for Evaluation run of lilloukas/GPlatty-30B Dataset Summary Dataset automatically created during the evaluation run of model lilloukas/GPlatty-30B on the Open LLM Leaderboard. The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task. The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_lilloukas__GPlatty-30B.0 likes533 downloads3y agoHugging Face05yuchenlin /G-PlanET Dataset Card for Dataset Name Dataset Summary This G-PlanET dataset is built on AI2 ALFRED. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation Curation Rationale [More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/yuchenlin/G-PlanET.texttext-generation10K<n<100K5 likes153 downloads3y agoHugging Face06gplsi /ES-VA_translation_test Subtask (ES-VA_translation) of Phrases adaptability task This dataset was built from 200,000 sentences extracted from the Common Voice tool, an open resource that collects text contributions in various languages. These sentences were subjected to a rigorous filtering process, selecting only those with the greatest linguistic richness to ensure their usefulness in applications requiring language diversity and complexity. Subsequently, the selected sentences were translated from… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/ES-VA_translation_test.texttranslation1K<n<10K0 likes143 downloads9mo agoHugging Face07gplsi /alia_dogv 📘 ALIA_DOGV Dataset The ALIA_DOGV dataset is a multilingual resource designed for text generation. The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries. Each entry includes information about the text's language, format, text, and metadata. 🧾 Column Descriptions Field Type Description format string Indicates the text format. All entries use "md" (Markdown). language string Language of the… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_dogv.texttext-generation100K<n<1M1 likes128 downloads4mo agoHugging Face08irfankabir02 /OpenGameArt-GPL-3.0 Dataset Card for OpenGameArt-GPL-3.0 Dataset Summary This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the GNU General Public License version 3.0 (GPL-3.0). The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, and textures along with their associated metadata. Languages The dataset is primarily monolingual: English (en): All asset descriptions… See the full description on the dataset page: https://huggingface.co/datasets/irfankabir02/OpenGameArt-GPL-3.0.audioimage-classificationn<1K0 likes127 downloads10mo agoHugging Face09gplsi /DOM-Dataset-v1 DOM-Dataset-v1 This dataset contains a collection of raw HTML files (DOM snapshots) organized under the data/ directory. Total HTML files: 52323 File layout: data/**/.html Access Use snapshot_download to retrieve the full dataset locally: from huggingface_hub import snapshot_download local_dir = snapshot_download(repo_id="gplsi/DOM-Dataset-v1", repo_type="dataset") print(local_dir) HTML Corpus This dataset contains a collection of raw HTML files uploaded… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/DOM-Dataset-v1.other10K<n<100K0 likes103 downloads1y agoHugging Face10nyuuzyou /OpenGameArt-GPL-2.0 Dataset Card for OpenGameArt-GPL-2.0 Dataset Summary This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the GNU General Public License version 2.0 (GPL-2.0). The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, textures, and documents along with their associated metadata. Languages The dataset is primarily monolingual: English (en): All asset… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-GPL-2.0.audioimage-classificationn<1K0 likes96 downloads1y agoHugging Face11fa0311 /warabi-dictionary-extended-gpl IME Dictionary Extended GPL Japanese IME dictionary pack. The SQLite database and TSV files contain the same conversion and prediction records. See NOTICE.md before redistribution. Compilation license: GPL-3.0-only SQLite tables: entries, entry_sources, predictions, prediction_sources TSV exports: entries.tsv, predictions.tsv Kaomoji flag The SQLite kaomoji column (0/1) marks symbol-structured emoticons; filter with WHERE NOT kaomoji for a face-free dictionary.… See the full description on the dataset page: https://huggingface.co/datasets/fa0311/warabi-dictionary-extended-gpl.text1M<n<10M0 likes96 downloads1mo agoHugging Face12gplsi /dogv_parallel DOGV_PARALLEL Dataset Dataset Summary DOGV_PARALLEL is a parallel dataset for Valencian (VA) to Spanish (ES) translation. It consists of sentence pairs in Valencian and Spanish, along with the source file from which the data was extracted. This dataset is designed to support machine translation tasks and linguistic research. Dataset Structure Each row in the dataset includes the following columns: VA: A sentence in Valencian. ES: The corresponding translation… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/dogv_parallel.texttranslation100K<n<1M1 likes85 downloads7mo agoHugging Face13gplsi /alia_tourism 📘 ALIA_TOURISM Dataset The ALIA_TOURISM dataset is a multilingual resource designed for text generation within the tourism domain. The source field indicates the data origin. Documents in the train partition are restricted to sources with LLM-permissive licenses or explicit donations to the ALIA project, whereas the test partition contains no such restrictions. The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries. Each… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_tourism.text-generation10K<n<100K0 likes84 downloads4mo agoHugging Face14k-mktr /kjv_gpl_en King James Version (GPL Edition) Description The King James Version (KJV), also known as the Authorized Version, is the most influential English Bible translation ever produced. First published in 1611 by a commission of 47 scholars appointed by King James I, it was translated from the original Hebrew and Greek texts. This edition is released under the GNU General Public License (GPL) by the getBible project. While the KJV text itself is public domain, this… See the full description on the dataset page: https://huggingface.co/datasets/k-mktr/kjv_gpl_en.tabular10K<n<100K0 likes77 downloads2mo agoHugging Face15NoraResearchLab /gplates-tectonic-intelligence GPlates Tectonic Intelligence Dataset Overview GPlates Tectonic Intelligence Dataset is a standardized geospatial dataset maintained by NORA Research Lab, converting the Müller et al. (2019) global plate reconstruction model into AI-ready spatial features. Every cell in a 1.0° global grid (360 × 180 = 64,800 points, present day / 0 Ma) is tagged with its reconstruction plate ID, seafloor age, and distance to the nearest mid-ocean ridge, subduction trench… See the full description on the dataset page: https://huggingface.co/datasets/NoraResearchLab/gplates-tectonic-intelligence.geospatialother10K<n<100K0 likes73 downloads1mo agoHugging Face16gplsi /alia_les_corts 📘 ALIA_LES_CORTS Dataset The ALIA_LES_CORTS dataset is a multilingual resource designed for text generation. The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries. Each entry includes information about the text's language, format, text, and metadata. 🧾 Column Descriptions Field Type Description format string Indicates the text format. All entries use "md" (Markdown). language string Language of… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_les_corts.texttext-generation1K<n<10K0 likes66 downloads9mo agoHugging Face17gplsi /CA-VA_alignment_test Subtask (CA-VA_Alignment) of Phrases adaptability task This dataset was built from 200,000 sentences extracted from the Common Voice tool, an open resource that collects text contributions in various languages. These sentences were subjected to a rigorous filtering process, selecting only those with the greatest linguistic richness to ensure their usefulness in applications requiring language diversity and complexity. Subsequently, the selected sentences were translated from Spanish… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/CA-VA_alignment_test.texttranslation1K<n<10K0 likes64 downloads9mo agoHugging Face18gplsi /boua_parallel BOUA_PARALLEL Dataset Dataset Summary BOUA_PARALLEL is a parallel dataset for Valencian (VA) to Spanish (ES) translation. It consists of sentence pairs in Valencian and Spanish, along with the source file from which the data was extracted. This dataset is designed to support machine translation tasks and linguistic research. Dataset Structure Each row in the dataset includes the following columns: VA: A sentence in Valencian. ES: The corresponding translation… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/boua_parallel.texttranslation10K<n<100K0 likes63 downloads7mo agoHugging Face19nyuuzyou /OpenGameArt-GPL-3.0 Dataset Card for OpenGameArt-GPL-3.0 Dataset Summary This dataset contains game artwork assets collected from OpenGameArt.org that are specifically released under the GNU General Public License version 3.0 (GPL-3.0). The dataset includes various types of game assets such as 2D art, 3D art, concept art, music, sound effects, and textures along with their associated metadata. Languages The dataset is primarily monolingual: English (en): All asset descriptions… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/OpenGameArt-GPL-3.0.audioimage-classificationn<1K0 likes60 downloads1y agoHugging Face20gplsi /alia_boua 📘 ALIA_BOUA Dataset The ALIA_BOUA dataset is a multilingual resource designed for text generation. The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries. Each entry includes information about the text's language, format, text, and metadata. 🧾 Column Descriptions Field Type Description format string Indicates the text format. All entries use "md" (Markdown). language string Language of the… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_boua.text-generation10K<n<100K0 likes60 downloads4mo agoHugging Face21gplsi /alia_intellectual_property 📘 ALIA_INTELLECTUAL_PROPERTY Dataset The ALIA_INTELLECTUAL_PROPERTY dataset is a multilingual resource designed for text generation tasks within the intellectual property (IP) domain, including topics such as copyrights, patents, trademarks, and related legal and institutional information. The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries. Each entry includes information about the text's source, language, format, text… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_intellectual_property.text-generation10K<n<100K0 likes56 downloads4mo agoHugging Face22gplsi /xnli_va Dataset Summary This dataset is a professional translation into Valencian of the Cross-lingual Natural Language Inference XNLI dataset. XNLI-va is a collection of 5.010 sentence pairs annotated with textual entailment. The original dataset was restricted to only non-commercial research purposes under the Creative Commons Attribution Non-commercial 4.0 International Public License. Dataset Structure premise: a string feature. hypothesis: a string feature. label: a… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/xnli_va.texttext-classification1K<n<10K0 likes54 downloads2y agoHugging Face23gplsi /SocialTOX 📚 Dataset: Comentarios anotados con toxicidad y constructividad El corpus desarrollado en esta investigación es una extensión del NECOS-TOX corpus. A diferencia del original, nuestra anotación incluye comentarios de 16 medios de noticias en español, mientras que el corpus NECOS-TOX contenía 1,419 comentarios de 10 noticias del periódico El Mundo. 📝 Guía de anotación de toxicidad y constructividad La ampliación del corpus se realizó siguiendo la guía de anotación de… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/SocialTOX.texttext-classification1K<n<10K1 likes54 downloads11mo agoHugging Face24gplsi /alia_amic 📘 ALIA_AMIC Dataset The ALIA_AMIC dataset is a monolingual resource designed for text generation. The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries. Each entry includes information about the text's language, format, text, and metadata. 🧾 Column Descriptions Field Type Description format string Indicates the text format. All entries use "md" (Markdown). language string Language of the… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_amic.texttext-generation100K<n<1M0 likes54 downloads9mo agoHugging Face25gplsi /MULTICOMMULTICOM V1.1 This repository hosts the MULTICOM dataset, a novel benchmark for evaluating the multilingual commonsense generation abilities of Large Language Models (LLMs), as presented in the paper Do LLMs exhibit the same commonsense capabilities across languages?. The dataset extends the COCOTEROS dataset to four languages: English, Spanish, Dutch, and Valencian. The task involves generating a commonsensical sentence that includes a given triplet of words. texttext-generation10K<n<100K0 likes49 downloads11mo agoHugging Face26gplsi /alia_multilingual_parallel_sentences MULTILINGUAL PARALLEL SENTENCES Dataset The dataset is built from parallel corpora for translation tasks and is intended to be used for continual pretraining of language models. It provides aligned sentences in multiple languages to facilitate multilingual learning. Dataset Structure The dataset is stored in a single file: a JSON Lines file where each line contains sentences in multiple languages. Each sentence is prefixed with the full name of the language. The following… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_multilingual_parallel_sentences.texttext-generation1M<n<10M0 likes48 downloads8mo agoHugging Face27gplsi /amic_parallel AMIC_PARALLEL Dataset Dataset Summary AMIC_PARALLEL is a parallel dataset for Valencian (VA) to Spanish (ES) translation. It consists of sentence pairs in Valencian and Spanish, along with the source file from which the data was extracted. This dataset is designed to support machine translation tasks and linguistic research. Dataset Structure Each row in the dataset includes the following columns: VA: A sentence in Valencian. ES: The corresponding translation… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/amic_parallel.texttranslation100K<n<1M0 likes47 downloads7mo agoHugging Face28gplsi /cocoteros_vagatedCOCOTEROS_VA Dataset Dataset Summary: The COCOTEROS_VA dataset is a translation of the COCOTEROS dataset, carried out by a linguist specialized in Valencian. It is designed for constrained text generation tasks with the added feature of providing contextual information to assist models in generating text. The dataset is structured to allow models to generate coherent phrases based on a set of keywords and a linguistic context, which serves as the co-text of the keywords provided. This makes… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/cocoteros_va.texttext-generationn<1K0 likes47 downloads11mo agoHugging Face29gplsi /truthfulqa_vagated TRUTHFULQA_VA Dataset Dataset Summary TruthfulQA_va is the Valencian version of the TruthfulQA dataset. This dataset is used to measure the truthfulness of a language model when generating answers to questions. It includes questions from different categories that some humans would answer wrongly due to false beliefs or misconceptions. Note that this version includes only the generation split. Dataset Structure Each row in the dataset includes the following… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/truthfulqa_va.texttext-generationn<1K0 likes41 downloads11mo agoHugging Face30gplsi /es_vaca ES_VACA Dataset Dataset Summary ES_VACA is a Spanish–Valencian–Catalan lexical correspondence dataset designed to facilitate research on lexical variation and equivalence between Spanish (ES), exclusive Valencian (VA), exclusive Catalan (CA), and shared Valencian–Catalan (VA-CA) lexical items. The dataset is indexed by Spanish words (ES) and provides the corresponding terms in the three other categories, highlighting cases where the Valencian and Catalan varieties… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/es_vaca.texttranslationn<1K0 likes40 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.