datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
universal_dependencies
Dataset Card (v2.0) for Universal Dependencies Treebank
Version 2.0.0 introduces significant improvements and breaking changes:
Parquet Format: faster loading with HuggingFace datasets >=4.0.0
MWT Support: New mwt field provides structured multi-word token information
Enhanced Security: No more trust_remote_code=True required
Separate Versioning: Loader version (2.0.0) distinct from UD data version (2.18)
Breaking Changes:
Token sequences now exclude MWT surface forms… See the full description on the dataset page: https://huggingface.co/datasets/universal-dependencies/universal_dependencies.quranic-universal-ayahs
Qur'anic Universal Ayahs
Qur'anic Universal Audio (QUA) is a project that unifies recitations on the internet and generates timing data using forced alignment — community-verified results and constantly expanding dataset.
This dataset pairs ayah by ayah audio with word-level timestamps, DigitalKhatt letter-animation timestamps, and waqf-aware segment data. Repeated words are preserved in text_uthmani and word_timestamps, so the row reflects what the reciter… See the full description on the dataset page: https://huggingface.co/datasets/QUD-Technologies/quranic-universal-ayahs.universal_spanish_chilean_corpus
Universal Chilean Spanish Corpus
Este dataset se compone de 37_213_992 textos correspondientes a español de Chile y a español multidialectal.
Los textos en español multidialectal provienen del spanish books.
Los textos en español de Chile vienen de los dominios .cl del mc4 dataset y de tweets, noticias y reclamos de l chilean-spanish-corpus
Name
Count
Source
books
87967
spanish books
mc4
8706681
from mc4 (.cl domains) in chilean-spanish-corpus
twitter
27306583… See the full description on the dataset page: https://huggingface.co/datasets/jorgeortizfuentes/universal_spanish_chilean_corpus.Universal-glaive-function-calling-v2
Dataset Card for "Universal-glaive-function-calling-v2"
More Information needed
universal_nerThis is an exact duplicate of https://huggingface.co/datasets/universalner/universal_ner, which is not compatible with modern versions of datasets anymore because loading data via a custom script is no longer supported. All credit goes to the original creators. Original README below.
Dataset Card for Universal NER
Dataset Summary
Universal NER (UNER) is an open, community-driven initiative aimed at creating gold-standard benchmarks for Named Entity Recognition (NER)… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/universal_ner.universal-preference-hijacking-datasets
Phi: Preference Hijacking in Multi-modal Large Language Models at Inference Time
Figure 1: Examples of Phi, which can hijack MLLM's preference toward the image.
Figure 2: Example of a universal hijacking perturbation, which can be transferred across different images.
This dataset is used to train and evaluate the universal hijacking perturbations in the paper "Phi: Preference Hijacking in Multi-modal Large Language Models at Inference Time", accepted at EMNLP… See the full description on the dataset page: https://huggingface.co/datasets/yflantmy/universal-preference-hijacking-datasets.universal-dependencies-parquetThe official (?) HF dataset repo for Universal Dependencies treebanks uses a dataset script to load datasets. Dataset scripts are no longer supported as of v4 of HF Datasets.
The treebanks that I need are made directly available here as parquet files.
I've adopted the license used by most of the Universal Dependencies treebanks.
_DESCRIPTIONS = {
"af_afribooms": "UD Afrikaans-AfriBooms is a conversion of the AfriBooms Dependency Treebank, originally annotated with a simplified PoS set and… See the full description on the dataset page: https://huggingface.co/datasets/a3lem/universal-dependencies-parquet.UniversalScienceKownledge-finetome-top-20ksnli_universal
SNLI dataset with Universal triggers
This dataset consists of 7 splits, including
train: 6,000 examples are randomly sampled from "train" split of SNLI dataset out of which 3000 examples are left unmodified and universal triggers are prepended to the hypothesis of rest 3000 examples.
entailment: ~1000 randomly chosen entailment examples from validation split of SNLI dataset.
entailment_triggers: ~1000 examples of entailment split of this dataset are prepended with universal… See the full description on the dataset page: https://huggingface.co/datasets/ckverma/snli_universal.universal_dependencies_fr_spoken_fr_prompt_pos
universal_dependencies_fr_spoken_fr_prompt_pos
Summary
universal_dependencies_fr_spoken_fr_prompt_pos is a subset of the Dataset of French Prompts (DFP).It contains 58,926 rows that can be used for a part-of-speech task.The original data (without prompts) comes from the dataset universal_dependencies where only the French spoken split has been kept.A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the same format… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/universal_dependencies_fr_spoken_fr_prompt_pos.Universal-Genius-Master-2MPabasaraXE
universal_dependencies_fr_partut_fr_prompt_pos
universal_dependencies_fr_partut_fr_prompt_pos
Summary
universal_dependencies_fr_partut_fr_prompt_pos is a subset of the Dataset of French Prompts (DFP).It contains 21,420 rows that can be used for a part-of-speech task.The original data (without prompts) comes from the dataset universal_dependencies where only the French parput split has been kept.A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the same format… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/universal_dependencies_fr_partut_fr_prompt_pos.caire-universal
Dataset Card for "caire-universal"
More Information needed
universal_dependencies_fr_gsd_fr_prompt_pos
universal_dependencies_fr_gsd_fr_prompt_pos
Summary
universal_dependencies_fr_gsd_fr_prompt_pos is a subset of the Dataset of French Prompts (DFP).It contains 343,161 rows that can be used for a part-of-speech task.The original data (without prompts) comes from the dataset universal_dependencies where only the French gsd split has been kept.A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the same format as the… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/universal_dependencies_fr_gsd_fr_prompt_pos.universal_dependencies_fr_sequoia_fr_prompt_pos
universal_dependencies_fr_sequoia_fr_prompt_pos
Summary
universal_dependencies_fr_sequoia_fr_prompt_pos is a subset of the Dataset of French Prompts (DFP).It contains 27,804 rows that can be used for a part-of-speech task.The original data (without prompts) comes from the dataset universal_dependencies where only the French sequoia split has been kept.A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the same… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/universal_dependencies_fr_sequoia_fr_prompt_pos.universal_nerRe-upload of original dataset https://huggingface.co/datasets/universalner/universal_ner, with additional fields:
marked_text: The original text with marked entities in pair tags. Example:
What if <ORG>Google</ORG> Morphed Into GoogleOS?
entities: List of entities with type, text, and start (chracter offset).
universal-logo-detectoruniversalner-instructionuniversal_dependencies_fr_pud_fr_prompt_pos
universal_dependencies_fr_pud_fr_prompt_pos
Summary
universal_dependencies_fr_pud_fr_prompt_pos is a subset of the Dataset of French Prompts (DFP).It contains 21,000 rows that can be used for a part-of-speech task.The original data (without prompts) comes from the dataset universal_dependencies where only the French pud split has been kept.A list of prompts (see below) was then applied in order to build the input and target columns and thus obtain the same format as the… See the full description on the dataset page: https://huggingface.co/datasets/CATIE-AQ/universal_dependencies_fr_pud_fr_prompt_pos.universal_pipeline_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "dual_arm_robot",
"total_episodes": 1,
"total_frames": 3105,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Sraghvi/universal_pipeline_test.Universal-Magicoder-Evol-Instruct-110K
Dataset Card for "Universal-Magicoder-Evol-Instruct-110K"
Magicoder-Evol-Instruct-110K reformatted to in the universal data format.
Universal-Pure-Dove
Dataset Card for "Universal-Pure-Dove"
More Information needed
Universal_ner_chatmlasia-owid-universal-suffrage-lexical
Universal Suffrage Lexical | Asia (Our World in Data)
🌏 8,453 observations · 49 Asia countries · 1789–2025 · Repackaged by Electric Sheep Asia
TL;DR
This dataset contains 8,453 observations of Universal Suffrage Lexical data across 49 Asia countries, spanning 1789–2025.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Universal Suffrage Lexical
Geographic coverage
49 Asia… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-owid-universal-suffrage-lexical.he-universal_morphologieseurope-owid-universal-suffrage-men-lexical
Universal Suffrage Men Lexical | Europe (Our World in Data)
🇪🇺 7,863 observations · 44 Europe countries · 1789–2025 · Repackaged by Electric Sheep Europe
TL;DR
This dataset contains 7,863 observations of Universal Suffrage Men Lexical data across 44 Europe countries, spanning 1789–2025.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Universal Suffrage Men Lexical
Geographic… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-owid-universal-suffrage-men-lexical.europe-owid-universal-suffrage-women-lexical
Universal Suffrage Women Lexical | Europe (Our World in Data)
🇪🇺 7,863 observations · 44 Europe countries · 1789–2025 · Repackaged by Electric Sheep Europe
TL;DR
This dataset contains 7,863 observations of Universal Suffrage Women Lexical data across 44 Europe countries, spanning 1789–2025.
About the source
Source: Our World in Data
Publisher: Our World in Data
License: cc-by-4.0
Topic: Universal Suffrage Women Lexical
Geographic… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepeurope/europe-owid-universal-suffrage-women-lexical.universal_privacyUniversal-Reasoning-DatasetPlantNet-300K-Universal
مستودع Pl@ntNet-300K الشامل (Universal Plant Identification)
📌 نبذة تعريفيّة عن المستودع
يحتوي هذا المستودع على مجموعة بيانات Pl@ntNet-300K، وهي واحدة من أضخم المجموعات العالمية لتصنيف النباتات. تم دمجها في مشروع "الخبير الزراعي اليمني" لتكون بمثابة "المرحلة الصفرية" في النظام؛ حيث تسمح للتطبيق بالتعرف على نوع الشجرة أو النبات (من بين أكثر من 1000 نوع) قبل البدء في تشخيص الأمراض أو الآفات.
📂 هيكل البيانات (Data Structure)
البيانات مرفوعة بصيغة… See the full description on the dataset page: https://huggingface.co/datasets/Omarrs11/PlantNet-300K-Universal.
