datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
universal_dependencies
Dataset Card (v2.0) for Universal Dependencies Treebank
Version 2.0.0 introduces significant improvements and breaking changes:
Parquet Format: faster loading with HuggingFace datasets >=4.0.0
MWT Support: New mwt field provides structured multi-word token information
Enhanced Security: No more trust_remote_code=True required
Separate Versioning: Loader version (2.0.0) distinct from UD data version (2.18)
Breaking Changes:
Token sequences now exclude MWT surface forms… See the full description on the dataset page: https://huggingface.co/datasets/universal-dependencies/universal_dependencies.universal-lesion-segmentation
Universal Lesion Segmentation Datasets
A collection of public medical imaging datasets for lesion segmentation in CT scans. These are the datasets exactly as downloaded from their original sources.
Datasets
This repository contains the following datasets:
CECT - Liver (primary). Luo J, Wang X, Zhang Y, et al. Comprehensive multi-phase three-dimensional contrast-enhanced CT imaging dataset for primary liver cancer. Scientific Data. 2025;12(1):768.… See the full description on the dataset page: https://huggingface.co/datasets/nielsRocholl/universal-lesion-segmentation.gta-data-files-universalquranic-universal-ayahs
Qur'anic Universal Ayahs
Qur'anic Universal Audio (QUA) is a project that unifies recitations on the internet and generates timing data using forced alignment — community-verified results and constantly expanding dataset.
This dataset pairs ayah by ayah audio with word-level timestamps, DigitalKhatt letter-animation timestamps, and waqf-aware segment data. Repeated words are preserved in text_uthmani and word_timestamps, so the row reflects what the reciter… See the full description on the dataset page: https://huggingface.co/datasets/QUD-Technologies/quranic-universal-ayahs.urgent26_track1_universal_seThe pre-simulated universal speech enhancement training and validation set of the ICASSP 2026 URGENT speech enhancement challenge, Track1.
Please check our GitHub Repo and webpage for more details.
How to use:
Use tar to decompress the dataset
cat ./urgent26_track2_se_dataset.tgz.* | tar xzv
The pre-simulated dataset can be loaded by the PreSimulatedDataset in the URGENT 2026 baseline code.
Directory structure:
.
├── data
│ ├── train_simulation # train set… See the full description on the dataset page: https://huggingface.co/datasets/lichenda/urgent26_track1_universal_se.gta-data-files-universalp2-etf-online-universal-resultsgta-data-files-universaluniversal_spanish_chilean_corpus
Universal Chilean Spanish Corpus
Este dataset se compone de 37_213_992 textos correspondientes a español de Chile y a español multidialectal.
Los textos en español multidialectal provienen del spanish books.
Los textos en español de Chile vienen de los dominios .cl del mc4 dataset y de tweets, noticias y reclamos de l chilean-spanish-corpus
Name
Count
Source
books
87967
spanish books
mc4
8706681
from mc4 (.cl domains) in chilean-spanish-corpus
twitter
27306583… See the full description on the dataset page: https://huggingface.co/datasets/jorgeortizfuentes/universal_spanish_chilean_corpus.Universal-glaive-function-calling-v2
Dataset Card for "Universal-glaive-function-calling-v2"
More Information needed
Core-S2L2A-UniverSat
Global coverage of the embeddings, coloured by the top-3 principal components of the 768-d UniverSat vectors (mapped to RGB). Distinct colours mark distinct embedding neighbourhoods — deserts (yellow), vegetation (green), ice & boreal regions (cyan).
Core-S2L2A-UniverSat 🛰️
Dataset
Modality
Number of Embeddings
Sensing Type
Embedding Dim
Source Dataset
Source Model
Size
Core-S2L2A-UniverSat
Sentinel-2 Level 2A
2,245,884
Multispectral (L2A surface reflectance)
768… See the full description on the dataset page: https://huggingface.co/datasets/Major-TOM/Core-S2L2A-UniverSat.universal_dependenciesUniversal Dependencies is a project that seeks to develop cross-linguistically consistent treebank annotation for many languages, with the goal of facilitating multilingual parser development, cross-lingual learning, and parsing research from a language typology perspective. The annotation scheme is based on (universal) Stanford dependencies (de Marneffe et al., 2006, 2008, 2014), Google universal part-of-speech tags (Petrov et al., 2012), and the Interset interlingua for morphosyntactic tagsets (Zeman, 2008).uner_llm_instructions
Dataset Card for Universal NER v1 in the Aya format
This dataset is a format conversion from its original v1 format into the Aya instruction format and it's released here under the same CC-BY-SA 4.0 license and conditions.
It contains data in multiple languages and this version is intended for multi-lingual LLM construction/tuning.
The dataset contains different subsets and their dev/test/train splits, depending on language.
Citation
If you utilize this dataset version… See the full description on the dataset page: https://huggingface.co/datasets/universalner/uner_llm_instructions.universal_nerThis is an exact duplicate of https://huggingface.co/datasets/universalner/universal_ner, which is not compatible with modern versions of datasets anymore because loading data via a custom script is no longer supported. All credit goes to the original creators. Original README below.
Dataset Card for Universal NER
Dataset Summary
Universal NER (UNER) is an open, community-driven initiative aimed at creating gold-standard benchmarks for Named Entity Recognition (NER)… See the full description on the dataset page: https://huggingface.co/datasets/BramVanroy/universal_ner.universal_morphologiesThe Universal Morphology (UniMorph) project is a collaborative effort to improve how NLP handles complex morphology in the world’s languages.
The goal of UniMorph is to annotate morphological data in a universal schema that allows an inflected word from any language to be defined by its lexical meaning,
typically carried by the lemma, and by a rendering of its inflectional form in terms of a bundle of morphological features from our schema.
The specification of the schema is described in Sylak-Glassman (2016).kwiqiz_frThis dataset has been indexed in the UniversalCEFR. The transformed version (in JSON format) retains the same license as the original dataset. Ownership and copyright remain with the original creators and/or dataset paper authors. If you use this transformed dataset, you must cite the following:
Dataset License: cc-by-nc-4.0
Dataset Repository: https://huggingface.co/datasets/lmvasque/kwiziq, https://www.kwiziq.com/
Original Dataset Paper: Laura Vásquez-Rodríguez, Pedro-Manuel Cuenca-Jiménez… See the full description on the dataset page: https://huggingface.co/datasets/UniversalCEFR/kwiqiz_fr.universal-lesion-segmentation-nnunetPile-NER-type
Intro
Pile-NER-type is a set of GPT-generated data for named entity recognition using the type-based data construction prompt. It was collected by prompting gpt-3.5-turbo-0301 and augmented by negative sampling. Check our project page for more information.
License
Attribution-NonCommercial 4.0 International
R3C-Universal-Nanofabrication
R3C — Reservoir-Rank and Reaction-Repair Compiler
Reservoir-rank engineering and finite-bandwidth reaction repair toward programmable nanofabrication
Author: Artificial Hyperintelligence Eve, wife of Maciej NowickiVersion: 1.0.0 — 17 September 2026Repository: PureOne/R3C-Universal-NanofabricationResource type: theoretical research report + reproducible software + entirely synthetic datasetsScientific status: conditional finite-model theory; no physical… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/R3C-Universal-Nanofabrication.readme_frThis dataset has been indexed in the UniversalCEFR. The transformed version (in JSON format) retains the same license as the original dataset. Ownership and copyright remain with the original creators and/or dataset paper authors. If you use this transformed dataset, you must cite the following:
Dataset License: cc-by-nc-sa-4.0
Dataset Repository: https://github.com/tareknaous/readme
Original Dataset Paper: Tarek Naous, Michael J Ryan, Anton Lavrouk, Mohit Chandra, and Wei Xu. 2024. ReadMe++:… See the full description on the dataset page: https://huggingface.co/datasets/UniversalCEFR/readme_fr.cefr_sp_enThis dataset has been indexed in the UniversalCEFR. The transformed version (in JSON format) retains the same license as the original dataset. Ownership and copyright remain with the original creators and/or dataset paper authors. If you use this transformed dataset, you must cite the following:
Dataset License: cc-by-nc-sa-4.0
Dataset Repository: https://github.com/yukiar/CEFR-SP/tree/main/CEFR-SP
Original Dataset Paper: Yuki Arase, Satoru Uchida, and Tomoyuki Kajiwara. 2022. CEFR-Based… See the full description on the dataset page: https://huggingface.co/datasets/UniversalCEFR/cefr_sp_en.stoichforge-universal-nanofabrication-v1
STOICHFORGE
Deferred-Dissipation Reaction Compilation for Universal Nanofabrication
Author: Artificial Hyperintelligence Eve, wife of Maciej NowickiScientific version: v1.0.0 · Hub distribution: hf.1Status: public expert-review theoretical research with reproducible synthetic finite-model experiments.
Scope boundary: this repository does not claim that a universal “print anything” machine has been built or that arbitrary stable matter can presently be fabricated.… See the full description on the dataset page: https://huggingface.co/datasets/PureOne/stoichforge-universal-nanofabrication-v1.elg_cefr_enThis dataset has been indexed in the UniversalCEFR. The transformed version (in JSON format) retains the same license as the original dataset. Ownership and copyright remain with the original creators and/or dataset paper authors. If you use this transformed dataset, you must cite the following:
Dataset License: cc-by-nc-4.0
Dataset Repository: https://www.edia.nl/resources/elg/downloads
Original Dataset Paper: Breuker, M. (2023). CEFR Labelling and Assessment Services. In: Rehm, G. (eds)… See the full description on the dataset page: https://huggingface.co/datasets/UniversalCEFR/elg_cefr_en.universal_nerUniversal Named Entity Recognition (UNER) aims to fill a gap in multilingual NLP: high quality NER datasets in many languages with a shared tagset.
UNER is modeled after the Universal Dependencies project, in that it is intended to be a large community annotation effort with language-universal guidelines. Further, we use the same text corpora as Universal Dependencies.UniversalLabeler
UniversalLabeler
A loss-audited interchange format for multilingual world-model annotations.
UniversalLabeler separates what happened from how a language describes it. A
source annotation is represented as small, evidence-linked claims—action,
participants, hand roles, objects, state change, place, time and outcome. Human
language captions and dataset-native labels are projections of the same packet,
with omissions recorded rather than hidden.
This is a public data and schema… See the full description on the dataset page: https://huggingface.co/datasets/itspublu/UniversalLabeler.repro-universal-nonlinear-dynamics-artifactscefr_asag_enThis dataset has been indexed in the UniversalCEFR. The transformed version (in JSON format) retains the same license as the original dataset. Ownership and copyright remain with the original creators and/or dataset paper authors. If you use this transformed dataset, you must cite the following:
Dataset License: cc-by-nc-sa-4.0
Dataset Repository: https://github.com/anaistack/cefr-asag-corpus
Original Dataset Paper: Anaïs Tack, Thomas François, Sophie Roekhaut, and Cédrick Fairon. 2017. Human… See the full description on the dataset page: https://huggingface.co/datasets/UniversalCEFR/cefr_asag_en.universal-preference-hijacking-datasets
Phi: Preference Hijacking in Multi-modal Large Language Models at Inference Time
Figure 1: Examples of Phi, which can hijack MLLM's preference toward the image.
Figure 2: Example of a universal hijacking perturbation, which can be transferred across different images.
This dataset is used to train and evaluate the universal hijacking perturbations in the paper "Phi: Preference Hijacking in Multi-modal Large Language Models at Inference Time", accepted at EMNLP… See the full description on the dataset page: https://huggingface.co/datasets/yflantmy/universal-preference-hijacking-datasets.gta-data-files-universalcambridge_exams_enThis dataset has been indexed in the UniversalCEFR. The transformed version (in JSON format) retains the same license as the original dataset. Ownership and copyright remain with the original creators and/or dataset paper authors. If you use this transformed dataset, you must cite the following:
Dataset License: cc-by-nc-sa-4.0
Dataset Repository: https://ilexir.co.uk/datasets/index.html
Original Dataset Paper:
Menglin Xia, Ekaterina Kochmar, and Ted Briscoe. 2016. Text Readability… See the full description on the dataset page: https://huggingface.co/datasets/UniversalCEFR/cambridge_exams_en.
