universa
Datasets
All datasets matching “universa”universal_dependencies
Dataset Card (v2.0) for Universal Dependencies Treebank
Version 2.0.0 introduces significant improvements and breaking changes:
Parquet Format: faster loading with HuggingFace datasets >=4.0.0
MWT Support: New mwt field provides structured multi-word token information
Enhanced Security: No more trust_remote_code=True required
Separate Versioning: Loader version (2.0.0) distinct from UD data version (2.18)
Breaking Changes:
Token sequences now exclude MWT surface forms… See the full description on the dataset page: https://huggingface.co/datasets/universal-dependencies/universal_dependencies.universal-lesion-segmentation
Universal Lesion Segmentation Datasets
A collection of public medical imaging datasets for lesion segmentation in CT scans. These are the datasets exactly as downloaded from their original sources.
Datasets
This repository contains the following datasets:
CECT - Liver (primary). Luo J, Wang X, Zhang Y, et al. Comprehensive multi-phase three-dimensional contrast-enhanced CT imaging dataset for primary liver cancer. Scientific Data. 2025;12(1):768.… See the full description on the dataset page: https://huggingface.co/datasets/nielsRocholl/universal-lesion-segmentation.gta-data-files-universalquranic-universal-ayahs
Qur'anic Universal Ayahs
Qur'anic Universal Audio (QUA) is a project that unifies recitations on the internet and generates timing data using forced alignment — community-verified results and constantly expanding dataset.
This dataset pairs ayah by ayah audio with word-level timestamps, DigitalKhatt letter-animation timestamps, and waqf-aware segment data. Repeated words are preserved in text_uthmani and word_timestamps, so the row reflects what the reciter… See the full description on the dataset page: https://huggingface.co/datasets/QUD-Technologies/quranic-universal-ayahs.urgent26_track1_universal_seThe pre-simulated universal speech enhancement training and validation set of the ICASSP 2026 URGENT speech enhancement challenge, Track1.
Please check our GitHub Repo and webpage for more details.
How to use:
Use tar to decompress the dataset
cat ./urgent26_track2_se_dataset.tgz.* | tar xzv
The pre-simulated dataset can be loaded by the PreSimulatedDataset in the URGENT 2026 baseline code.
Directory structure:
.
├── data
│ ├── train_simulation # train set… See the full description on the dataset page: https://huggingface.co/datasets/lichenda/urgent26_track1_universal_se.gta-data-files-universal
