datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
udio-POSTPROCESS-3fd79cfbUCI_HARdk-udvalg-skole-boern-unge
Udvalgsbilag — børn, skole og unge i alle kommuner
Møder, dagsordenspunkter og bilag fra de kommunale fagudvalg for børn,
skole og unge i hele landet. Det er her, tildelingsmodeller vedtages,
besparelser udmøntes og skolefordelte budgettal offentliggøres.
Bilagene er den værdifulde del. Et udvalgsbilag indeholder ofte den tabel med
beløb pr. skole, som ellers ikke publiceres nogen steder — og som er den
eneste måde at sammenligne skolers tildeling på tværs af kommuner.… See the full description on the dataset page: https://huggingface.co/datasets/Skoleatlas/dk-udvalg-skole-boern-unge.udio-POSTPROCESS-49a37219ud-treebank-tokens
Dataset Card for Dataset Name
Dataset Summary
This is a subset of the Universal Dependencies Treebanks dataset (version 2.12) which only contains raw sentences and their corresponding tokenized form.
This dataset is licensed under the Universal Dependencies v2.10 License Agreement. The user is reminded that some subsets only allow non-commercial use. Each row contains the license of the dataset it originated from, and a full listing of all included subsets and… See the full description on the dataset page: https://huggingface.co/datasets/tokenizer-eval/ud-treebank-tokens.udio-POSTPROCESS-8fe366afudio-POSTPROCESS-ec30566fudio-POSTPROCESS-d6ca28d5udio-POSTPROCESS-4a3a19c6Phonemized-UD
Phoneme-UD: A Multilingual Phonemized Universal Dependencies Corpus for 34+ Languages
G2P+ Phonemizer
We use G2P+ to phonemize Universal Dependencies. Here is an example usage:
# Install required packages
!apt-get install -y espeak-ng
!pip install phonemizer g2p-plus
# Set the environment variable from Python
import os
os.environ["PHONEMIZER_ESPEAK_LIBRARY"] = "/usr/lib/x86_64-linux-gnu/libespeak-ng.so.1"
# Now run your transcription
from g2p_plus import… See the full description on the dataset page: https://huggingface.co/datasets/suchirsalhan/Phonemized-UD.EA-UDudio-POSTPROCESS-7caf4357uds-governance-receipts
Part of the SZL Holdings governed estate — claims are designed to carry checkable receipts. Verification proves integrity & origin, never accuracy or performance.
UDS Governance Receipts — Decision Audit Log
Doctrine v11 LOCKED. No marketing. Every number resolves to a CI log, a Lean proof, or a Zenodo DOI.
Append-only log of DSSE-signed governance decision receipts for the Unified Deployment Substrate (UDS) mesh. Each record captures:… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/uds-governance-receipts.udio-POSTPROCESS-6f8468d1UDM_cleaned_docs
UDM cleaned docs
6,029,052 web pages reduced to just their mathematical content, extracted verbatim by oklenAI/udm_doc_extract_qwen3.5_2B — a 2B model distilled from GPT-5.6.
Every row is model output, not human-curated text. The extract field is what the model returned for that page; the source page text is not included. Read Two repetition flags below before filtering — the obvious flag is not the one you want.
How it was built
step
pages… See the full description on the dataset page: https://huggingface.co/datasets/oklenAI/UDM_cleaned_docs.udio-POSTPROCESS-1cee75c6UDR_CosmosQA
Dataset Card for "UDR_CosmosQA"
More Information needed
udio-POSTPROCESS-51913732udio-POSTPROCESS-3655f4ecudio-POSTPROCESS-126e7471udtrees
Dataset Card for "udtrees"
More Information needed
udio-POSTPROCESS-21d9538cuds-spans-receipts
Part of the SZL Holdings governed estate — claims are designed to carry checkable receipts. Verification proves integrity & origin, never accuracy or performance.
UDS Spans Receipts — OTel Governance Audit Log
Doctrine v11 LOCKED. No marketing. Every number resolves to a CI log, a Lean proof, or a Zenodo DOI.
Append-only audit log of DSSE-signed OpenTelemetry spans emitted by the UDS mesh governance layer. Each span record includes: operation… See the full description on the dataset page: https://huggingface.co/datasets/SZLHOLDINGS/uds-spans-receipts.UDA-QA
Dataset Card for Dataset Name
[NIPS-2024] UDA: A Benchmark Suite for Retrieval Augmented Generation in Real-world Document Analysis (https://arxiv.org/abs/2406.15187)
UDA (Unstructured Document Analysis) is a benchmark suite for Retrieval Augmented Generation (RAG) in real-world document analysis.
Each entry in the UDA dataset is organized as a document-question-answer triplet, where a question is raised from the document, accompanied by a corresponding ground-truth answer.
The… See the full description on the dataset page: https://huggingface.co/datasets/qinchuanhui/UDA-QA.udio-POSTPROCESS-f9a2b58dxtreme-r-udpos
UDPOS of XTREME-R
Generated by build_parquet.py
XTREME-R: Towards More Challenging and Nuanced Multilingual Evaluation
https://arxiv.org/abs/2104.07412
datasets
UdonPred datasets
Per-target protein intrinsic-disorder datasets for UdonPred: train/valid/test as jsonl ({id, y, x_0}) and FASTA, plus precomputed per-pLM embeddings under <target>/embeddings/<plm>/<split>.h5 (keyed by jsonl id).
swe-bench-coding-tasks
SWE-Bench Dataset - 8,712 files
The dataset comprises 8,712 files across 6 programming languages, featuring verified tasks and benchmarks for evaluating coding agents and language models. It supports coding agents, language models, and developer tools with verified benchmark scores and multi-language test sets. - Get the data
Dataset characteristics:
Characteristic
Data
Description
An extended benchmark of real-world software engineering tasks with enhanced… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/swe-bench-coding-tasks.UDR_COPA
Dataset Card for "UDR_COPA"
More Information needed
udemy-corpus
