datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
UDM_cleaned_docs
UDM cleaned docs
6,029,052 web pages reduced to just their mathematical content, extracted verbatim by oklenAI/udm_doc_extract_qwen3.5_2B — a 2B model distilled from GPT-5.6.
Every row is model output, not human-curated text. The extract field is what the model returned for that page; the source page text is not included. Read Two repetition flags below before filtering — the obvious flag is not the one you want.
How it was built
step
pages… See the full description on the dataset page: https://huggingface.co/datasets/oklenAI/UDM_cleaned_docs.UDA-QA
Dataset Card for Dataset Name
[NIPS-2024] UDA: A Benchmark Suite for Retrieval Augmented Generation in Real-world Document Analysis (https://arxiv.org/abs/2406.15187)
UDA (Unstructured Document Analysis) is a benchmark suite for Retrieval Augmented Generation (RAG) in real-world document analysis.
Each entry in the UDA dataset is organized as a document-question-answer triplet, where a question is raised from the document, accompanied by a corresponding ground-truth answer.
The… See the full description on the dataset page: https://huggingface.co/datasets/qinchuanhui/UDA-QA.swe-bench-coding-tasks
SWE-Bench Dataset - 8,712 files
The dataset comprises 8,712 files across 6 programming languages, featuring verified tasks and benchmarks for evaluating coding agents and language models. It supports coding agents, language models, and developer tools with verified benchmark scores and multi-language test sets. - Get the data
Dataset characteristics:
Characteristic
Data
Description
An extended benchmark of real-world software engineering tasks with enhanced… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/swe-bench-coding-tasks.task584_udeps_eng_fine_pos_tagging
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task584_udeps_eng_fine_pos_tagging
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task584_udeps_eng_fine_pos_tagging.task583_udeps_eng_coarse_pos_tagging
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task583_udeps_eng_coarse_pos_tagging
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task583_udeps_eng_coarse_pos_tagging.taskweft-fbd-udon-train
taskweft-fbd-udon-train
Intents and the IEC 61131-3 Function Block Diagrams that carry them out, as an
EditScore-shaped corpus: one root row per intent, three candidates per row (rank1 the
reference diagram, rank3 one that compiles and does the wrong thing, rank5 one the
compiler refuses), and one score row per candidate from the compiler and a runner that performed the plan. Every row is
constructed from a template and a seed, so the labels are true by construction and
the… See the full description on the dataset page: https://huggingface.co/datasets/chibifire/taskweft-fbd-udon-train.redsm5-sample
⚕️💬 ReDSM5: A Reddit Dataset for DSM-5 Depression Detection
📝 Dataset Summary
ReDSM5-Sample is a public, fully paraphrased, and anonymized sample of the ReDSM5 dataset.It contains 25 entries from the original dataset, each one rewritten to ensure no original user content is present and full privacy is maintained.
Each sample includes sentence-level clinical annotations for presence/absence of DSM-5 major depressive episode symptoms, together with an expert-written… See the full description on the dataset page: https://huggingface.co/datasets/irlab-udc/redsm5-sample.redsm5
⚕️💬 ReDSM5: A Reddit Dataset for DSM-5 Depression Detection
ℹ️ Looking for a quick preview?A fully paraphrased, anonymized sample with 25 entries is publicly available on the Hugging Face Hub — no user agreement required!
🚦 Access Conditions
This dataset is gated. To obtain access, please complete the access request form at ReDSM5 Agreement Form and submit it via email to eliseo.bao@udc.es. Your request will be reviewed and you’ll receive approval or further… See the full description on the dataset page: https://huggingface.co/datasets/irlab-udc/redsm5.PromptDataset-v2-Complete
Prompt Dataset v2 Complete
Dataset Description
A comprehensive collection of prompts for LLM fine-tuning and testing, including adversarial examples, jailbreaks, and safety test cases.
Dataset Statistics
Total Samples: 182,473
Training Samples: 179,378
Evaluation Samples: 3,095
Train/Eval Ratio: 58.0:1
Data Sources
The dataset is compiled from the following sources:
jailbreak_prompts_2023_12_25.csv
qualifire/prompt-injections-benchmark… See the full description on the dataset page: https://huggingface.co/datasets/UdayGattu23/PromptDataset-v2-Complete.alpaca_data_galician
Galician version of alpaca_data.json
This is a Galician-translated with Python package googletranslatepy version of the Stanford alpaca_data.json dataset. Our working notes are available here.
Dataset Structure
The dataset contains 52K instruction-following elements in a JSON file with a list of dictionaries. Each dictionary contains the following fields:
instruction: str, describes the task the model should perform. Each of the 52K instructions is unique.
input: str… See the full description on the dataset page: https://huggingface.co/datasets/irlab-udc/alpaca_data_galician.civilian-hazard-lifecycle-instruct
Hazards Dataset
This dataset contains comprehensive information about various hazards, including preparation, reaction, and recovery steps. It is designed for fine-tuning Large Language Models (LLMs) on safety and emergency response procedures.
Dataset Structure
The dataset is provided in a format compatible with the Hugging Face datasets library.
Features
hazard_type (string): The high-level category of the hazard (e.g., "Wildfire", "Active Shooter").
phase… See the full description on the dataset page: https://huggingface.co/datasets/Uday/civilian-hazard-lifecycle-instruct.LLM-Text-Generation-Dataset
Generated Text Dataset - 4 Millions+ Logs
Dataset comprises 4 million+ logs of synthetic texts generated by large language models (LLMs) across 32 languages, leveraging 3 different GPT models for diverse, high-quality training data. Designed for text generation tasks, language model training, and NLP applications, supporting generative AI and text classification.- Get the data
Dataset characteristics:
Characteristic
Data
Description
Generated texts to achieve… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/LLM-Text-Generation-Dataset.UD_Thai-PUD-prompt
Dataset Card for "UD_Thai-PUD-prompt"
This dataset is the test set from the Parallel Universal Dependencies (PUD) treebanks.
See more https://github.com/UniversalDependencies/UD_Thai-PUD
Template
Inputs: จงสร้างประโยคตามโครงสร้าง {pos}:
Targets: Thai sentence
pos: All tag
Source code for create dataset: https://github.com/PyThaiNLP/support-aya-datasets/blob/main/pos/ud_pud_thai.ipynb
Multilingual-UD-Bronze-Extension-Delex
Delexicalized Universal Dependencies (UD 2.17 + Silver v2)
A multilingual language-modelling dataset built from Universal Dependencies 2.17
gold treebanks and UDPipe-parsed silver data. Lexemes (NOUN, VERB, ADJ, ADV, PROPN, NUM)
are replaced with structured tokens encoding part-of-speech, morphological features,
and a dependency-frame cluster label. Function words are kept as lowercased surface
forms, namespaced by language.
Languages
ar, de, en, es, eu, fi, hi, hy, id… See the full description on the dataset page: https://huggingface.co/datasets/NotQuiteHereYet/Multilingual-UD-Bronze-Extension-Delex.maritime-sft-mixed-formats
Maritime SFT Mixed Formats Dataset
High-quality supervised fine-tuning (SFT) dataset for the maritime domain,
generated from maritime books and technical documents using GLM-4.7 via NVIDIA NIM API.
Dataset Configs
Config
Format
Description
alpaca
Instruction / Input / Output
Standard Alpaca SFT format
chat
System / User / Assistant
ChatML multi-turn format
rag
Context / Question / Answer
RAG triad format for retrieval-augmented fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/uday568/maritime-sft-mixed-formats.parameter-golf-caseops-v1-webish-ud24
parameter-golf CaseOps v1 webish ud24
This dataset repo contains a retokenized CaseOps-style dataset and tokenizer for
parameter-golf experiments.
Contents:
tokenizers/fineweb_8192_bpe_lossless_caps_caseops_v1_webish_ud24.model
tokenizers/fineweb_8192_bpe_lossless_caps_caseops_v1_webish_ud24.vocab
datasets/fineweb10B_sp8192_lossless_caps_caseops_v1_webish_ud24/
fineweb_train_000000.bin through fineweb_train_000079.bin
fineweb_val_000000.bin
fineweb_val_bytes_000000.bin
Notes:… See the full description on the dataset page: https://huggingface.co/datasets/romeerp/parameter-golf-caseops-v1-webish-ud24.
