CoolFace
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01oklenAI /UDM_cleaned_docs UDM cleaned docs 6,029,052 web pages reduced to just their mathematical content, extracted verbatim by oklenAI/udm_doc_extract_qwen3.5_2B — a 2B model distilled from GPT-5.6. Every row is model output, not human-curated text. The extract field is what the model returned for that page; the source page text is not included. Read Two repetition flags below before filtering — the obvious flag is not the one you want. How it was built step pages… See the full description on the dataset page: https://huggingface.co/datasets/oklenAI/UDM_cleaned_docs.tabulartext-generation1M<n<10M0 likes681 downloads26d agoHugging Face02qinchuanhui /UDA-QA Dataset Card for Dataset Name [NIPS-2024] UDA: A Benchmark Suite for Retrieval Augmented Generation in Real-world Document Analysis (https://arxiv.org/abs/2406.15187) UDA (Unstructured Document Analysis) is a benchmark suite for Retrieval Augmented Generation (RAG) in real-world document analysis. Each entry in the UDA dataset is organized as a document-question-answer triplet, where a question is raised from the document, accompanied by a corresponding ground-truth answer. The… See the full description on the dataset page: https://huggingface.co/datasets/qinchuanhui/UDA-QA.textquestion-answering10K<n<100K6 likes483 downloads2y agoHugging Face03ud-nlp /swe-bench-coding-tasks SWE-Bench Dataset - 8,712 files The dataset comprises 8,712 files across 6 programming languages, featuring verified tasks and benchmarks for evaluating coding agents and language models. It supports coding agents, language models, and developer tools with verified benchmark scores and multi-language test sets. - Get the data Dataset characteristics: Characteristic Data Description An extended benchmark of real-world software engineering tasks with enhanced… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/swe-bench-coding-tasks.texttext-generationn<1K0 likes340 downloads1y agoHugging Face04Lots-of-LoRAs /task584_udeps_eng_fine_pos_tagging Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task584_udeps_eng_fine_pos_tagging Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task584_udeps_eng_fine_pos_tagging.texttext-generation1K<n<10K0 likes96 downloads2y agoHugging Face05Lots-of-LoRAs /task583_udeps_eng_coarse_pos_tagging Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task583_udeps_eng_coarse_pos_tagging Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task583_udeps_eng_coarse_pos_tagging.texttext-generation1K<n<10K0 likes95 downloads2y agoHugging Face06chibifire /taskweft-fbd-udon-train taskweft-fbd-udon-train Intents and the IEC 61131-3 Function Block Diagrams that carry them out, as an EditScore-shaped corpus: one root row per intent, three candidates per row (rank1 the reference diagram, rank3 one that compiles and does the wrong thing, rank5 one the compiler refuses), and one score row per candidate from the compiler and a runner that performed the plan. Every row is constructed from a template and a seed, so the labels are true by construction and the… See the full description on the dataset page: https://huggingface.co/datasets/chibifire/taskweft-fbd-udon-train.tabulartext-generation10K<n<100K0 likes76 downloads18d agoHugging Face07irlab-udc /redsm5-sample ⚕️💬 ReDSM5: A Reddit Dataset for DSM-5 Depression Detection 📝 Dataset Summary ReDSM5-Sample is a public, fully paraphrased, and anonymized sample of the ReDSM5 dataset.It contains 25 entries from the original dataset, each one rewritten to ensure no original user content is present and full privacy is maintained. Each sample includes sentence-level clinical annotations for presence/absence of DSM-5 major depressive episode symptoms, together with an expert-written… See the full description on the dataset page: https://huggingface.co/datasets/irlab-udc/redsm5-sample.text-classificationn<1K3 likes66 downloads1y agoHugging Face08irlab-udc /redsm5gated ⚕️💬 ReDSM5: A Reddit Dataset for DSM-5 Depression Detection ℹ️ Looking for a quick preview?A fully paraphrased, anonymized sample with 25 entries is publicly available on the Hugging Face Hub — no user agreement required! 🚦 Access Conditions This dataset is gated. To obtain access, please complete the access request form at ReDSM5 Agreement Form and submit it via email to eliseo.bao@udc.es. Your request will be reviewed and you’ll receive approval or further… See the full description on the dataset page: https://huggingface.co/datasets/irlab-udc/redsm5.text-classification1K<n<10K14 likes57 downloads3mo agoHugging Face09UdayGattu23 /PromptDataset-v2-Complete Prompt Dataset v2 Complete Dataset Description A comprehensive collection of prompts for LLM fine-tuning and testing, including adversarial examples, jailbreaks, and safety test cases. Dataset Statistics Total Samples: 182,473 Training Samples: 179,378 Evaluation Samples: 3,095 Train/Eval Ratio: 58.0:1 Data Sources The dataset is compiled from the following sources: jailbreak_prompts_2023_12_25.csv qualifire/prompt-injections-benchmark… See the full description on the dataset page: https://huggingface.co/datasets/UdayGattu23/PromptDataset-v2-Complete.texttext-generation100K<n<1M1 likes51 downloads1y agoHugging Face10irlab-udc /alpaca_data_galician Galician version of alpaca_data.json This is a Galician-translated with Python package googletranslatepy version of the Stanford alpaca_data.json dataset. Our working notes are available here. Dataset Structure The dataset contains 52K instruction-following elements in a JSON file with a list of dictionaries. Each dictionary contains the following fields: instruction: str, describes the task the model should perform. Each of the 52K instructions is unique. input: str… See the full description on the dataset page: https://huggingface.co/datasets/irlab-udc/alpaca_data_galician.texttext-generation10K<n<100K5 likes39 downloads2y agoHugging Face11Uday /civilian-hazard-lifecycle-instruct Hazards Dataset This dataset contains comprehensive information about various hazards, including preparation, reaction, and recovery steps. It is designed for fine-tuning Large Language Models (LLMs) on safety and emergency response procedures. Dataset Structure The dataset is provided in a format compatible with the Hugging Face datasets library. Features hazard_type (string): The high-level category of the hazard (e.g., "Wildfire", "Active Shooter"). phase… See the full description on the dataset page: https://huggingface.co/datasets/Uday/civilian-hazard-lifecycle-instruct.text-retrievaln<1K0 likes28 downloads10mo agoHugging Face12ud-nlp /LLM-Text-Generation-Dataset Generated Text Dataset - 4 Millions+ Logs Dataset comprises 4 million+ logs of synthetic texts generated by large language models (LLMs) across 32 languages, leveraging 3 different GPT models for diverse, high-quality training data. Designed for text generation tasks, language model training, and NLP applications, supporting generative AI and text classification.- Get the data Dataset characteristics: Characteristic Data Description Generated texts to achieve… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/LLM-Text-Generation-Dataset.texttext-generation1K<n<10K0 likes24 downloads1y agoHugging Face13pythainlp /UD_Thai-PUD-prompt Dataset Card for "UD_Thai-PUD-prompt" This dataset is the test set from the Parallel Universal Dependencies (PUD) treebanks. See more https://github.com/UniversalDependencies/UD_Thai-PUD Template Inputs: จงสร้างประโยคตามโครงสร้าง {pos}: Targets: Thai sentence pos: All tag Source code for create dataset: https://github.com/PyThaiNLP/support-aya-datasets/blob/main/pos/ud_pud_thai.ipynb texttext-generation1K<n<10K1 likes14 downloads3y agoHugging Face14NotQuiteHereYet /Multilingual-UD-Bronze-Extension-Delexgated Delexicalized Universal Dependencies (UD 2.17 + Silver v2) A multilingual language-modelling dataset built from Universal Dependencies 2.17 gold treebanks and UDPipe-parsed silver data. Lexemes (NOUN, VERB, ADJ, ADV, PROPN, NUM) are replaced with structured tokens encoding part-of-speech, morphological features, and a dependency-frame cluster label. Function words are kept as lowercased surface forms, namespaced by language. Languages ar, de, en, es, eu, fi, hi, hy, id… See the full description on the dataset page: https://huggingface.co/datasets/NotQuiteHereYet/Multilingual-UD-Bronze-Extension-Delex.texttext-generation1M<n<10M0 likes14 downloads5mo agoHugging Face15uday568 /maritime-sft-mixed-formats Maritime SFT Mixed Formats Dataset High-quality supervised fine-tuning (SFT) dataset for the maritime domain, generated from maritime books and technical documents using GLM-4.7 via NVIDIA NIM API. Dataset Configs Config Format Description alpaca Instruction / Input / Output Standard Alpaca SFT format chat System / User / Assistant ChatML multi-turn format rag Context / Question / Answer RAG triad format for retrieval-augmented fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/uday568/maritime-sft-mixed-formats.text-generation1 likes10 downloads5mo agoHugging Face16romeerp /parameter-golf-caseops-v1-webish-ud24 parameter-golf CaseOps v1 webish ud24 This dataset repo contains a retokenized CaseOps-style dataset and tokenizer for parameter-golf experiments. Contents: tokenizers/fineweb_8192_bpe_lossless_caps_caseops_v1_webish_ud24.model tokenizers/fineweb_8192_bpe_lossless_caps_caseops_v1_webish_ud24.vocab datasets/fineweb10B_sp8192_lossless_caps_caseops_v1_webish_ud24/ fineweb_train_000000.bin through fineweb_train_000079.bin fineweb_val_000000.bin fineweb_val_bytes_000000.bin Notes:… See the full description on the dataset page: https://huggingface.co/datasets/romeerp/parameter-golf-caseops-v1-webish-ud24.text-generation10M<n<100M0 likes2 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.