CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01oklenAI /UDM_cleaned_docs UDM cleaned docs 6,029,052 web pages reduced to just their mathematical content, extracted verbatim by oklenAI/udm_doc_extract_qwen3.5_2B — a 2B model distilled from GPT-5.6. Every row is model output, not human-curated text. The extract field is what the model returned for that page; the source page text is not included. Read Two repetition flags below before filtering — the obvious flag is not the one you want. How it was built step pages… See the full description on the dataset page: https://huggingface.co/datasets/oklenAI/UDM_cleaned_docs.tabulartext-generation1M<n<10M0 likes681 downloads26d agoHugging Face02qinchuanhui /UDA-QA Dataset Card for Dataset Name [NIPS-2024] UDA: A Benchmark Suite for Retrieval Augmented Generation in Real-world Document Analysis (https://arxiv.org/abs/2406.15187) UDA (Unstructured Document Analysis) is a benchmark suite for Retrieval Augmented Generation (RAG) in real-world document analysis. Each entry in the UDA dataset is organized as a document-question-answer triplet, where a question is raised from the document, accompanied by a corresponding ground-truth answer. The… See the full description on the dataset page: https://huggingface.co/datasets/qinchuanhui/UDA-QA.textquestion-answering10K<n<100K6 likes483 downloads2y agoHugging Face03ud-nlp /swe-bench-coding-tasks SWE-Bench Dataset - 8,712 files The dataset comprises 8,712 files across 6 programming languages, featuring verified tasks and benchmarks for evaluating coding agents and language models. It supports coding agents, language models, and developer tools with verified benchmark scores and multi-language test sets. - Get the data Dataset characteristics: Characteristic Data Description An extended benchmark of real-world software engineering tasks with enhanced… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/swe-bench-coding-tasks.texttext-generationn<1K0 likes340 downloads1y agoHugging Face04Lots-of-LoRAs /task584_udeps_eng_fine_pos_tagging Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task584_udeps_eng_fine_pos_tagging Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task584_udeps_eng_fine_pos_tagging.texttext-generation1K<n<10K0 likes96 downloads2y agoHugging Face05Lots-of-LoRAs /task583_udeps_eng_coarse_pos_tagging Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task583_udeps_eng_coarse_pos_tagging Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task583_udeps_eng_coarse_pos_tagging.texttext-generation1K<n<10K0 likes95 downloads2y agoHugging Face06chibifire /taskweft-fbd-udon-train taskweft-fbd-udon-train Intents and the IEC 61131-3 Function Block Diagrams that carry them out, as an EditScore-shaped corpus: one root row per intent, three candidates per row (rank1 the reference diagram, rank3 one that compiles and does the wrong thing, rank5 one the compiler refuses), and one score row per candidate from the compiler and a runner that performed the plan. Every row is constructed from a template and a seed, so the labels are true by construction and the… See the full description on the dataset page: https://huggingface.co/datasets/chibifire/taskweft-fbd-udon-train.tabulartext-generation10K<n<100K0 likes76 downloads18d agoHugging Face07UdayGattu23 /PromptDataset-v2-Complete Prompt Dataset v2 Complete Dataset Description A comprehensive collection of prompts for LLM fine-tuning and testing, including adversarial examples, jailbreaks, and safety test cases. Dataset Statistics Total Samples: 182,473 Training Samples: 179,378 Evaluation Samples: 3,095 Train/Eval Ratio: 58.0:1 Data Sources The dataset is compiled from the following sources: jailbreak_prompts_2023_12_25.csv qualifire/prompt-injections-benchmark… See the full description on the dataset page: https://huggingface.co/datasets/UdayGattu23/PromptDataset-v2-Complete.texttext-generation100K<n<1M1 likes51 downloads1y agoHugging Face08irlab-udc /alpaca_data_galician Galician version of alpaca_data.json This is a Galician-translated with Python package googletranslatepy version of the Stanford alpaca_data.json dataset. Our working notes are available here. Dataset Structure The dataset contains 52K instruction-following elements in a JSON file with a list of dictionaries. Each dictionary contains the following fields: instruction: str, describes the task the model should perform. Each of the 52K instructions is unique. input: str… See the full description on the dataset page: https://huggingface.co/datasets/irlab-udc/alpaca_data_galician.texttext-generation10K<n<100K5 likes39 downloads2y agoHugging Face09ud-nlp /LLM-Text-Generation-Dataset Generated Text Dataset - 4 Millions+ Logs Dataset comprises 4 million+ logs of synthetic texts generated by large language models (LLMs) across 32 languages, leveraging 3 different GPT models for diverse, high-quality training data. Designed for text generation tasks, language model training, and NLP applications, supporting generative AI and text classification.- Get the data Dataset characteristics: Characteristic Data Description Generated texts to achieve… See the full description on the dataset page: https://huggingface.co/datasets/ud-nlp/LLM-Text-Generation-Dataset.texttext-generation1K<n<10K0 likes24 downloads1y agoHugging Face10pythainlp /UD_Thai-PUD-prompt Dataset Card for "UD_Thai-PUD-prompt" This dataset is the test set from the Parallel Universal Dependencies (PUD) treebanks. See more https://github.com/UniversalDependencies/UD_Thai-PUD Template Inputs: จงสร้างประโยคตามโครงสร้าง {pos}: Targets: Thai sentence pos: All tag Source code for create dataset: https://github.com/PyThaiNLP/support-aya-datasets/blob/main/pos/ud_pud_thai.ipynb texttext-generation1K<n<10K1 likes14 downloads3y agoHugging Face11NotQuiteHereYet /Multilingual-UD-Bronze-Extension-Delexgated Delexicalized Universal Dependencies (UD 2.17 + Silver v2) A multilingual language-modelling dataset built from Universal Dependencies 2.17 gold treebanks and UDPipe-parsed silver data. Lexemes (NOUN, VERB, ADJ, ADV, PROPN, NUM) are replaced with structured tokens encoding part-of-speech, morphological features, and a dependency-frame cluster label. Function words are kept as lowercased surface forms, namespaced by language. Languages ar, de, en, es, eu, fi, hi, hy, id… See the full description on the dataset page: https://huggingface.co/datasets/NotQuiteHereYet/Multilingual-UD-Bronze-Extension-Delex.texttext-generation1M<n<10M0 likes14 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.