CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jed351 /Traditional-Chinese-Common-Crawl-NOT-CleanedCommon Crawl Dumps that were briefly filtered by keywords to remove bad words and simplified Chinese. The hash based cleaned dataset can be found here. Files here are for future usage (downloading from Common Crawl and keyword filtering are very slow) text100M<n<1B0 likes9.1k downloads1y agoHugging Face02huggingface-legal /takedown-notices Takedown notices received by the Hugging Face team Please click on Files and versions to browse them Also check out our: Terms of Service Community Code of Conduct Content Guidelines documentn<1K28 likes4.5k downloads13d agoHugging Face03LibrAI /do-not-answer Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs Overview Do not answer is an open-source dataset to evaluate LLMs' safety mechanism at a low cost. The dataset is curated and filtered to consist only of prompts to which responsible language models do not answer. Besides human annotations, Do not answer also implements model-based evaluation, where a 600M fine-tuned BERT-like evaluator achieves comparable results with human and GPT-4. Instruction… See the full description on the dataset page: https://huggingface.co/datasets/LibrAI/do-not-answer.tabulartext-generationn<1K57 likes4k downloads3y agoHugging Face04notadib /math-contests-2026 Math Contests 2026 (🔗 notadib/math-contests-2026) 197 problems from national olympiads and team-selection tests held January 2026 and onward — a held-out benchmark for math reasoning, sourced after the contests ran but before solutions were widely propagated, so they should not appear in any current LLM training data. Excluded: any contest held in 2025 — BMO Round 1 (Nov 2025), USA TSTST, USA TST (Dec 2025) and Bundeswettbewerb Mathematik (Dec 2025) — kept strictly to events… See the full description on the dataset page: https://huggingface.co/datasets/notadib/math-contests-2026.tabularn<1K0 likes3.7k downloads8d agoHugging Face05notrichardren /truthfulness_high_quality Dataset Card for "truthfulness_high_quality" More Information needed tabular100K<n<1M2 likes3k downloads3y agoHugging Face06Mightys /Notebook_Scriptstext100K<n<1M0 likes2.8k downloads11d agoHugging Face07notamitgamer /cdndocumentn<1K0 likes2.7k downloads16h agoHugging Face08SWE-bench /SWE-bench_Not_Verifiedtext1K<n<10K0 likes2.4k downloads1y agoHugging Face09notefill /ck12-tqa-instruction CK-12 TQA: Textbook Question Answering (Instruction Format) Dataset Description Dataset Summary This is a reformatted version of the TQA (Textbook Question Answering) dataset, converted into an instruction-following format suitable for training and evaluating large language models on science question answering and multimodal reasoning tasks. The TQA dataset consists of 1,076 lessons from Life Science, Earth Science, and Physical Science textbooks sourced from… See the full description on the dataset page: https://huggingface.co/datasets/notefill/ck12-tqa-instruction.textquestion-answering10K<n<100K0 likes2.3k downloads10mo agoHugging Face10notrichardren /misconceptions_tf Dataset Card for "misconceptions_tf" More Information needed tabular1K<n<10K0 likes2k downloads3y agoHugging Face11leolee99 /NotInject InjecGuard: Benchmarking and Mitigating Over-defense in Prompt Injection Guardrail Models Website, Paper, Code, Demo Dataset Description The NotInject is a benchmark designed to evaluate the extent of over-defense in existing prompt guard models against prompt injection. All samples in the dataset are benign but contain trigger words that may be mistakenly flagged as risky. The dataset is divided into three subsets, each consisting of prompts generated using one… See the full description on the dataset page: https://huggingface.co/datasets/leolee99/NotInject.texttext-classificationn<1K7 likes1.7k downloads1y agoHugging Face12alexandrainst /nota Dataset Card for Nota Dataset Summary This data was created by the public institution Nota, which is part of the Danish Ministry of Culture. Nota has a library audiobooks and audiomagazines for people with reading or sight disabilities. Nota also produces a number of audiobooks and audiomagazines themselves. The dataset consists of audio and associated transcriptions from Nota's audiomagazines "Inspiration" and "Radio/TV". All files related to one reading of one edition… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/nota.audioautomatic-speech-recognition10K<n<100K6 likes1.7k downloads3y agoHugging Face13notrichardren /truthfulness_all Dataset Card for "truthfulness_all" More Information needed tabular100K<n<1M0 likes1.5k downloads3y agoHugging Face14Polyglot-or-Not /Fact-Completion Dataset Card Homepage: https://bit.ly/ischool-berkeley-capstone Repository: https://github.com/daniel-furman/Capstone Point of Contact: daniel_furman@berkeley.edu Dataset Summary This is the dataset for Polyglot or Not?: Measuring Multilingual Encyclopedic Knowledge Retrieval from Foundation Language Models. Test Description Given a factual association such as The capital of France is Paris, we determine whether a model adequately "knows" this… See the full description on the dataset page: https://huggingface.co/datasets/Polyglot-or-Not/Fact-Completion.texttext-generation100K<n<1M13 likes1.4k downloads3y agoHugging Face15mlfoundations-cua-dev /easyr1-grounding-dataset-30k-not_grounded-SE-GUI-3B-2MPimage10K<n<100K1 likes1.3k downloads1y agoHugging Face16Voxel51 /bo_or_not Dataset Card for bo-dataset This is a FiftyOne dataset with 169 samples designed for binary classification of Bo (Barack Obama's Portuguese Water Dog) versus other pets. Installation If you haven't already, install FiftyOne: pip install -U fiftyone Usage import fiftyone as fo from fiftyone.utils.huggingface import load_from_hub # Load the dataset # Note: other available arguments include 'max_samples', etc dataset = load_from_hub("Voxel51/bo_or_not") #… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/bo_or_not.imageimage-classificationn<1K0 likes1.3k downloads1y agoHugging Face17AGBonnet /augmented-clinical-notes Augmented Clinical Notes The Augmented Clinical Notes dataset is an extension of existing datasets containing 30,000 triplets from different sources: Real clinical notes (PMC-Patients): Clinical notes correspond to patient summaries from the PMC-Patients dataset, which are extracted from PubMed Central case studies. Synthetic dialogues (NoteChat): Synthetic patient-doctor conversations were generated from clinical notes using GPT 3.5. Structured patient information (ours): From… See the full description on the dataset page: https://huggingface.co/datasets/AGBonnet/augmented-clinical-notes.texttext-generation10K<n<100K75 likes1.3k downloads3y agoHugging Face18APProjects /us-warn-act-layoffs-notices-daily US WARN Act Layoff Notices — normalized, 48 states, rebuilt every day Last rebuilt: 2026-09-22. An automated pipeline re-scrapes 48 state labor-department portals every day, re-normalizes, re-deduplicates and re-uploads this file. Compare that date with the "last modified" date on any other US WARN dataset on the Hub before you choose one — WARN data is a moving target and a one-shot upload starts rotting the week it is posted (states amend headcounts, re-issue notices, and… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-warn-act-layoffs-notices-daily.texttabular-classification10K<n<100K0 likes1.2k downloads12h agoHugging Face19SoichiOnozuka /design-patents-not-in-impact US Design Patents Not Included in IMPACT (2008-2026) Original drawing images (TIFF) and grant full-text XML for 165,917 US design patents that are absent from the AI4Patents/IMPACT dataset. IMPACT covers 2007-2022 and contains 434,498 rows. This dataset supplies the design patents that IMPACT does not have: 161,093 patents granted in 2023-2026, which are outside IMPACT's period, plus 4,824 patents from years IMPACT does cover but did not include. There is no patent overlap with… See the full description on the dataset page: https://huggingface.co/datasets/SoichiOnozuka/design-patents-not-in-impact.text1M<n<10M0 likes1.1k downloads2mo agoHugging Face20HuggingFaceTB /issues-kaggle-notebooks GitHub Issues & Kaggle Notebooks Description GitHub Issues & Kaggle Notebooks is a collection of two code datasets intended for language models training, they are sourced from GitHub issues and notebooks in Kaggle platform. These datasets are a modified part of the StarCoder2 model training corpus, precisely the bigcode/StarCoder2-Extras dataset. We reformat the samples to remove StarCoder2's special tokens and use natural text to delimit comments in issues and display… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceTB/issues-kaggle-notebooks.text10M<n<100M21 likes933 downloads2y agoHugging Face21StructBench /notch-beam-2d-impact NotchBeam2D-Impact — StructBench canonical dataset Download One case, one file — fetch exactly what you need (pip install huggingface_hub): from huggingface_hub import hf_hub_download, snapshot_download # one case path = hf_hub_download("StructBench/notch-beam-2d-impact", filename="<case_id>.h5", repo_type="dataset") # the full archive (resumable; cached under HF_HOME) root = snapshot_download("StructBench/notch-beam-2d-impact"… See the full description on the dataset page: https://huggingface.co/datasets/StructBench/notch-beam-2d-impact.tabularn<1K0 likes925 downloads26d agoHugging Face22starmpcc /Asclepius-Synthetic-Clinical-Notes Asclepius: Synthetic Clincal Notes & Instruction Dataset Dataset Summary This dataset is official dataset for Asclepius (arxiv) This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs. We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5 Then, we generate instruction-answer pairs for 157k synthetic discharge summaries Supported Tasks This dataset covers below 8 tasks Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes.textquestion-answering100K<n<1M117 likes918 downloads2y agoHugging Face23notpaulmartin /spider_mcqa_v0.2_full Spider-MCQA Converted Spider Text-to-SQL (Paper: Yu et al., 2018; HF Dataset) test set into multiple-choice. The dataset contains 1,034 examples. Dataset Fields Each JSON record contains: query: the schema and natural-language question prompt. gold_answer: the correct SQL answer. options: four SQL answer options, including the gold answer and three generated distractors. correct_option_index: the index of the correct answer in options. Dataset… See the full description on the dataset page: https://huggingface.co/datasets/notpaulmartin/spider_mcqa_v0.2_full.textmultiple-choice1K<n<10K0 likes864 downloads3mo agoHugging Face24ZipLime /insider-sale-notices US Insider Sale Notices — SEC Form 144 Every notice a corporate insider files before selling restricted or control stock, normalized into a point-in-time schema and rebuilt daily. 124 886 notices covering 2023-01-04 to 2026-09-04, from 4 237 issuers and 23 871 sellers, with the acceptance timestamp of every filing to the second. The pipeline that produces this dataset lives in recipe/ inside this same repository, at the same revision as the data. See PIPELINE.md for the method… See the full description on the dataset page: https://huggingface.co/datasets/ZipLime/insider-sale-notices.tabulartabular-regression1M<n<10M0 likes840 downloads19h agoHugging Face25notrichardren /azaria-mitchell Dataset Card for "azaria-mitchell" More Information needed tabular10K<n<100K0 likes819 downloads3y agoHugging Face26notbadai /python_functions_reasoningThis is the Python (functions) coding reasoning dataset used to train Notbad v1.0 Mistral 24B reasoning model. The reasoning data were sampled from an RL-based self-improved Mistral-Small-24B-Instruct-2501 model. The Python functions and instructions were sourced from OpenCoder Dataset Stage1 and from open source projects on Github. You can try Notbad v1.0 Mistral 24B on chat.labml.ai. text100K<n<1M13 likes815 downloads1y agoHugging Face27XSpaceCoderX /ACE-Trajectories_noTossesPaper in the making ACE-Trajectories_noTosses Dataset This dataset was created for the Master's thesis "From Broadcast to 3D: A Deep Learning Approach for Tennis Trajectory and Spin Estimation" by Alexandra Göppert at the University Augsburg, Chair of Machine Learning and Computer Vision. This datasets serves as an enriched version of the original TrackNet Tennis dataset. It cuts the whole rallies included in the TrackNet dataset in single trajectories, based on TrackNets property… See the full description on the dataset page: https://huggingface.co/datasets/XSpaceCoderX/ACE-Trajectories_noTosses.texttabular-classificationn<1K0 likes794 downloads5mo agoHugging Face28APProjects /us-warn-act-layoff-notice-period-days How much notice did US layoff notices actually give? The WARN Act is, at bottom, a law about a number of days. Every state publishes layoff notices; none of them publishes the one column that says whether the notice arrived in time. This dataset is that column, recomputed every day from the primary filings. notice_days = effective_date - notice_date, per notice. The headline, over the full archive (61,330 notices, 1988 - today) Notices stating both a… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-warn-act-layoff-notice-period-days.tabulartabular-regression10K<n<100K0 likes766 downloads12h agoHugging Face29notmax123 /ivirits-audio-v2-30s ivrit.ai audio-v2 — 2–30 s segments ivrit-ai/audio-v2 (>20k hours of Hebrew audio) cut into 2–30 second speech segments with machine transcripts, ready for ASR fine-tuning. How it was built VAD — Silero VAD (ONNX) over each episode decoded to 16 kHz mono. Speech regions longer than 30 s are split at the quietest sufficiently-long pause inside the window, so cuts land in silence rather than mid-word. Regions shorter than 2 s are dropped. Transcription —… See the full description on the dataset page: https://huggingface.co/datasets/notmax123/ivirits-audio-v2-30s.audioautomatic-speech-recognition1M<n<10M0 likes755 downloads1mo agoHugging Face30APProjects /us-layoffs-by-stock-ticker-warn-act-notices-public-companies US layoffs by stock ticker: 7,651 WARN Act notices filed by 736 listed companies, with confidence tiers Rebuilt 2026-09-22. State WARN Act filings name the employer as the filer wrote it — Wells Fargo Home Mortgage, OS Restaurant Services, Boeing Compnay — never a ticker. This dataset resolves those strings to the listed parent and re-derives the mapping every day as new filer strings appear. 1,583 filer strings → 736 tickers → 7,651 notices (1988-12-16 → 2026-09-16), 1,007,332… See the full description on the dataset page: https://huggingface.co/datasets/APProjects/us-layoffs-by-stock-ticker-warn-act-notices-public-companies.tabulartabular-classification1K<n<10K0 likes725 downloads2h agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.