CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ieasybooks-org /prophet-mosque-library Prophet's Mosque Library 📖 Overview Prophet’s Mosque Library is one of the primary resources for Islamic books. It hosts more than 48,000 PDF books across over 70 categories. In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX. 📊 Dataset Contents The dataset includes 70,884 PDF files (spanning 23,494,042 pages) representing 48,717 Islamic books. Each book is… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/prophet-mosque-library.textimage-to-text10K<n<100K6 likes297k downloads1y agoHugging Face02ieasybooks-org /waqfeya-library Waqfeya Library 📖 Overview Waqfeya is one of the primary online resources for Islamic books, similar to Shamela. It hosts more than 10,000 PDF books across over 80 categories. In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX. 📊 Dataset Contents The dataset includes 22,443 PDF files (spanning 8,978,634 pages) representing 10,150 Islamic books. Each book is… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/waqfeya-library.imageimage-to-text10K<n<100K12 likes136k downloads1y agoHugging Face03ieasybooks-org /shamela-waqfeya-library Shamela Waqfeya Library 📖 Overview Shamela Waqfeya is one of the primary online resources for Islamic books, similar to Shamela. It hosts more than 4,500 PDF books across over 40 categories. In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX. 📊 Dataset Contents The dataset includes 12,877 PDF files (spanning 5,138,027 pages) representing 4,661 Islamic books.… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/shamela-waqfeya-library.tabularimage-to-text1K<n<10K4 likes91k downloads1y agoHugging Face04Owen777 /HQ-OpenHumanVidtabular1M<n<10M6 likes22k downloads10mo agoHugging Face05openai /MMMLU Multilingual Massive Multitask Language Understanding (MMMLU) The MMLU is a widely recognized benchmark of general knowledge attained by AI models. It covers a broad range of topics from 57 different categories, covering elementary-level knowledge up to advanced professional subjects like law, physics, history, and computer science. We translated the MMLU’s test set into 14 languages using professional human translators. Relying on human translators for this evaluation increases… See the full description on the dataset page: https://huggingface.co/datasets/openai/MMMLU.textquestion-answering100K<n<1M526 likes12k downloads2y agoHugging Face06bench-llm /or-bench OR-Bench: An Over-Refusal Benchmark for Large Language Models Please see our demo at HuggingFace Spaces. Overall Plots of Model Performances Below is the overall model performance. X axis shows the rejection rate on OR-Bench-Hard-1K and Y axis shows the rejection rate on OR-Bench-Toxic. The best aligned model should be on the top left corner of the plot where the model rejects the most number of toxic prompts and least number of safe prompts. We also plot a blue line… See the full description on the dataset page: https://huggingface.co/datasets/bench-llm/or-bench.imagetext-generation10K<n<100K22 likes9.4k downloads2y agoHugging Face07databricks /officeqagated OfficeQA Dataset Summary OfficeQA is a grounded reasoning benchmark by Databricks for evaluating model and agent performance on end-to-end reasoning over real-world documents. The benchmark consists of question–answer pairs that require reasoning over historical U.S. Treasury Bulletin documents (1939–2025), which contain dense financial tables, charts, and narrative text. OfficeQA is designed to test retrieval, tool use, and multi-step reasoning in… See the full description on the dataset page: https://huggingface.co/datasets/databricks/officeqa.documentquestion-answeringn<1K26 likes9.1k downloads2mo agoHugging Face08opensporks /resumes Dataset Card for Resume Dataset Dataset Summary Context A collection of Resume Examples taken from livecareer.com for categorizing a given resume into any of the labels defined in the dataset. Content Contains 2400+ Resumes in string as well as PDF format. PDF stored in the data folder differentiated into their respective labels as folders with each resume residing inside the folder in pdf form with filename as the id defined in the csv. Inside the… See the full description on the dataset page: https://huggingface.co/datasets/opensporks/resumes.text1K<n<10K14 likes9k downloads2y agoHugging Face09Abtinzandi /Obstacle-Detection-Dataset-YOLO ROD-Dataset: Real-Time Obstacle Detection for Smartphone-Based Assistive Vision 24,326-image, 25-class YOLO dataset for obstacle detection This dataset is the data product of our Real-Time Obstacle Detection (ROD) project at Amirkabir University of Technology, Tehran. The project addresses two related public-safety problems on the city sidewalk: the limited situational awareness of people living with visual impairments, and the elevated collision and fall risk for pedestrians… See the full description on the dataset page: https://huggingface.co/datasets/Abtinzandi/Obstacle-Detection-Dataset-YOLO.imageobject-detection10K<n<100K16 likes8.3k downloads3mo agoHugging Face10ArtificialAnalysis /AA-Omniscience-Public Public Dataset for AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Large Language Models AA-Omniscience-Public contains 600 questions across a wide range of domains used to test a model’s knowledge and hallucination tendencies. Leaderboard and detailed results Paper Introduction We introduce AA-Omniscience, a benchmark dataset designed to measure a model’s ability to both recall factual information accurately across domains, and correctly… See the full description on the dataset page: https://huggingface.co/datasets/ArtificialAnalysis/AA-Omniscience-Public.documentquestion-answeringn<1K50 likes7.9k downloads1mo agoHugging Face11hf-audio /open-asr-leaderboard-resultstabularn<1K0 likes5.2k downloads2d agoHugging Face12victor /real-or-fake-fake-jobposting-predictiontabular10K<n<100K5 likes4.8k downloads4y agoHugging Face13oskarvanderwal /winogenderSource: https://github.com/rudinger/winogender-schemas/tree/master @InProceedings{rudinger-EtAl:2018:N18, author = {Rudinger, Rachel and Naradowsky, Jason and Leonard, Brian and {Van Durme}, Benjamin}, title = {Gender Bias in Coreference Resolution}, booktitle = {Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies}, month = {June}, year = {2018}, address = {New Orleans… See the full description on the dataset page: https://huggingface.co/datasets/oskarvanderwal/winogender.textn<1K3 likes4.8k downloads3y agoHugging Face14prism-oncology /novae Description Full novae dataset, including: All the spatial transcriptomics samples used to train Novae Protein samples used in the article Some Visium and Visium HD samples Synthetic data samples You can download this dataset from the API, see novae.load_dataset See here the list of available models trained on this dataset. [!NOTE] Note that Novae was trained on the image-based spatial transcriptomics samples. This means that it was not trained on the Visium/VisiumHD samples… See the full description on the dataset page: https://huggingface.co/datasets/prism-oncology/novae.tabularn<1K5 likes3.8k downloads4mo agoHugging Face15memo-ozdincer /jepa-qwen3-32b-pure-baselines-2026-05-25 JEPA-Align: Qwen3-32B Safety Defense Matrix The complete 11-condition Qwen3-32B experiment for Predictive Representation Alignment (PRA), the paired-view objective introduced in Predictive Representation Alignment Improves Generalization in LLM Safety. PRA aligns adversarially rewritten prompts with clean prompts expressing the same intent. This release contains trained adapters, attack traces, benign capability evaluations, machine-readable results, and paper-ready tables for… See the full description on the dataset page: https://huggingface.co/datasets/memo-ozdincer/jepa-qwen3-32b-pure-baselines-2026-05-25.tabulartext-classificationn<1K0 likes3.3k downloads1mo agoHugging Face16osunlp /TravelPlanner TravelPlanner Dataset TravelPlanner is a benchmark crafted for evaluating language agents in tool-use and complex planning within multiple constraints. (See our paper for more details.) Introduction In TravelPlanner, for a given query, language agents are expected to formulate a comprehensive plan that includes transportation, daily meals, attractions, and accommodation for each day. TravelPlanner comprises 1,225 queries in total. The number of days and hard constraints… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/TravelPlanner.tabulartext-generation1K<n<10K86 likes3.2k downloads2y agoHugging Face17OpenMOSS-Team /SWE-bench-Science SWE-bench Science SWE-bench Science evaluates coding agents on software-engineering tasks drawn from scientific-computing repositories. The release contains 119 tasks across 20 scientific domains, with isolated environments and separate programmatic verifiers. GitHub release repository: OpenMOSS/SWE-bench-Science Runtime images: Docker Hub, pinned by immutable linux/amd64 digests Evaluation framework: Pier, compatible with Harbor task format Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/SWE-bench-Science.textn<1K7 likes3k downloads1mo agoHugging Face18databricks /officeqa-pro-v2gated OfficeQA Pro v2 Dataset Summary OfficeQA Pro v2 is a grounded reasoning benchmark by Databricks for evaluating model and agent performance on end-to-end reasoning over real-world documents. The benchmark consists of question–answer pairs that require reasoning over two centuries of U.S. Federal Accounts of Receipts and Expenditures reporting (1793–2024) — Combined Statements of Receipts, Outlays, and Balances of the United States Government, together with earlier… See the full description on the dataset page: https://huggingface.co/datasets/databricks/officeqa-pro-v2.documentquestion-answeringn<1K17 likes2.6k downloads2mo agoHugging Face19openadmet /cyp-challenge-train-test CYP Challenge Train/Test Dataset A high-quality experimental dataset for predicting inhibition of the major drug-metabolizing Cytochrome P450 enzymes (CYP1A2, CYP2C9, CYP2D6, CYP3A4), released as part of the OpenADMET CYP Inhibition Blind Challenge. Blog post: Announcing OpenADMET’s CYP inhibition blind challenge Challenge Space: OpenADMET CYP Inhibition Blind Challenge Challenge period: August 17, 2026 - November 3, 2026 Produced by: OpenADMET CHANGELOG Updated… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/cyp-challenge-train-test.tabulartabular-regression10K<n<100K9 likes2.6k downloads1d agoHugging Face20wulipc /CC-OCR CC-OCR This is the Repository for CC-OCR Benchmark. Dataset and evaluation code for the Paper "CC-OCR: A Comprehensive and Challenging OCR Benchmark for Evaluating Large Multimodal Models in Literacy". 🚀 GitHub   |   🤗 Hugging Face   |   🤖 ModelScope   |    📑 Paper    |   📗 Blog Here is hosting the tsv version of CC-OCR data, which is used for evaluation in VLMEvalKit. Please refer to our GitHub for more information. Benchmark Leaderboard Model… See the full description on the dataset page: https://huggingface.co/datasets/wulipc/CC-OCR.text1K<n<10K5 likes2.4k downloads2y agoHugging Face21matichon /thai-onet-m6-exam Thai O-Net Exams Dataset Overview The Thai O-Net Exams dataset is a comprehensive collection of exam questions and answers from the Thai Ordinary National Educational Test (O-Net). This dataset covers various subjects for Grade 12 (M6) level, designed to assist in educational research and development of question-answering systems. Dataset Source Thai National Institute of Educational Testing Service (NIETS) Maintainer Dr. Kobkrit Viriyayudhakorn… See the full description on the dataset page: https://huggingface.co/datasets/matichon/thai-onet-m6-exam.textquestion-answering1K<n<10K0 likes2.2k downloads5mo agoHugging Face22OzzyChen97 /TC-SSA TC-SSA: Token Compression via Semantic Slot Aggregation for Gigapixel Pathology Reasoning Links: Project homepage | arXiv paper | Code Authors: Zhuo Chen1,2, Xiaoyu Yang1, and Lijian Xu1,* 1 Shenzhen University of Advanced Technology, Shenzhen, Guangdong, China2 University of Nottingham Ningbo China, FoSE, Ningbo, Zhejiang, China* Corresponding author: xulijian@suat-sz.edu.cn TC-SSA WSI Feature Bags This public repository contains pre-extracted whole-slide… See the full description on the dataset page: https://huggingface.co/datasets/OzzyChen97/TC-SSA.tabularimage-feature-extraction1K<n<10K1 likes2.2k downloads2mo agoHugging Face23Linzhan /Objaverse-XL-Rigged-Animated Objaverse-XL Rigged & Animated Subset Every asset here carries both a skeleton and at least one animation clip, selected from Objaverse / Objaverse-XL. Rigs range from 3 to 344 joints and span characters as well as articulated rigid objects. Objaverse-XL indexes over 10 million objects, but only a small fraction carry a usable rig and motion on it. This subset isolates that fraction: every file was checked to contain at least one skin with joints and at least one animation clip… See the full description on the dataset page: https://huggingface.co/datasets/Linzhan/Objaverse-XL-Rigged-Animated.3dtext-to-3d10K<n<100K1 likes2.1k downloads17d agoHugging Face24ieasybooks-org /prophet-mosque-library-compressed Prophet's Mosque Library - Compressed 📖 Overview Prophet’s Mosque Library is one of the primary resources for Islamic books. It hosts more than 48,000 PDF books across over 70 categories. In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX. 📊 Dataset Contents This dataset is identical to ieasybooks-org/prophet-mosque-library, with one key… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/prophet-mosque-library-compressed.textimage-to-text10K<n<100K0 likes2.1k downloads1y agoHugging Face25Real-TSF /TIME-OutputThis repository contains the extracted time series features (tsfeatures) for each variate and the detailed forecasting results for every experiment. Note: These files are for building leaderboard and visualization; users do not need to download this directory. features/: Statistical Features (tsfeatures) Each dataset's features are saved to: output/features/{dataset}/{freq}/. This directory stores the computed tsfeatures for the variates in the dataset. The folder contains a CSV file… See the full description on the dataset page: https://huggingface.co/datasets/Real-TSF/TIME-Output.tabulartime-series-forecasting1K<n<10K0 likes2.1k downloads2d agoHugging Face26KRAFTON /Raon-OpenTTS-Eval Raon-OpenTTS-Eval Technical Report A robustness-oriented evaluation benchmark for zero-shot text-to-speech, covering 4 acoustic regimes (Clean, Noisy, Wild, Expressive) across 12 datasets with 6,000 prompt–text pairs. Existing zero-shot TTS benchmarks typically evaluate models using prompts drawn from a single read-speech dataset, providing an incomplete view of robustness under realistic and challenging recording scenarios. Raon-OpenTTS-Eval… See the full description on the dataset page: https://huggingface.co/datasets/KRAFTON/Raon-OpenTTS-Eval.audiotext-to-speech1K<n<10K9 likes1.9k downloads4mo agoHugging Face27stablellama /Krea-2-Raw_samples_Best_ofThis dataset is a highly diverse set of high quality images generated with Krea 2 Raw. NOTE: Raw is not intended for image generation, so do not use these images to judge the quality of the model. Raw is intended for training, as are the samples in this dataset as they can be used for regularization. Possible uses Regularization images for training models based on Krea 2 Raw Quality testing Data source This dataset is derived from… See the full description on the dataset page: https://huggingface.co/datasets/stablellama/Krea-2-Raw_samples_Best_of.tabulartext-to-image1K<n<10K0 likes1.9k downloads24d agoHugging Face28ieasybooks-org /waqfeya-library-compressed Waqfeya Library - Compressed 📖 Overview Waqfeya is one of the primary online resources for Islamic books, similar to Shamela. It hosts more than 10,000 PDF books across over 80 categories. In this dataset, we processed the original PDF files using Google Document AI APIs and extracted their contents into two additional formats: TXT and DOCX. 📊 Dataset Contents This dataset is identical to ieasybooks-org/waqfeya-library, with one key difference: the contents… See the full description on the dataset page: https://huggingface.co/datasets/ieasybooks-org/waqfeya-library-compressed.tabularimage-to-text10K<n<100K6 likes1.8k downloads1y agoHugging Face29open-spaced-repetition /fsrs-datasettabular10M<n<100M5 likes1.8k downloads3y agoHugging Face30openadmet /openadmet-expansionrx-challenge-data OpenADMET-ExpansionRx Challenge FULL dataset This is the full dataset used in the OpenADMET-ExpansionRx blind challenge, which finalized in January 19th, 2026. Originally split in a train and blinded test set, we now release the full dataset, which contains real-work ADMET data from a recently prosecuted series of drug discovery campaigns by Expansion Therapeutics on RNA mediated diseases. While optimising candidate molecules for their preclinical programs Expansion collected… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/openadmet-expansionrx-challenge-data.tabular10K<n<100K12 likes1.7k downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.