CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Owen777 /HQ-OpenHumanVidtabular1M<n<10M6 likes12k downloads10mo agoHugging Face02openai /MMMLU Multilingual Massive Multitask Language Understanding (MMMLU) The MMLU is a widely recognized benchmark of general knowledge attained by AI models. It covers a broad range of topics from 57 different categories, covering elementary-level knowledge up to advanced professional subjects like law, physics, history, and computer science. We translated the MMLU’s test set into 14 languages using professional human translators. Relying on human translators for this evaluation increases… See the full description on the dataset page: https://huggingface.co/datasets/openai/MMMLU.textquestion-answering100K<n<1M526 likes12k downloads2y agoHugging Face03opensporks /resumes Dataset Card for Resume Dataset Dataset Summary Context A collection of Resume Examples taken from livecareer.com for categorizing a given resume into any of the labels defined in the dataset. Content Contains 2400+ Resumes in string as well as PDF format. PDF stored in the data folder differentiated into their respective labels as folders with each resume residing inside the folder in pdf form with filename as the id defined in the csv. Inside the… See the full description on the dataset page: https://huggingface.co/datasets/opensporks/resumes.text1K<n<10K14 likes8.9k downloads2y agoHugging Face04hf-audio /open-asr-leaderboard-resultstabularn<1K0 likes4.5k downloads2d agoHugging Face05OpenMOSS-Team /SWE-bench-Science SWE-bench Science SWE-bench Science evaluates coding agents on software-engineering tasks drawn from scientific-computing repositories. The release contains 119 tasks across 20 scientific domains, with isolated environments and separate programmatic verifiers. GitHub release repository: OpenMOSS/SWE-bench-Science Runtime images: Docker Hub, pinned by immutable linux/amd64 digests Evaluation framework: Pier, compatible with Harbor task format Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/SWE-bench-Science.textn<1K7 likes3.4k downloads28d agoHugging Face06openadmet /cyp-challenge-train-test CYP Challenge Train/Test Dataset A high-quality experimental dataset for predicting inhibition of the major drug-metabolizing Cytochrome P450 enzymes (CYP1A2, CYP2C9, CYP2D6, CYP3A4), released as part of the OpenADMET CYP Inhibition Blind Challenge. Blog post: Announcing OpenADMET’s CYP inhibition blind challenge Challenge Space: OpenADMET CYP Inhibition Blind Challenge Challenge period: August 17, 2026 - November 3, 2026 Produced by: OpenADMET CHANGELOG Updated… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/cyp-challenge-train-test.tabulartabular-regression10K<n<100K9 likes2.8k downloads25d agoHugging Face07KRAFTON /Raon-OpenTTS-Eval Raon-OpenTTS-Eval Technical Report A robustness-oriented evaluation benchmark for zero-shot text-to-speech, covering 4 acoustic regimes (Clean, Noisy, Wild, Expressive) across 12 datasets with 6,000 prompt–text pairs. Existing zero-shot TTS benchmarks typically evaluate models using prompts drawn from a single read-speech dataset, providing an incomplete view of robustness under realistic and challenging recording scenarios. Raon-OpenTTS-Eval… See the full description on the dataset page: https://huggingface.co/datasets/KRAFTON/Raon-OpenTTS-Eval.audiotext-to-speech1K<n<10K9 likes1.9k downloads4mo agoHugging Face08open-spaced-repetition /fsrs-datasettabular10M<n<100M5 likes1.7k downloads3y agoHugging Face09openadmet /openadmet-expansionrx-challenge-data OpenADMET-ExpansionRx Challenge FULL dataset This is the full dataset used in the OpenADMET-ExpansionRx blind challenge, which finalized in January 19th, 2026. Originally split in a train and blinded test set, we now release the full dataset, which contains real-work ADMET data from a recently prosecuted series of drug discovery campaigns by Expansion Therapeutics on RNA mediated diseases. While optimising candidate molecules for their preclinical programs Expansion collected… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/openadmet-expansionrx-challenge-data.tabular10K<n<100K12 likes1.6k downloads8mo agoHugging Face10openadmet /pxr-challenge-train-test PXR Challenge Train/Test Dataset A high-quality experimental dataset for predicting human Pregnane-X Receptor (PXR) induction, comprising over 11,000 compounds screened using a high-fidelity in-house assay. This is the largest publicly available PXR activity dataset, released as part of the OpenADMET PXR Induction Blind Challenge. Blog post: Announcing the Next OpenADMET Blind Challenge: Predicting PXR Induction Challenge Space: openadmet/pxr-challenge Challenge period: April 1… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/pxr-challenge-train-test.tabulartabular-regression10K<n<100K17 likes1.6k downloads20d agoHugging Face11openlifescienceai /Med-HALT Med-HALT: Medical Domain Hallucination Test for Large Language Models This is a dataset used in the Med-HALT research paper. This research paper focuses on the challenges posed by hallucinations in large language models (LLMs), particularly in the context of the medical domain. We propose a new benchmark and dataset, Med-HALT (Medical Domain Hallucination Test), designed specifically to evaluate hallucinations. Med-HALT provides a diverse multinational dataset derived from medical… See the full description on the dataset page: https://huggingface.co/datasets/openlifescienceai/Med-HALT.tabular10K<n<100K11 likes1.5k downloads3y agoHugging Face12openthaigpt /thai-onet-m6-exam Thai O-Net Exams Dataset Overview The Thai O-Net Exams dataset is a comprehensive collection of exam questions and answers from the Thai Ordinary National Educational Test (O-Net). This dataset covers various subjects for Grade 12 (M6) level, designed to assist in educational research and development of question-answering systems. Dataset Source Thai National Institute of Educational Testing Service (NIETS) Maintainer Dr. Kobkrit… See the full description on the dataset page: https://huggingface.co/datasets/openthaigpt/thai-onet-m6-exam.textquestion-answering1K<n<10K7 likes1.4k downloads3y agoHugging Face13leodriesch /open-hdri-1k Open HDRI 1K A consolidated, public-domain (CC0-1.0) collection of 3,491 equirectangular HDR environment maps at 1K resolution, gathered from five free HDRI libraries: Poly Haven, BlenderKit, ambientCG, CGEES and Open HDRI. Every map is stored as a linear, high-dynamic-range .exr file alongside a tonemapped .jpg preview, with a per-asset metadata row (dimensions, source, author, license, SHA-256 checksum and tags). Contents Source Assets Author(s) License… See the full description on the dataset page: https://huggingface.co/datasets/leodriesch/open-hdri-1k.imageimage-to-image1K<n<10K1 likes1.4k downloads2mo agoHugging Face14openai /genebench-pro-public-package GeneBench-Pro Public Case Studies This repository contains public GeneBench-Pro case studies. It is the self-contained package intended for public distribution, including Hugging Face publication. Package Layout <repo-root>/ ├── .gitattributes ├── README.md ├── LICENSE ├── problems.csv ├── checksums.sha256 ├── manifest.json ├── reference_definitions.md ├── reference_grader.py └── problems/ └── <eval_id>/ ├── eval_config.json ├── data_files/… See the full description on the dataset page: https://huggingface.co/datasets/openai/genebench-pro-public-package.documentn<1K14 likes1.3k downloads3mo agoHugging Face15openadmet /openadmet-expansionrx-challenge-train-data OpenADMET-ExpansionRx Challenge training dataset This dataset contains real-work ADMET data from a recently prosecuted series of drug discovery campaigns by Expansion Therapeutics on RNA mediated diseases. While optimising candidate molecules for their preclinical programs Expansion collected a variety of ADMET data for off-targets and properties of interest in the traditional game of “whack-a-mole” familiar to all drug hunters. Now, they’ve made the bold and generous decision to… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/openadmet-expansionrx-challenge-train-data.tabular10K<n<100K10 likes876 downloads10mo agoHugging Face16tomdoyo /open-command OpenCommand Catcher targets and pitcher command for MLB, inferred from broadcast video. OpenCommand scores command using the pitch location's distance from target. This dataset contains the 2024/2025/2026 computer vision object detections, every intermediate the pipeline writes, and the resulting command scores. The pipeline itself and the full method write-up live at github.com/tomdoyo/open-command. Download hf download tomdoyo/open-command --repo-type dataset… See the full description on the dataset page: https://huggingface.co/datasets/tomdoyo/open-command.tabular1M<n<10M1 likes761 downloads25d agoHugging Face17opensporks /crunchbasetabular1M<n<10M14 likes749 downloads2y agoHugging Face18blairducrayoppat /openvino-arc140v-lunarlake OpenVINO local-inference on an Intel Arc 140V (Lunar Lake) iGPU Reference performance data for running local models on a single Intel Core Ultra 7 258V (Lunar Lake) laptop with the integrated Intel Arc 140V (Xe2) GPU, via OpenVINO. All inference runs on the iGPU; the NPU stays idle throughout, confirmed by the telemetry here. This is reference characterization shared by a non-expert contributor — careful measurements on one machine, offered so others can compare and correct, not… See the full description on the dataset page: https://huggingface.co/datasets/blairducrayoppat/openvino-arc140v-lunarlake.tabulartext-generationn<1K0 likes696 downloads6d agoHugging Face19openadmet /openadmet-expansionrx-challenge-test-data-blinded OpenADMET-ExpansionRx Challenge blinded test dataset This dataset contains real-work ADMET data from a recently prosecuted series of drug discovery campaigns by Expansion Therapeutics on RNA mediated diseases. While optimising candidate molecules for their preclinical programs Expansion collected a variety of ADMET data for off-targets and properties of interest in the traditional game of “whack-a-mole” familiar to all drug hunters. Now, they’ve made the bold and generous decision… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/openadmet-expansionrx-challenge-test-data-blinded.text1K<n<10K5 likes584 downloads11mo agoHugging Face20OpenGVLab /MMT-Bench Dataset Card for MMT-Bench Repository: https://github.com/OpenGVLab/MMT-Bench Paper: https://openreview.net/forum?id=R4Ng8zYaiz Point of Contact: Wenqi Shao Introduction Large Vision-Language Models (LVLMs) show significant strides in general-propose multimodal applications such as visual dialogue and embodied navigation. However, existing multimodal evaluation benchmarks cover a limited number of multimodal tasks testing rudimentary capabilities, falling short in… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/MMT-Bench.text10K<n<100K7 likes463 downloads2y agoHugging Face21openbrain-anon /openbrain_v1_0 OpenBrain v1.0 OpenBrain v1.0 is a public-ready release of brain-extracted T1-weighted MRI images, SynthStrip-derived brain masks, and automated whole-brain segmentation labels. Release Contents Cases: 35,838 Source OpenNeuro datasets: 607 License: CC0 Artifacts per case: image.nii.gz: revised brain-extracted T1w image brain_mask.nii.gz: SynthStrip-derived brain mask whole_brain_segmentation.nii.gz: automated whole-brain segmentation label All released cases are… See the full description on the dataset page: https://huggingface.co/datasets/openbrain-anon/openbrain_v1_0.tabularn<1K0 likes462 downloads5mo agoHugging Face22dongludeeplearning /OpenASL_3D This is a Large-Scale 3D Datset for Continuous American Sign Language tabular10K<n<100K0 likes441 downloads1y agoHugging Face23lodestone-horizon /OpenVid-1M Summary This is the dataset proposed in our paper "OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation". OpenVid-1M is a high-quality text-to-video dataset designed for research institutions to enhance video quality, featuring high aesthetics, clarity, and resolution. It can be used for direct training or as a quality tuning complement to other video datasets. All videos in the OpenVid-1M dataset have resolutions of at least 512×512. Furthermore, we… See the full description on the dataset page: https://huggingface.co/datasets/lodestone-horizon/OpenVid-1M.texttext-to-video1M<n<10M0 likes430 downloads2y agoHugging Face24zd1949 /openfwi-preprocessed-72x72textn<1K0 likes380 downloads1y agoHugging Face25open-spaced-repetition /FSRS-Anki-20kgated Update We have released a new dataset: anki-revlogs-10k. Introduction FSRS-Anki-20k is a dataset of 20k collections from Anki for FSRS project. It is a random sample of collections with 5000+ revlog entries, so it should contain a mix of older (still active) users, and newer users. Entries are pre-sorted in (cid, id) order. There are two versions of the dataset: ./revlogs and ./dataset. The ./revlogs version contains the raw revlog entries, while the ./dataset version… See the full description on the dataset page: https://huggingface.co/datasets/open-spaced-repetition/FSRS-Anki-20k.tabular1B<n<10B22 likes361 downloads2y agoHugging Face26scikit-fingerprints /ExpansionRx_OpenADMET_RLM_CLint ExpansionRx-OpenADMET RLM CLint RLM CLint (rat liver microsomal intrinsic clearance) dataset from the ExpansionRx-OpenADMET Blind Challenge [1] [2]. It is intended to be used through scikit-fingerprints library. The task is to predict the rat liver microsomal intrinsic clearance (RLM CLint) of molecules. Note that this dataset was not part of the original challenge. It was provided by the organizers afterward as an additional endpoint. Characteristic Description Tasks 1… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/ExpansionRx_OpenADMET_RLM_CLint.texttabular-regressionn<1K0 likes358 downloads6mo agoHugging Face27namanvats /harbor-goose-openhands-benchmark Same Model, Opposite Results: Goose vs OpenHands Turn Budget Study on Harbor Terminal-Bench-Pro Trial-level results from a small controlled study comparing two agent harnesses — Goose and OpenHands-SDK — on a frozen 40-task Harbor Terminal-Bench-Pro slice. All runs used minimax/minimax-m2.5 via OpenRouter with Daytona as the sandbox backend. Key Findings Reducing the turn budget from 100 to 60 pushed the two harnesses in opposite directions under the base setup:… See the full description on the dataset page: https://huggingface.co/datasets/namanvats/harbor-goose-openhands-benchmark.tabularn<1K3 likes349 downloads5mo agoHugging Face28openadmet /Octant_CYP_inhibition_reactivity_blog_release OpenADMET Octant CYP Inhibition & Reactivity Data release from the OpenADMET consortium, generated by Octant Bio. This dataset accompanies the blog post Building the OpenADMET Data Engine. Source code, assay protocols, and raw TSV files are on GitHub. Overview Cytochrome P450 (CYP) enzymes drive the oxidative metabolism of most drugs and are a primary cause of drug-drug interactions (DDIs). Despite their importance, public CYP datasets are sparse, noisy, and… See the full description on the dataset page: https://huggingface.co/datasets/openadmet/Octant_CYP_inhibition_reactivity_blog_release.tabular10K<n<100K3 likes309 downloads2mo agoHugging Face29scikit-fingerprints /ExpansionRx_OpenADMET_KSOL ExpansionRx-OpenADMET KSOL KSOL dataset from the ExpansionRx-OpenADMET Blind Challenge [1] [2]. It is intended to be used through scikit-fingerprints library. The task is to predict KSOL of molecules. Characteristic Description Tasks 1 Task type regression Total samples 7298 Recommended split time Recommended metric MAE References [1] OpenADMET team "Announcement 1: ExpansionRx-OpenADMET Blind Challenge"… See the full description on the dataset page: https://huggingface.co/datasets/scikit-fingerprints/ExpansionRx_OpenADMET_KSOL.texttabular-regression1K<n<10K0 likes298 downloads6mo agoHugging Face30LaconicAI /text_message_function_calling_open_chatThis is a small synthetic dataset to model a function call for text messaging someone from a cell phone. This has been tested with and used to finetune a set of smaller models and deployed directly on the pixel 8 pro and Fold 4 phones. texttext-generation10K<n<100K2 likes260 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.