CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01gOLIVES /OLIVES_Dataset OLIVES_Dataset Abstract Clinical diagnosis of the eye is performed over multifarious data modalities including scalar clinical labels, vectorized biomarkers, two-dimensional fundus images, and three-dimensional Optical Coherence Tomography (OCT) scans. While the clinical labels, fundus images and OCT scans are instrumental measurements, the vectorized biomarkers are interpreted attributes from the other measurements. Clinical practitioners use all these data modalities… See the full description on the dataset page: https://huggingface.co/datasets/gOLIVES/OLIVES_Dataset.image100K<n<1M6 likes2.1k downloads6mo agoHugging Face02Oliver-Ma /Real-3DQA Real-3DQA Do 3D Large Language Models Really Understand 3D Spatial Relationships? 🌐 Project Page · 📄 Paper · 💻 GitHub Overview Real-3DQA is a debiased 3D spatial QA benchmark with viewpoint rotation consistency evaluation. It addresses two key shortcomings of existing benchmarks: Language Shortcut Filtering — Questions answerable through linguistic priors alone are removed by comparing 3D-LLMs against blind text-only counterparts. Viewpoint Rotation Score (VRS) — Each… See the full description on the dataset page: https://huggingface.co/datasets/Oliver-Ma/Real-3DQA.textimage-text-to-text1K<n<10K8 likes1.9k downloads6mo agoHugging Face03SOHAIBSUL123 /OLIVES_Dataset OLIVES_Dataset Abstract Clinical diagnosis of the eye is performed over multifarious data modalities including scalar clinical labels, vectorized biomarkers, two-dimensional fundus images, and three-dimensional Optical Coherence Tomography (OCT) scans. While the clinical labels, fundus images and OCT scans are instrumental measurements, the vectorized biomarkers are interpreted attributes from the other measurements. Clinical practitioners use all these data modalities… See the full description on the dataset page: https://huggingface.co/datasets/SOHAIBSUL123/OLIVES_Dataset.image100K<n<1M0 likes655 downloads6mo agoHugging Face04oliveryanzuolu /crafter-cache-81f-256pxtabular10K<n<100K0 likes601 downloads2mo agoHugging Face05olive5 /ml-lectures ML/Math Lecture Archive Archived lecture videos (1080p MP4) with English subtitles (.vtt) from publicly available university course recordings on YouTube. 268 videos, ~72 GB. Contents Folder Course Videos 18.065_Strang/ MIT 18.065 — Matrix Methods in Data Analysis, Signal Processing, and Machine Learning (Gilbert Strang) 36 18.06SC_LinearAlgebra/ MIT 18.06SC — Linear Algebra, Fall 2011 (Gilbert Strang) 74 CS109_Piech/ Stanford CS109 — Introduction… See the full description on the dataset page: https://huggingface.co/datasets/olive5/ml-lectures.videovideo-text-to-textn<1K0 likes581 downloads14d agoHugging Face06Oliver-Ma /IDPP-Data IDPP-Data Video archives for IDPP / SVD physics-benchmark reproduction: friction, viscosity, and elasticity, each with absolute and relative splits (train / test1 / test2 / test3 where applicable). Repository: Oliver-Ma/IDPP-DataSnapshot (README revision): 2026-04-12 — documents 26 .zip blobs under video_zips/ (mirroring local dataset/<lane>/). Important: archives, not frame folders Everything under video_zips/ is a .zip file. Training and evaluation code (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/Oliver-Ma/IDPP-Data.video0 likes322 downloads5mo agoHugging Face07oliverkinch /da-bird DA-BIRD: Danish NL2SQL Benchmark DA-BIRD is a Danish text-to-SQL benchmark for evaluating large language models on natural language to SQL generation. All tasks are in Danish and use SQLite databases. The dataset is designed for use with the Harbor evaluation framework. It combines two corpora — 363 tasks across 22 unique databases: Corpus Tasks Databases Language Difficulty bird_* 150 11 Danish (translated) easy / medium / hard dst_* 213 11 Danish (original) medium /… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/da-bird.table-question-answering0 likes291 downloads4mo agoHugging Face08Oliver1515 /ProGAN-Eval ProGAN eval set This repository hosts the test set used in the FPBA on the LSUN Synthesis dataset, consisting of 1,000 test images from the LSUN dataset and 1,000 images synthesized by ProGAN. imagen<1K0 likes253 downloads9mo agoHugging Face09OliverHausdoerfer /shelf_jimUSE FIRST COMMIT FOR ORIGINAL DATASET. LATEST COMMIT IS MAHATHI-SHELF DATASET image1K<n<10K0 likes245 downloads2mo agoHugging Face10oliverdk /nl_gameable_programmatic_graderstextn<1K0 likes207 downloads3mo agoHugging Face11OliverHausdoerfer /libero_90_sawyer_defaultCams_failuresThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "sawyer", "total_episodes": 4278, "total_frames": 637138, "total_tasks": 74, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 200, "fps": 20, "splits": { "train": "0:4278" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/OliverHausdoerfer/libero_90_sawyer_defaultCams_failures.tabularrobotics100K<n<1M0 likes186 downloads3mo agoHugging Face12OliverHausdoerfer /mahathi_bottleimagen<1K0 likes179 downloads2mo agoHugging Face13OliverZhao /dreamvla_il_datavideon<1K0 likes175 downloads1y agoHugging Face14oliverdk /impossible_mbpp_natural_diversetextn<1K0 likes175 downloads4mo agoHugging Face15OliverHuang1998 /3DRS1 likes156 downloads1y agoHugging Face16OliverUrbann /HumanoidRobotSoccer Fall Prediction Dataset for Humanoid Robots Dataset Summary This dataset consists of 37.9 hours of real-world sensor data collected from 20 Nao humanoid robots over the course of one year in various test environments, including RoboCup soccer matches. The dataset includes 18.3 hours of walking data, featuring 2519 falls. It captures a wide range of activities such as omni-directional walking, collisions, standing up, and falls on various surfaces like artificial turf and… See the full description on the dataset page: https://huggingface.co/datasets/OliverUrbann/HumanoidRobotSoccer.tabular10M<n<100M2 likes155 downloads1y agoHugging Face17oliverwang15 /news_with_gpt_instructions Dataset Card for "news_with_gpt_instructions" More Information needed tabular10K<n<100K11 likes148 downloads3y agoHugging Face18olivertzeng /NekoQA-tw What this dataset does A catgirl roleplay QA translated into Traditional Chinese(Taiwan) forked from NekoQA-10k Motivation There aren't a lot of Traditional Chinese(Taiwan) datasets on huggingface and it pretty much ruins the mood when the AI spits out Simplified Chinese to people from Taiwan(especially those who uses qwen3 as the base training model) How this works It's pretty much well known that Taiwan uses a different variant of Chinese, different… See the full description on the dataset page: https://huggingface.co/datasets/olivertzeng/NekoQA-tw.texttext-generation10K<n<100K0 likes147 downloads9mo agoHugging Face19OliveiraJLT /gigaverbo-v2-rec-sft GigaVerbo-v2 REC SFT A model should not merely know how to reason; it should learn when reasoning is worth the cost. Dataset repository: OliveiraJLT/gigaverbo-v2-rec-sftBase dataset: Polygl0t/gigaverbo-v2-sftAnswer-generation model: openai/gpt-oss-20bQuality classifier: Polygl0t/portuguese-qwen3-4b-instruct-quality-classifierReasoning translation model and token accounting tokenizer: Qwen/Qwen3.5-9B Dataset Summary GigaVerbo-v2 REC SFT — short for GigaVerbo-v2… See the full description on the dataset page: https://huggingface.co/datasets/OliveiraJLT/gigaverbo-v2-rec-sft.tabulartext-generation100K<n<1M0 likes138 downloads4mo agoHugging Face20oliversayshi /big-finance-benchmark BigFinanceBench Public Release arXiv | Website | GitHub | Blog post Finance answers are only useful when another analyst can audit how they were produced. BigFinanceBench evaluates that full workflow: agents must produce a numerical answer, and their traces are graded against point-weighted rubrics for source choice, period, accounting definition, assumptions, adjustments, and calculation. This release contains a 50-question stratified subset of the 928-item BigFinanceBench… See the full description on the dataset page: https://huggingface.co/datasets/oliversayshi/big-finance-benchmark.textquestion-answeringn<1K0 likes130 downloads2mo agoHugging Face21oliverkinch /multi-wiki-qa-high-quality-subset multi-wiki-qa-high-quality-subset A quality-filtered subset of the Danish (da) split of alexandrainst/multi-wiki-qa, a Wikipedia-based extractive question-answering dataset. Configs Config Samples Description da 4,767 All LLM-verified correct samples da-short 3,527 Correct samples where the answer is at most 3 words Filtering methodology Starting from the 5,000 samples in the original Danish split: Span validation -- deterministic check that… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/multi-wiki-qa-high-quality-subset.textquestion-answering1K<n<10K0 likes128 downloads6mo agoHugging Face22SamBP069 /olives-multimodal-dataset1 likes127 downloads1mo agoHugging Face23Olivercool /lsm360videon<1K1 likes127 downloads8d agoHugging Face24olivenet /thermoqa ThermoQA — A Benchmark for Evaluating Thermodynamic Reasoning in Large Language Models ThermoQA evaluates how well large language models can solve engineering thermodynamics problems — from steam table property lookups to multi-step component analysis with exergy destruction. 293 questions across three tiers, all grounded in CoolProp 7.2.0 (IAPWS-IF97 + Helmholtz EOS). No other benchmark covers applied engineering thermodynamics at this depth. Leaderboard (v0.4) All… See the full description on the dataset page: https://huggingface.co/datasets/olivenet/thermoqa.textquestion-answeringn<1K1 likes106 downloads6mo agoHugging Face25olivertzeng /catgirl-zhtw-uncensored What this dataset does catgirl-dataset where the AI acts as a catgirl maid to service the user(referred as master) The problem that the original catgirl-dataset has There aren't a lot of Traditional Chinese(Taiwan) datasets on huggingface and it pretty much ruins the mood when the AI spits out Simplified Chinese to people from Taiwan(especially those who uses qwen3 as the base training model) The original dataset is designed to be censored. For example when the user ask… See the full description on the dataset page: https://huggingface.co/datasets/olivertzeng/catgirl-zhtw-uncensored.text-generation1K<n<10K2 likes95 downloads9mo agoHugging Face26OliverSlivka /itemset-extraction-v2 Itemset Extraction Training Data v2 3-phase training dataset for fine-tuning LLMs to extract frequent itemsets from CSV transaction data. Overview Config Purpose Train Val Format sft SFT with Chain-of-Thought 245 27 messages (ChatML) dpo DPO with real LLM failures 546 60 prompt / chosen / rejected grpo GRPO with Apriori rewards 245 27 prompt / ground_truth Training Pipeline (v2 — council-corrected) Phase 1: SFT-CoT (5 epochs) → Teach… See the full description on the dataset page: https://huggingface.co/datasets/OliverSlivka/itemset-extraction-v2.texttext-generation1K<n<10K0 likes92 downloads6mo agoHugging Face27Anonymous-OLiVES /OLiVES OLiVES: An Outdoor Low-Light Video Benchmark for Enhancement and Segmentation OLiVES is a new benchmark for low-light video enhancement and video object segmentation. It contains over 25,000 aligned normal/low-light frames and over 46,000 video object segmentation annotations. LLVE The video folder names and frame names for aligned frames are identical in input/ (low-light videos) and gt/ (normal-light videos) VOS Please run python VOS_dataset_mapper.py to… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous-OLiVES/OLiVES.imagevideo-to-video10K<n<100K0 likes92 downloads5mo agoHugging Face28oliverdk /impossible_apps_introtextn<1K0 likes90 downloads4mo agoHugging Face29oliveirabruno01 /acebench ACEBench Dataset This repository contains the ACEBench dataset, formatted for evaluating and training tool-using language models. The dataset has been processed into a unified structure, with problem descriptions merged with their corresponding ground-truth rubrics. Notebook used to format the dataset: Open in Colab Dataset Structure The dataset is provided under a single configuration, en, which contains three distinct splits: normal: Standard tool-use scenarios.… See the full description on the dataset page: https://huggingface.co/datasets/oliveirabruno01/acebench.text1K<n<10K1 likes88 downloads1y agoHugging Face30oliver-camp /beatles_maroon5_katyperry_rhcp_nirvana_pre2016_instrumental_no_bass_no_drumsaudion<1K0 likes86 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.