datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
OLIVES_Dataset
OLIVES_Dataset
Abstract
Clinical diagnosis of the eye is performed over multifarious data modalities including scalar clinical labels, vectorized biomarkers, two-dimensional fundus images, and three-dimensional Optical Coherence Tomography (OCT) scans. While the clinical labels, fundus images and OCT scans are instrumental measurements, the vectorized biomarkers are interpreted attributes from the other measurements. Clinical practitioners use all these data modalities… See the full description on the dataset page: https://huggingface.co/datasets/gOLIVES/OLIVES_Dataset.Real-3DQA
Real-3DQA
Do 3D Large Language Models Really Understand 3D Spatial Relationships?
🌐 Project Page · 📄 Paper · 💻 GitHub
Overview
Real-3DQA is a debiased 3D spatial QA benchmark with viewpoint rotation consistency evaluation. It addresses two key shortcomings of existing benchmarks:
Language Shortcut Filtering — Questions answerable through linguistic priors alone are removed by comparing 3D-LLMs against blind text-only counterparts.
Viewpoint Rotation Score (VRS) — Each… See the full description on the dataset page: https://huggingface.co/datasets/Oliver-Ma/Real-3DQA.OLIVES_Dataset
OLIVES_Dataset
Abstract
Clinical diagnosis of the eye is performed over multifarious data modalities including scalar clinical labels, vectorized biomarkers, two-dimensional fundus images, and three-dimensional Optical Coherence Tomography (OCT) scans. While the clinical labels, fundus images and OCT scans are instrumental measurements, the vectorized biomarkers are interpreted attributes from the other measurements. Clinical practitioners use all these data modalities… See the full description on the dataset page: https://huggingface.co/datasets/SOHAIBSUL123/OLIVES_Dataset.crafter-cache-81f-256pxml-lectures
ML/Math Lecture Archive
Archived lecture videos (1080p MP4) with English subtitles (.vtt) from publicly
available university course recordings on YouTube.
268 videos, ~72 GB.
Contents
Folder
Course
Videos
18.065_Strang/
MIT 18.065 — Matrix Methods in Data Analysis, Signal Processing, and Machine Learning (Gilbert Strang)
36
18.06SC_LinearAlgebra/
MIT 18.06SC — Linear Algebra, Fall 2011 (Gilbert Strang)
74
CS109_Piech/
Stanford CS109 — Introduction… See the full description on the dataset page: https://huggingface.co/datasets/olive5/ml-lectures.IDPP-Data
IDPP-Data
Video archives for IDPP / SVD physics-benchmark reproduction: friction, viscosity, and elasticity, each with absolute and relative splits (train / test1 / test2 / test3 where applicable).
Repository: Oliver-Ma/IDPP-DataSnapshot (README revision): 2026-04-12 — documents 26 .zip blobs under video_zips/ (mirroring local dataset/<lane>/).
Important: archives, not frame folders
Everything under video_zips/ is a .zip file. Training and evaluation code (e.g.… See the full description on the dataset page: https://huggingface.co/datasets/Oliver-Ma/IDPP-Data.da-bird
DA-BIRD: Danish NL2SQL Benchmark
DA-BIRD is a Danish text-to-SQL benchmark for evaluating large language models on natural language to SQL generation. All tasks are in Danish and use SQLite databases. The dataset is designed for use with the Harbor evaluation framework.
It combines two corpora — 363 tasks across 22 unique databases:
Corpus
Tasks
Databases
Language
Difficulty
bird_*
150
11
Danish (translated)
easy / medium / hard
dst_*
213
11
Danish (original)
medium /… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/da-bird.ProGAN-Eval
ProGAN eval set
This repository hosts the test set used in the FPBA on the LSUN Synthesis dataset, consisting of 1,000 test images from the LSUN dataset and 1,000 images synthesized by ProGAN.
shelf_jimUSE FIRST COMMIT FOR ORIGINAL DATASET. LATEST COMMIT IS MAHATHI-SHELF DATASET
nl_gameable_programmatic_graderslibero_90_sawyer_defaultCams_failuresThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "sawyer",
"total_episodes": 4278,
"total_frames": 637138,
"total_tasks": 74,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 20,
"splits": {
"train": "0:4278"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/OliverHausdoerfer/libero_90_sawyer_defaultCams_failures.mahathi_bottledreamvla_il_dataimpossible_mbpp_natural_diverse3DRSHumanoidRobotSoccer
Fall Prediction Dataset for Humanoid Robots
Dataset Summary
This dataset consists of 37.9 hours of real-world sensor data collected from 20 Nao humanoid robots over the course of one year in various test environments, including RoboCup soccer matches. The dataset includes 18.3 hours of walking data, featuring 2519 falls. It captures a wide range of activities such as omni-directional walking, collisions, standing up, and falls on various surfaces like artificial turf and… See the full description on the dataset page: https://huggingface.co/datasets/OliverUrbann/HumanoidRobotSoccer.news_with_gpt_instructions
Dataset Card for "news_with_gpt_instructions"
More Information needed
NekoQA-tw
What this dataset does
A catgirl roleplay QA translated into Traditional Chinese(Taiwan) forked from
NekoQA-10k
Motivation
There aren't a lot of Traditional Chinese(Taiwan) datasets on huggingface and it pretty much
ruins the mood when the AI spits out Simplified Chinese to people from Taiwan(especially those
who uses qwen3 as the base training model)
How this works
It's pretty much well known that Taiwan uses a different variant of Chinese, different… See the full description on the dataset page: https://huggingface.co/datasets/olivertzeng/NekoQA-tw.gigaverbo-v2-rec-sft
GigaVerbo-v2 REC SFT
A model should not merely know how to reason; it should learn when reasoning is worth the cost.
Dataset repository: OliveiraJLT/gigaverbo-v2-rec-sftBase dataset: Polygl0t/gigaverbo-v2-sftAnswer-generation model: openai/gpt-oss-20bQuality classifier: Polygl0t/portuguese-qwen3-4b-instruct-quality-classifierReasoning translation model and token accounting tokenizer: Qwen/Qwen3.5-9B
Dataset Summary
GigaVerbo-v2 REC SFT — short for GigaVerbo-v2… See the full description on the dataset page: https://huggingface.co/datasets/OliveiraJLT/gigaverbo-v2-rec-sft.big-finance-benchmark
BigFinanceBench Public Release
arXiv | Website | GitHub | Blog post
Finance answers are only useful when another analyst can audit how they were produced. BigFinanceBench evaluates that full workflow: agents must produce a numerical answer, and their traces are graded against point-weighted rubrics for source choice, period, accounting definition, assumptions, adjustments, and calculation.
This release contains a 50-question stratified subset of the 928-item BigFinanceBench… See the full description on the dataset page: https://huggingface.co/datasets/oliversayshi/big-finance-benchmark.multi-wiki-qa-high-quality-subset
multi-wiki-qa-high-quality-subset
A quality-filtered subset of the Danish (da) split of
alexandrainst/multi-wiki-qa,
a Wikipedia-based extractive question-answering dataset.
Configs
Config
Samples
Description
da
4,767
All LLM-verified correct samples
da-short
3,527
Correct samples where the answer is at most 3 words
Filtering methodology
Starting from the 5,000 samples in the original Danish split:
Span validation -- deterministic check that… See the full description on the dataset page: https://huggingface.co/datasets/oliverkinch/multi-wiki-qa-high-quality-subset.olives-multimodal-datasetlsm360thermoqa
ThermoQA — A Benchmark for Evaluating Thermodynamic Reasoning in Large Language Models
ThermoQA evaluates how well large language models can solve
engineering thermodynamics problems — from steam table property
lookups to multi-step component analysis with exergy destruction.
293 questions across three tiers, all grounded in CoolProp 7.2.0
(IAPWS-IF97 + Helmholtz EOS). No other benchmark covers applied
engineering thermodynamics at this depth.
Leaderboard (v0.4)
All… See the full description on the dataset page: https://huggingface.co/datasets/olivenet/thermoqa.catgirl-zhtw-uncensored
What this dataset does
catgirl-dataset where the AI acts as a
catgirl maid to service the user(referred as master)
The problem that the original catgirl-dataset has
There aren't a lot of Traditional Chinese(Taiwan) datasets on huggingface and it pretty much
ruins the mood when the AI spits out Simplified Chinese to people from Taiwan(especially those
who uses qwen3 as the base training model)
The original dataset is designed to be censored. For example when the user ask… See the full description on the dataset page: https://huggingface.co/datasets/olivertzeng/catgirl-zhtw-uncensored.itemset-extraction-v2
Itemset Extraction Training Data v2
3-phase training dataset for fine-tuning LLMs to extract frequent itemsets from CSV transaction data.
Overview
Config
Purpose
Train
Val
Format
sft
SFT with Chain-of-Thought
245
27
messages (ChatML)
dpo
DPO with real LLM failures
546
60
prompt / chosen / rejected
grpo
GRPO with Apriori rewards
245
27
prompt / ground_truth
Training Pipeline (v2 — council-corrected)
Phase 1: SFT-CoT (5 epochs) → Teach… See the full description on the dataset page: https://huggingface.co/datasets/OliverSlivka/itemset-extraction-v2.OLiVES
OLiVES: An Outdoor Low-Light Video Benchmark for Enhancement and Segmentation
OLiVES is a new benchmark for low-light video enhancement and video object segmentation. It contains over 25,000 aligned normal/low-light frames and over 46,000 video object segmentation annotations.
LLVE
The video folder names and frame names for aligned frames are identical in input/ (low-light videos) and gt/ (normal-light videos)
VOS
Please run python VOS_dataset_mapper.py to… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous-OLiVES/OLiVES.impossible_apps_introacebench
ACEBench Dataset
This repository contains the ACEBench dataset, formatted for evaluating and training tool-using language models. The dataset has been processed into a unified structure, with problem descriptions merged with their corresponding ground-truth rubrics.
Notebook used to format the dataset: Open in Colab
Dataset Structure
The dataset is provided under a single configuration, en, which contains three distinct splits:
normal: Standard tool-use scenarios.… See the full description on the dataset page: https://huggingface.co/datasets/oliveirabruno01/acebench.beatles_maroon5_katyperry_rhcp_nirvana_pre2016_instrumental_no_bass_no_drums
