free
Datasets
All datasets matching “free”medical-o1-reasoning-SFT
News
[2025/04/22] We split the data and kept only the medical SFT dataset (medical_o1_sft.json). The file medical_o1_sft_mix.json contains a mix of medical and general instruction data.
[2025/02/22] We released the distilled dataset from Deepseek-R1 based on medical verifiable problems. You can use it to initialize your models with the reasoning chain from Deepseek-R1.
[2024/12/25] We open-sourced the medical reasoning dataset for SFT, built on medical verifiable problems and an… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/medical-o1-reasoning-SFT.CLAP_freesound
LAION-Audio-630K Freesound Dataset
LAION-Audio-630K is the largest audio-text dataset publicly available and a magnitude larger than previous audio-text datasets (by 2022-11-05). Notably, it combines eight distinct datasets, which includes the Freesound dataset.
Specifically, this Hugging face repository contains two versions of Freesound dataset. Details of each dataset (e.g. how captions are made etc.) could be found in the "datacard" column of the table below.
Freesound (full):… See the full description on the dataset page: https://huggingface.co/datasets/Meranti/CLAP_freesound.cad-gen-freecad-bench
Parametric CAD Bench — results dataset
Run-by-run results for Parametric CAD Bench, a benchmark that
measures whether AI agents can author editable FreeCAD models from
natural-language part descriptions. 1000 rows, one per
(agent, model, task_id, trial) over the
gnucleus-ai/cad-bench@v1
task suite. The public leaderboard view of this data lives at
cadbench.ai.
What's in here
data/cad-bench-v1.parquet — the row table. Each row carries the
composite + sub-scores… See the full description on the dataset page: https://huggingface.co/datasets/gnucleus-ai/cad-gen-freecad-bench.MMLU_ChineseChinese version of MMLU dataset tranlasted by gpt-3.5-turbo.The dataset is used in the research related to MultilingualSIFT.
FreeTacMan
📦 FreeTacman
Robot-free Visuo-Tactile Data Collection System for Contact-rich Manipulation [ICRA 2026]
🎯 Overview
This dataset supports the paper FreeTacman: Robot-free Visuo-Tactile Data Collection System for Contact-rich Manipulation.
It contains a large-scale, high-precision visuo-tactile manipulation dataset with over 3000k visuo-tactile image pairs, more than 10k trajectories across 50 tasks.
We provide 🤗 Script (Hugging Face) and 👾 Script… See the full description on the dataset page: https://huggingface.co/datasets/OpenDriveLab/FreeTacMan.BlenderLore
The target scope is 22,360 video-associated Blender project instances, not 22,360 distinct tutorial videos. Uploads are in progress, so the currently published files may be a subset of this target. The 44 biomedical project instances and one software-bundled Dome template are excluded.
Data Structure
The dataset is organized as a collection of sample-level directories under assets/. Each directory corresponds to one Blender creation task and follows the structure below:… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/BlenderLore.
