datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
medical-o1-reasoning-SFT
News
[2025/04/22] We split the data and kept only the medical SFT dataset (medical_o1_sft.json). The file medical_o1_sft_mix.json contains a mix of medical and general instruction data.
[2025/02/22] We released the distilled dataset from Deepseek-R1 based on medical verifiable problems. You can use it to initialize your models with the reasoning chain from Deepseek-R1.
[2024/12/25] We open-sourced the medical reasoning dataset for SFT, built on medical verifiable problems and an… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/medical-o1-reasoning-SFT.CLAP_freesound
LAION-Audio-630K Freesound Dataset
LAION-Audio-630K is the largest audio-text dataset publicly available and a magnitude larger than previous audio-text datasets (by 2022-11-05). Notably, it combines eight distinct datasets, which includes the Freesound dataset.
Specifically, this Hugging face repository contains two versions of Freesound dataset. Details of each dataset (e.g. how captions are made etc.) could be found in the "datacard" column of the table below.
Freesound (full):… See the full description on the dataset page: https://huggingface.co/datasets/Meranti/CLAP_freesound.cad-gen-freecad-bench
Parametric CAD Bench — results dataset
Run-by-run results for Parametric CAD Bench, a benchmark that
measures whether AI agents can author editable FreeCAD models from
natural-language part descriptions. 1000 rows, one per
(agent, model, task_id, trial) over the
gnucleus-ai/cad-bench@v1
task suite. The public leaderboard view of this data lives at
cadbench.ai.
What's in here
data/cad-bench-v1.parquet — the row table. Each row carries the
composite + sub-scores… See the full description on the dataset page: https://huggingface.co/datasets/gnucleus-ai/cad-gen-freecad-bench.MMLU_ChineseChinese version of MMLU dataset tranlasted by gpt-3.5-turbo.The dataset is used in the research related to MultilingualSIFT.
FreeTacMan
📦 FreeTacman
Robot-free Visuo-Tactile Data Collection System for Contact-rich Manipulation [ICRA 2026]
🎯 Overview
This dataset supports the paper FreeTacman: Robot-free Visuo-Tactile Data Collection System for Contact-rich Manipulation.
It contains a large-scale, high-precision visuo-tactile manipulation dataset with over 3000k visuo-tactile image pairs, more than 10k trajectories across 50 tasks.
We provide 🤗 Script (Hugging Face) and 👾 Script… See the full description on the dataset page: https://huggingface.co/datasets/OpenDriveLab/FreeTacMan.BlenderLore
The target scope is 22,360 video-associated Blender project instances, not 22,360 distinct tutorial videos. Uploads are in progress, so the currently published files may be a subset of this target. The 44 biomedical project instances and one software-bundled Dome template are excluded.
Data Structure
The dataset is organized as a collection of sample-level directories under assets/. Each directory corresponds to one Blender creation task and follows the structure below:… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/BlenderLore.ShareGPT-4o-Image
📚 ShareGPT-4o-Image
ShareGPT-4o-Image is a large-scale and high-quality image generation dataset, where all images are produced by GPT-4o’s image generation capabilities. This dataset is designed to align open multimodal models with GPT-4o’s strengths in visual content creation. It includes 45K text-to-image and 46K text-and-image-to-image samples, making it a useful resource for enhancing multimodal models in both image generation and editing tasks.
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ShareGPT-4o-Image.FreeMan
FreeMan: Towards 3D Human Pose Estimation In the Wild
🌏 Project Page •
🙋♂️ Download Procedure •
📄 Paper •
▶️ YouTube •
🖥️ Code
This is official release of FreeMan dataset. To access the dataset, please submit previous steps at HERE.
[❗️❗️❗️] MAKE SURE you finish required steps HERE before apply dataset access here. Otherwise access request will NOT be approved.For who are in mainland China, you can also apply & download from… See the full description on the dataset page: https://huggingface.co/datasets/wjwow/FreeMan.free-music-archive-full
FMA: A Dataset for Music Analysis
Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson.
International Society for Music Information Retrieval Conference (ISMIR), 2017.
We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/free-music-archive-full.CMB
CMB: A Comprehensive Medical Benchmark in Chinese
🌐 Github • 🌐 Website • 🤗 HuggingFace
🌈 Update
[2024.02.21] The answers to the CMB-Exam test has been updated and some errors caused by omissions in version management have been fixed.
[2024.01.08] In order to facilitate testing, we disclose the answers to the CMB-Exam test
[2023.09.22] CMB is included in OpenCompass.
[2023.08.21] Paper released.
[2023.08.01] 🎉🎉🎉 CMB is published!🎉🎉🎉
🌐… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/CMB.TalkVid
TalkVid Dataset
This repository hosts the TalkVid dataset.
Paper: TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis
Arxiv paper: https://arxiv.org/abs/2508.13618
Project Page: https://freedomintelligence.github.io/talk-vid
GitHub: https://github.com/FreedomIntelligence/TalkVid
Abstract
Audio-driven talking head synthesis has achieved remarkable photorealism, yet state-of-the-art (SOTA) models exhibit a critical failure: they lack… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/TalkVid.cad-gen-freecad
CAD Generation Dataset
Each row in this dataset describes one parametric CAD part. Columns:
id — row identifier (also the basename of the per-row asset files)
name — part family (e.g. flanges, spur_gear_stock)
description — natural-language description of the geometry
key_parameters — the dimensions that drive the parametric model
image — 512×512 PNG preview rendered from the FCStd
fcstd_path — relative path inside this repo to the parametric FreeCAD document (fcstd/<id>.FCStd)… See the full description on the dataset page: https://huggingface.co/datasets/gnucleus-ai/cad-gen-freecad.IG_CACHE_I3D_Freeze_1freesound-laion-640k
About this Repository
This repository is a re-upload of the FreeSound.org dataset as curated by LAION for the larger LAION-Audio-630k dataset, with the following changes:
Limited columns to only the audio and basic metadata.
Incorporated necessary information for licensing and attribution.
Removed ambiguously licensed samples, amounting to around 1,000 total samples.
What about download links?
Links were ommitted for the sake of size, as they can be constructed from… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/freesound-laion-640k.old-comfyui-freezebanking-concentration-data
Banking Concentration Data (1959–2025)
Public portion of the source data for the banking-concentration visualizations
project at https://github.com/aaron-freedman/banking-concentration. This dataset
assembles regulatory filings from three US bank and thrift supervisors covering
1959–2025 into a single, loadable collection of parquet and comma-separated
values (CSV) files.
Important: three additional files are needed by the full preprocessing
pipeline but are not included here… See the full description on the dataset page: https://huggingface.co/datasets/aaron-freedman/banking-concentration-data.cad-gen-freecad-bench-v2
Parametric CAD Bench v2 — results dataset
Run-by-run results for Parametric CAD Bench v2, a benchmark that measures
whether AI agents can author editable FreeCAD models from natural-language part
descriptions. This archive contains 1,000 rows: one trial for each of 100 tasks
across the 10 public jobs on the live
gnucleus-ai/cad-bench@v2
leaderboard.
What's in here
data/cad-bench-v2.parquet — the trial index. Each row carries the
continuous reward and its… See the full description on the dataset page: https://huggingface.co/datasets/gnucleus-ai/cad-gen-freecad-bench-v2.alpaca-gpt4-chineseThe dataset is used in the research related to MultilingualSIFT.
anima-tagger-artifacts
anima-tagger-artifacts
Pre-built retrieval artefacts for the Anima tag format of the
sd-webui-prompt-enhancer
Stable Diffusion WebUI extension. Lets the extension's Anima pipeline
do real-time embedding-based tag validation and shortlist retrieval
without users needing to rebuild a 270k+ entry FAISS index locally.
Contents
File
Size
Description
tags.sqlite
~30 MB
273,025 Danbooru tags (name, category, post count, aliases, wiki). Post-count floor 10.… See the full description on the dataset page: https://huggingface.co/datasets/freedumb2000/anima-tagger-artifacts.MileBench
MileBench
Introduction
We introduce MileBench, a pioneering benchmark designed to test the MultImodal Long-contExt capabilities of MLLMs.
This benchmark comprises not only multimodal long contexts, but also multiple tasks requiring both comprehension and generation.
We establish two distinct evaluation sets, diagnostic and realistic, to systematically assess MLLMs’ long-context adaptation capacity and their ability to completetasks in long-context scenarios
To… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/MileBench.freesound-laion-640k-commercial-16khz-full
About this Repository
This repository is the training split of the complete FreeSound LAION 640k dataset, limited only to licenses that permit commercial works, resampled to 16khz using torchaudio.transforms.Resample.
This is ideal for use cases where a variety of audio is desired but fidelity and labels are unnecessary, such as background audio for augmenting other datasets.
Dataset Versions
You are looking at the full dataset which contains 403,146 unique sounds… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/freesound-laion-640k-commercial-16khz-full.free-music-archive-medium
FMA: A Dataset for Music Analysis
Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson.
International Society for Music Information Retrieval Conference (ISMIR), 2017.
We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/free-music-archive-medium.freebase
Freebase
Dataset Description
Large-scale knowledge base (archived by Google)
Original Source: http://commondatastorage.googleapis.com/freebase-public/rdf/freebase-rdf-latest.gz
Dataset Summary
This dataset contains RDF triples from Freebase converted to HuggingFace
dataset format for easy use in machine learning pipelines.
Format: Originally ntriples, converted to HuggingFace Dataset
Size: 300.0 GB (extracted)
Entities: ~50M
Triples: ~3B
Original License:
CC… See the full description on the dataset page: https://huggingface.co/datasets/CleverThis/freebase.PubMedVision
News
[2025/02/18]: We add the original captions of PubMedVision in PubMedVision_Original_Caption.json, as well as the Chinese version of PubMedVision in PubMedVision_Chinese.json.
[2024/07/01]: We add annotations for 'body_part' and 'modality' of images, utilizing the HuatuoGPT-Vision-7B model.
PubMedVision
PubMedVision is a large-scale medical VQA dataset. We extracted high-quality image-text pairs from PubMed and used GPT-4V to reformat them to enhance their quality.… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/PubMedVision.cad-bench-freecad-agent-initial-runsfree-music-archive-small
FMA: A Dataset for Music Analysis
Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson.
International Society for Music Information Retrieval Conference (ISMIR), 2017.
We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/free-music-archive-small.freelancer-projects-100k-tracesfree-music-archive-large
FMA: A Dataset for Music Analysis
Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson.
International Society for Music Information Retrieval Conference (ISMIR), 2017.
We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/free-music-archive-large.SciFigPlag-Bench
SciFigPlag-Bench: A Benchmark for Provenance-Aware Scientific Figure Plagiarism Detection
🌐 Homepage |
📖 arXiv
SciFigPlag-Bench is a benchmark for provenance-aware scientific figure plagiarism detection. It evaluates whether a suspicious figure reuses evidence from a specific source figure, which figure is the original source, how the reused content has been transformed, and where the reused evidence appears.
The benchmark is designed to evaluate vision-language models… See the full description on the dataset page: https://huggingface.co/datasets/FreeLand123/SciFigPlag-Bench.ApolloCorpus
Multilingual Medicine: Model, Dataset, Benchmark, Code
Covering English, Chinese, French, Hindi, Spanish, Hindi, Arabic So far
👨🏻💻Github •📃 Paper • 🌐 Demo • 🤗 ApolloCorpus • 🤗 XMedBench
中文 | English
🌈 Update
[2024.03.07] Paper released.
[2024.02.12] ApolloCorpus and XMedBench is published!🎉
[2024.01.23] Apollo repo is published!🎉
Results
Apollo-0.5B • 🤗 Apollo-1.8B • 🤗 Apollo-2B • 🤗 Apollo-6B • 🤗 Apollo-7B… See the full description on the dataset page: https://huggingface.co/datasets/FreedomIntelligence/ApolloCorpus.
