datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fruit-vegetable-conceptsmsmarco-conceptsconcepts_v2_mergedtooltalk-samples
ToolTalk Samples - High Quality Duplex Speech and Tool-calling in Customer Service Domain
Two people improvise realistic customer-service calls while one operates a live, stateful tool environment—with synchronized speaker-separated audio, tool calls, and outcomes.
▶ Listen to Clean · ▶ Listen to Noisy · Discuss the full dataset
In this sample: 26 calls · 90.7 minutes · 7 sample domains · 207 tool calls
Technical specs: 48 kHz / 32-bit PCM speaker-separated source… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/tooltalk-samples.friend-bench
Can a model — or a human — tell how two people are related from a 20-second clip of how they interact?
🌐 Built on Seamless Interaction
FriendBench is a suite of benchmarks for social perception from thin-slice dyadic
interaction — inferring facts about two people's relationship from a brief clip of how they
interact, built on the Seamless Interaction
dataset. Each released set is a config of this repository.
🎧 Multi-modal — text, audio, and video for every clip
🎯 Objective label —… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/friend-bench.multimodal-peer-collaboration-samples
Multimodal Peer Collaboration Samples - Embodied Map Task with Two Camera Angles
Two non-experts collaborate to build working circuits under asymmetric information: the instructor has the manual, the student has the components, and synchronized audio and dual-camera video capture how shared understanding emerges.
▶ Watch the interactions · See Expert Instruction samples · Discuss the full collection
Sister collection: Expert Instruction, a teacher and a student in… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/multimodal-peer-collaboration-samples.CIEL-Clinical-Concepts-to-ICD-11TD;LR Run:
pip install torch==2.4.1+cu118 torchvision==0.19.2+cu118 torchaudio==2.4.1 --extra-index-url https://download.pytorch.org/whl/cu118
pip install -U packaging setuptools wheel ninja
pip install --no-build-isolation axolotl[flash-attn,deepspeed]
axolotl train axolotl_2_a40_runpod_config.yaml
📚 CIEL to ICD-11 Fine-tuning Dataset
This dataset was created to support the fine-tuning of open-source large language models (LLMs) specialized in ICD-11 terminology mapping.
It… See the full description on the dataset page: https://huggingface.co/datasets/filipelopesmedbr/CIEL-Clinical-Concepts-to-ICD-11.multimodal-expert-instruction-samples
Multimodal Expert Instruction Samples - Musical Instrument Lessons with Channel-separated Audio and Video
A music teacher and a student work through two one-on-one lessons: both voices and both instruments on separate tracks, the student on camera, with the lesson plans, the instructions given to each side and both sides' post-lesson ratings alongside.
▶ Watch the lessons · See Peer Collaboration samples · Discuss the full collection
Sister collection: Peer Collaboration… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/multimodal-expert-instruction-samples.neuronpedia-sae-concepts
Neuronpedia SAE Concepts
Complete extraction of all individual concepts from every Sparse Autoencoder (SAE) released on Neuronpedia, plus all public features from Anthropic's Towards Monosemanticity (2023) and Scaling Monosemanticity (2024) papers.
Quick Start
from datasets import load_dataset
# Full Neuronpedia dataset (77M rows, streaming recommended)
ds = load_dataset("hbe/neuronpedia-sae-concepts", split="train", streaming=True)
# Unique concepts with essential… See the full description on the dataset page: https://huggingface.co/datasets/hbe/neuronpedia-sae-concepts.chess-positions-conceptshuman-conceptsDataset Summary
This dataset provides digitized versions of classic human categorization benchmarks from seminal cognitive psychology studies by Rosch (1973, 1975) and McCloskey & Glucksberg (1978). These datasets capture human judgments about semantic categories and typicality, offering high-fidelity insights into how humans organize conceptual knowledge.
This dataset was released as part of the study "From Tokens to Thoughts: How LLMs and Humans Trade Compression for Meaning"
(Shani et al.… See the full description on the dataset page: https://huggingface.co/datasets/CShani/human-concepts.PromptCoT-2.0-Concepts
🧩 PromptCoT 2.0 – Concepts Dataset
PromptCoT 2.0 Concepts provides the foundational conceptual inputs for large-scale prompt synthesis in mathematics and programming.These concept files are used to generate high-quality synthetic problems through the PromptCoT 2.0 Prompt Generation Model.
📘 Overview
Each file (e.g., math.jsonl, code.jsonl) contains a list of concept prompts that serve as the input for the problem generation stage.By feeding these prompts into the… See the full description on the dataset page: https://huggingface.co/datasets/xl-zhao/PromptCoT-2.0-Concepts.Statements-Of-Federal-Financial-Accounting-Concepts-And-Standards
Statements of Federal Financial Accounting Concepts and Standards
Dataset Summary
This dataset contains document-grounded question-and-answer samples based on the Statements of Federal Financial Accounting Concepts and Standards issued within the Federal accounting framework.
The source material establishes the concepts, principles, definitions, recognition criteria, measurement requirements, presentation standards, and disclosure expectations used in Federal… See the full description on the dataset page: https://huggingface.co/datasets/leeroy-jankins/Statements-Of-Federal-Financial-Accounting-Concepts-And-Standards.openalex-concepts
OpenAlex L1 + L2 Concepts
What this is
A snapshot of OpenAlex's Level 1 (broad fields) and Level 2 (subfields) concepts: 21,739 records across 284 broad fields and 21,455 subfields.
Level
Count
Examples
1
284
Computer science, Physics, Biology, Sociology
2
21,455
Machine learning, Quantum mechanics, Convolutional neural network
Each record:
Column
Type
Description
id
string
OpenAlex concept ID (e.g., C121955636)
name
string
Concept… See the full description on the dataset page: https://huggingface.co/datasets/barissozudogru/openalex-concepts.Multi_Subject_Concepts
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/ImagenHub/Multi_Subject_Concepts.self-oss-instruct-sc2-conceptsadaption-tech-concepts-explained
Adaption Tech Concepts Explained
A High-Quality Instruction Tuning Dataset for Large Language Models
A high-quality instruction tuning dataset designed for fine-tuning Large Language Models (LLMs) to generate clear, structured, and beginner-friendly explanations of technical concepts.
This dataset was enhanced using Adaption's Adaptive Data Platform, which improves instruction quality, response consistency, and educational value for supervised fine-tuning… See the full description on the dataset page: https://huggingface.co/datasets/ujjawalbansal/adaption-tech-concepts-explained.ioai2025-onsite-concepts-hint-descriptionsabstract-concepts-diffusiondb-imagesumls-concepts
Dataset Description
This dataset contains a list of biomedical concepts from the UMLS database.
Data Fields
Each concept has its corresponding CUI ID, Name, Aliases, and definition if available.
ioai2025-onsite-concepts-trainmanumoi-conceptsioai2025-onsite-concepts-testphoto_concepts_dataset_smallioai2025-onsite-concepts-validationioai2025-onsite-concepts-trainsneaker_concepts_prompts
Dataset Card for "sneaker_concepts_prompts"
More Information needed
conceptsrobots-human-concepts
Robots — Human Concepts
Synthetic benchmark for evaluating Concept Bottleneck Models (CBMs) under finer-grained, human-annotated concepts. Same underlying robot images and labels as juliannski/robots-true-concepts, but the foot_shape ground-truth concept is replaced by 6 one-hot subtypes that a human annotator would actually see, modelling concept specification mismatch between annotators and the latent labeling rule.
Generated from
This dataset is the exact… See the full description on the dataset page: https://huggingface.co/datasets/juliannski/robots-human-concepts.mining_concepts
