datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
quranic-universal-ayahs
Qur'anic Universal Ayahs
Qur'anic Universal Audio (QUA) is a project that unifies recitations on the internet and generates timing data using forced alignment — community-verified results and constantly expanding dataset.
This dataset pairs ayah by ayah audio with word-level timestamps, DigitalKhatt letter-animation timestamps, and waqf-aware segment data. Repeated words are preserved in text_uthmani and word_timestamps, so the row reflects what the reciter… See the full description on the dataset page: https://huggingface.co/datasets/QUD-Technologies/quranic-universal-ayahs.Treble10-Speech
Treble10-Speech (16 kHz)
The Treble10-Speech dataset is a dataset for automatic speech recognition (ASR), containing pre-convolved speech files using high fidelity room-acoustic simulations from the Treble10-RIR dataset with 10 different furnished rooms: 2 bathrooms, 2 bedrooms, 2 living rooms with hallway, 2 living rooms without hallway, 2 meeting rooms.
The room volumes range between 14 and 46 m3, resulting in reverberation times between 0.17 and 0.84 s.
Examples:… See the full description on the dataset page: https://huggingface.co/datasets/treble-technologies/Treble10-Speech.stack-v3-train
🥞 The Stack v3
What is it?
What is being released
How to download and use it
Dataset statistics
Dataset structure
Dataset creation
Considerations for using the data
Additional information
What is it?
The Stack v3 is the largest, most up-to-date open dataset of source code, crawled directly from GitHub and built to pre-train code LLMs with full-repository context. It is the successor to The Stack v2 and, like its predecessor, is released to make the training… See the full description on the dataset page: https://huggingface.co/datasets/TechnoBaptist/stack-v3-train.Treble10-RIR
Treble10-RIR (32 kHz)
The Treble10-RIR dataset is a dataset for automatic speech recognition (ASR), containing high fidelity room-acoustic simulations from 10 different furnished rooms: 2 bathrooms, 2 bedrooms, 2 living rooms with hallway, 2 living rooms without hallway, 2 meeting rooms.
The room volumes range between 14 and 46 m3, resulting in reverberation times between 0.17 and 0.84 s.
Illustrative plots of the rooms and device included in this dataset may be found in the… See the full description on the dataset page: https://huggingface.co/datasets/treble-technologies/Treble10-RIR.euler-structural-equationsriddle_senseriddle_sense dataset formatted into an alpaca format dataset for instruction tuning LLMs for reasoning capabilities.
claude-merged-traceslibrispeech_asr_sliced
Librispeech Slices
Description
Librispeech is a large corpus of read English utterances derived from the LibriVox public domain audiobook project.
It was assembled to assist in Automatic Speech Recognition tasks, and contains approximately 1000 Hours of utterances recorded at 16kHz.
A subset of the original Librispeech dataset was created to support the development of automatic audio scene creation in the Treble SDK environment.
To better imitate the natural flow of… See the full description on the dataset page: https://huggingface.co/datasets/treble-technologies/librispeech_asr_sliced.quran-alignment-benchmark
Quran Recitation Alignment Benchmark
Audio recordings of Quran recitation with a reviewed word-level ground truth: every recited word, in the order it was recited, with its start and end time, plus the reviewed segmentation and non-Quran regions. This is the corpus behind the Quran Recitation Alignment Benchmark; the task, scoring rules, leaderboard and submission format are documented there, not here.
16 recordings · 357 minutes · 18,421 recited words · Hafs ʿan ʿĀṣim ·… See the full description on the dataset page: https://huggingface.co/datasets/QUD-Technologies/quran-alignment-benchmark.ryan-test-white-tableThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"trossen_subversion": "v1.0",
"robot_type": "trossen_ai_stationary",
"total_episodes": 1,
"total_frames": 893,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/allday-technology/ryan-test-white-table.betty-testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"trossen_subversion": "v1.0",
"robot_type": "trossen_ai_stationary",
"total_episodes": 4,
"total_frames": 3807,
"total_tasks": 1,
"total_videos": 16,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:4"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/allday-technology/betty-test.asia-science-technology-world-bank-science-and-technology-indica
Maldives - Science and Technology
Publisher: World Bank Group · Source: HDX · License: cc-by · Updated: 2026-04-28
Abstract
Contains data from the World Bank's data portal. There is also a consolidated country dataset on HDX.
Technological innovation, often fueled by governments, drives industrial growth and helps raise living standards. Data here aims to shed light on countries technology base: research and development, scientific and technical journal articles… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-science-technology-world-bank-science-and-technology-indica.harmonia-infinity-corpusafrica-owid-access-to-clean-fuels-and-technologies-for-cooking
Access To Clean Fuels And Technologies For Cooking | Africa (Our World in Data) | Africa (Electric Sheep Africa metadata inventory)
Size category: 1K<n<10K - Formats: parquet - Sector: other_unclassified - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-owid-access-to-clean-fuels-and-technologies-for-cooking.harmonia-triples-rust-code-traversal
harmonia-triples-rust
Triples for source rust emitted by the ingest pipeline (current wave: v0.7). Schema: (s, p, o, src) with full provenance per ADR-0011. Pre-HHEC.
Provenance
Each parquet shard carries the full provenance chain per ADR-0011:
s, p, o, src columns (when this is a triples-stage dataset)
src = "<dataset>:<version>:<file>" for triples
Causal registry events recorded at causal_registry/master.jsonl chain
Architecture
Part of Harmonia… See the full description on the dataset page: https://huggingface.co/datasets/Hailstone-Technologies/harmonia-triples-rust-code-traversal.touch-cokeThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"trossen_subversion": "v1.0",
"robot_type": "trossen_ai_stationary",
"total_episodes": 10,
"total_frames": 2133,
"total_tasks": 1,
"total_videos": 40,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/allday-technology/touch-coke.fix-handle-angle-leftThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"trossen_subversion": "v1.0",
"robot_type": "trossen_ai_stationary",
"total_episodes": 10,
"total_frames": 2256,
"total_tasks": 1,
"total_videos": 40,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/allday-technology/fix-handle-angle-left.Locate-pickup-measuring-cup-2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"trossen_subversion": "v1.0",
"robot_type": "trossen_ai_stationary",
"total_episodes": 30,
"total_frames": 7048,
"total_tasks": 1,
"total_videos": 120,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/allday-technology/Locate-pickup-measuring-cup-2.pickup-sauce-bottle-smolvlaThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "bi_widowxai_follower",
"total_episodes": 10,
"total_frames": 2950,
"total_tasks": 1,
"total_videos": 40,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/allday-technology/pickup-sauce-bottle-smolvla.pick-up-color-measuring-cupThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"trossen_subversion": "v1.0",
"robot_type": "trossen_ai_stationary",
"total_episodes": 1,
"total_frames": 297,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/allday-technology/pick-up-color-measuring-cup.touch-coke-moreThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"trossen_subversion": "v1.0",
"robot_type": "trossen_ai_stationary",
"total_episodes": 1,
"total_frames": 239,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/allday-technology/touch-coke-more.cupHolderTrainingDataThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"trossen_subversion": "v1.0",
"robot_type": "trossen_ai_stationary",
"total_episodes": 1,
"total_frames": 297,
"total_tasks": 1,
"total_videos": 3,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/allday-technology/cupHolderTrainingData.Locate-pickup-soda-canThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"trossen_subversion": "v1.0",
"robot_type": "trossen_ai_stationary",
"total_episodes": 2,
"total_frames": 295,
"total_tasks": 1,
"total_videos": 8,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/allday-technology/Locate-pickup-soda-can.chatdoctor-embedded
Chat Doctor with Embeddings
This dataset is post-processed version of xzuyn/chatdoctor-200k-stripped:
Add embeddings for input and output columns using BAAI/bge-small-en-v1.5
Details
Sample Count
414k
Token Count
1.7b
Origin
https://drive.google.com/file/d/1lyfqIwlLSClhgrCutWuEe_IACNq6XNUt/view
Source of raw data
?
Processing details
paper
Embedding Model
BAAI/bge-small-en-v1.5
Data Diversity
index
Example Output
GPT-4 Rationale
GPT-4… See the full description on the dataset page: https://huggingface.co/datasets/Technoculture/chatdoctor-embedded.Locate-pickup-measuring-cupThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"trossen_subversion": "v1.0",
"robot_type": "trossen_ai_stationary",
"total_episodes": 1,
"total_frames": 147,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/allday-technology/Locate-pickup-measuring-cup.harmonia-graph-causal
harmonia-graph-causal
Causal graph triples from structured sources: bnlearn Bayesian networks, Reactome pathways, STRING protein interactions, SIGNOR signaling, Wikidata causal properties, ConceptNet causal relations, Tübingen cause-effect pairs, and 60+ Harmonia Structural Causal Models covering the full human experience.
Schema
Every row: s (subject), p (predicate), o (object), src (provenance)
src format: dataset:version:file
Part of the Harmonia… See the full description on the dataset page: https://huggingface.co/datasets/Hailstone-Technologies/harmonia-graph-causal.touch_coke_fastThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"trossen_subversion": "v1.0",
"robot_type": "trossen_ai_stationary",
"total_episodes": 5,
"total_frames": 743,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/allday-technology/touch_coke_fast.harmonia-triples-stackexchange-document-traversal
harmonia-triples-stackexchange-slice
Triples for source stackexchange-slice emitted by the ingest pipeline (current wave: v0.6). Schema: (s, p, o, src) with full provenance per ADR-0011. Pre-HHEC.
Provenance
Each parquet shard carries the full provenance chain per ADR-0011:
s, p, o, src columns (when this is a triples-stage dataset)
src = "<dataset>:<version>:<file>" for triples
Causal registry events recorded at causal_registry/master.jsonl chain… See the full description on the dataset page: https://huggingface.co/datasets/Hailstone-Technologies/harmonia-triples-stackexchange-document-traversal.pickup_chip_bagThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"trossen_subversion": "v1.0",
"robot_type": "trossen_ai_stationary",
"total_episodes": 1,
"total_frames": 235,
"total_tasks": 1,
"total_videos": 4,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:1"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet"… See the full description on the dataset page: https://huggingface.co/datasets/allday-technology/pickup_chip_bag.soarm_amazing_hand_pick
SO-ARM101 + AmazingHand Grasping Dataset
Teleoperated demonstration dataset for training a grasping policy on an SO-ARM101 arm with an AmazingHand dexterous hand (imitation learning).
Task
Pick up the cube with the dexterous hand — grasp a cube on the table using the dexterous hand.
Hardware & Collection
Component
Description
Follower arm
SO-ARM101, 5 × STS3215 (IDs 1–5); original gripper servo #6 removed
End-effector
AmazingHand, 8 ×… See the full description on the dataset page: https://huggingface.co/datasets/Juxi-Technology/soarm_amazing_hand_pick.
