datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ASMR-Archive-Processed
ASMR-Archive-Processed (WIP)
Update (2026-04-03): This dataset has reached the Hugging Face Public Storage Limit. After contacting support, we were informed that the only option is to pay for a storage expansion. Consequently, updates to this dataset are now suspended.
Work in Progress — expect breaking changes while the pipeline and data layout stabilize.
This dataset contains ASMR audio data sourced from DeliberatorArchiver/asmr-archive-data-01 and… See the full description on the dataset page: https://huggingface.co/datasets/OmniAICreator/ASMR-Archive-Processed.ikea_asm
IKEA ASM (RGB mirror)
This repository is a re-upload of the public RGB distribution of the IKEA ASM
dataset. It preserves the category/view organization of the source archives and
includes the official annotation, calibration, and split files needed to use the
RGB videos.
This is a dataset mirror, not a training or experiment artifact. It does not
contain model checkpoints, extracted features, project code, or training
configuration. The repository currently covers the RGB… See the full description on the dataset page: https://huggingface.co/datasets/NeqCene/ikea_asm.asmr-yt-chaptersasm_co_bench_extasmr-archive-data-01
ASMR Media Archive Storage
This repository contains an archive of ASMR works.
All data in this repository is uploaded for educational and research purposes only. All use is at your own risk.
[!IMPORTANT]
This repository contains >= 64 TiB of files.Git LFS consumes twice as much disk space because of the way it works, so git clone is not recommended. Hugging Face CLI or Python libraries allow you to select and download only a subset of files.
>>> CLICK HERE or on the IMAGE BELOW… See the full description on the dataset page: https://huggingface.co/datasets/DeliberatorArchiver/asmr-archive-data-01.ASMR-Archive-Processed-SFW
ASMR-Archive-Processed-SFW
Overview
This dataset is an “educational” subset of the original OmniAICreator/ASMR-Archive-Processed dataset.
We filtered the original dataset to include only records where the nsfw metadata flag is false.
To maintain the randomness and anonymity of the entries, multiple directories were combined and shuffled.
The nsfw tag in the original dataset is inherited from the tags of the original audio works before they were passed through the… See the full description on the dataset page: https://huggingface.co/datasets/noxwano/ASMR-Archive-Processed-SFW.asm_all_include_mistraldecompile-dataset-large-asmasm_all_experimentsasmr-zh-r18
asmr-zh-r18
Chinese R18 ASMR audio dataset for TTS/voice cloning fine-tuning.
Works: 5876 RJ-coded works
Total size: ~1.1TB raw (10 compressed parts)
Format: MP3/WAV/FLAC, 3 tracks per work
Source: asmr.one API, Chinese R18 category
Extract
for f in packs/*.tar.zst; do
zstd -d "$f" --stdout | tar -xf -
done
asmr-archive-data-02
ASMR Media Archive Storage
This repository contains an archive of ASMR works.
All data in this repository is uploaded for educational and research purposes only. All use is at your own risk.
[!IMPORTANT]
This repository contains >= 64 TiB of files.Git LFS consumes twice as much disk space because of the way it works, so git clone is not recommended. Hugging Face CLI or Python libraries allow you to select and download only a subset of files.
>>> CLICK HERE or on the IMAGE BELOW… See the full description on the dataset page: https://huggingface.co/datasets/DeliberatorArchiver/asmr-archive-data-02.ASM-steer-mistrallinux-asm-pairs
Linux Kernel Assembly → Explanation Dataset
A dataset of disassembled Linux kernel functions paired with structured natural-language explanations, built from the Linux Commits Dataset (Zenodo 10654193).
Each row corresponds to a single compiled function extracted from a .c or .h file touched by a Bug-Fix Commit (BFC) or Bug-Introducing Commit (BIC) in the Linux kernel git history. Source files are compiled with gcc -O2 -g -fno-inline -fno-omit-frame-pointer and disassembled with… See the full description on the dataset page: https://huggingface.co/datasets/theelderemo/linux-asm-pairs.movienet318-keyframes-240pASM_codeasm_cuda_to_amdscene-segmentation-embeddingsosworld_file_cacheasmr
Dataset Card for ASMR Audio Dataset
Dataset Summary
This dataset contains a large collection of ASMR (Autonomous Sensory Meridian Response) audio clips with corresponding machine-generated transcriptions. The dataset includes approximately 283,132 audio segments totaling over 307 hours of content, with an average duration of 3.92 seconds per clip. All audio files are provided in WAV format at 24 kHz sampling rate, making them suitable for various audio processing and… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/asmr.A-S-Mdreadditchunked-asm2asm-fulljapanese_asmrbringup_asm
BringUpBench C and Assembly
This dataset pairs C programs from BringUpBench 1.9 with assembly generated by Clang 17 for four CPU targets and seven optimization levels. It is intended for research on compilation, decompilation, assembly understanding, and cross-architecture code translation.
Dataset structure
The dataset contains 108 programs. Each optimization level is stored as a separate Hugging Face split, with 108 rows per split.
These are compiler… See the full description on the dataset page: https://huggingface.co/datasets/murodbek/bringup_asm.asm2asm_O0_1000000_gnueabi_gcc
Dataset Card for "asm2asm_O0_1000000_gnueabi_gcc"
More Information needed
asm2asm_100000
Dataset Card for "asm2asm_100000"
More Information needed
kakugo-asm
Kakugo Assamese dataset
[Paper] [Code] [Model]
A synthetically generated conversation dataset for training in Assamese.
This dataset contains synthetic conversational data and translated instructions designed to train Small Language Models (SLMs) for Assamese. It was generated using the Kakugo pipeline, a method for distilling high-quality capabilities from a large teacher model into low-resource language models. The teacher model used to generate this dataset was… See the full description on the dataset page: https://huggingface.co/datasets/ptrdvn/kakugo-asm.humaneval_asm
HumanEval-C and Assembly
This dataset pairs C functions from the HumanEval-Decompile benchmark with assembly generated by Clang 17 for four CPU targets and seven optimization levels. It is intended for research on compilation, decompilation, assembly understanding, and cross-architecture code translation.
Dataset structure
The dataset contains 164 programs. Each optimization level is stored as a separate Hugging Face split, with 164 rows per split.
These are… See the full description on the dataset page: https://huggingface.co/datasets/murodbek/humaneval_asm.mceval_asm
McEval-C and Assembly
This dataset pairs the C-language tasks from McEval with assembly generated by Clang 17 for four CPU targets and seven optimization levels. It is intended for research on compilation, decompilation, assembly understanding, and cross-architecture code translation.
Dataset structure
The dataset contains 50 programs. Each optimization level is stored as a separate Hugging Face split, with 50 rows per split.
These are compiler configurations, not… See the full description on the dataset page: https://huggingface.co/datasets/murodbek/mceval_asm.asmr-mami
