CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01noxwano /ASMR-Archive-Processed-SFW ASMR-Archive-Processed-SFW Overview This dataset is an “educational” subset of the original OmniAICreator/ASMR-Archive-Processed dataset. We filtered the original dataset to include only records where the nsfw metadata flag is false. To maintain the randomness and anonymity of the entries, multiple directories were combined and shuffled. The nsfw tag in the original dataset is inherited from the tags of the original audio works before they were passed through the… See the full description on the dataset page: https://huggingface.co/datasets/noxwano/ASMR-Archive-Processed-SFW.audioautomatic-speech-recognition1M<n<10M9 likes730 downloads5mo agoHugging Face02Neo111x /decompile-dataset-large-asmtext1M<n<10M3 likes597 downloads2y agoHugging Face03theelderemo /linux-asm-pairs Linux Kernel Assembly → Explanation Dataset A dataset of disassembled Linux kernel functions paired with structured natural-language explanations, built from the Linux Commits Dataset (Zenodo 10654193). Each row corresponds to a single compiled function extracted from a .c or .h file touched by a Bug-Fix Commit (BFC) or Bug-Introducing Commit (BIC) in the Linux kernel git history. Source files are compiled with gcc -O2 -g -fno-inline -fno-omit-frame-pointer and disassembled with… See the full description on the dataset page: https://huggingface.co/datasets/theelderemo/linux-asm-pairs.texttext-generation1K<n<10K0 likes242 downloads5mo agoHugging Face04bxiong /ASM_codetabularn<1K0 likes141 downloads5mo agoHugging Face05ahmedheakl /asm_cuda_to_amdtabular10K<n<100K1 likes126 downloads2y agoHugging Face06nyuuzyou /asmr Dataset Card for ASMR Audio Dataset Dataset Summary This dataset contains a large collection of ASMR (Autonomous Sensory Meridian Response) audio clips with corresponding machine-generated transcriptions. The dataset includes approximately 283,132 audio segments totaling over 307 hours of content, with an average duration of 3.92 seconds per clip. All audio files are provided in WAV format at 24 kHz sampling rate, making them suitable for various audio processing and… See the full description on the dataset page: https://huggingface.co/datasets/nyuuzyou/asmr.textautomatic-speech-recognition100K<n<1M5 likes80 downloads1y agoHugging Face07luckyor /A-S-Mtabularn<1K0 likes78 downloads18d agoHugging Face08ahmedheakl /asm2asm_mac_x86_datatext100K<n<1M0 likes74 downloads2y agoHugging Face09asmaab /dreaddittabular1K<n<10K3 likes71 downloads2y agoHugging Face10Raniahossam33 /chunked-asm2asm-fulltext100K<n<1M0 likes70 downloads2y agoHugging Face11murodbek /bringup_asm BringUpBench C and Assembly This dataset pairs C programs from BringUpBench 1.9 with assembly generated by Clang 17 for four CPU targets and seven optimization levels. It is intended for research on compilation, decompilation, assembly understanding, and cross-architecture code translation. Dataset structure The dataset contains 108 programs. Each optimization level is stored as a separate Hugging Face split, with 108 rows per split. These are compiler… See the full description on the dataset page: https://huggingface.co/datasets/murodbek/bringup_asm.texttext-generationn<1K0 likes65 downloads24d agoHugging Face12ahmedheakl /asm2asm_O0_1000000_gnueabi_gcc Dataset Card for "asm2asm_O0_1000000_gnueabi_gcc" More Information needed text100K<n<1M0 likes63 downloads2y agoHugging Face13ahmedheakl /asm2asm_100000 Dataset Card for "asm2asm_100000" More Information needed text10K<n<100K0 likes58 downloads2y agoHugging Face14ptrdvn /kakugo-asm Kakugo Assamese dataset [Paper] [Code] [Model] A synthetically generated conversation dataset for training in Assamese. This dataset contains synthetic conversational data and translated instructions designed to train Small Language Models (SLMs) for Assamese. It was generated using the Kakugo pipeline, a method for distilling high-quality capabilities from a large teacher model into low-resource language models. The teacher model used to generate this dataset was… See the full description on the dataset page: https://huggingface.co/datasets/ptrdvn/kakugo-asm.texttext-generation10K<n<100K0 likes56 downloads8mo agoHugging Face15murodbek /humaneval_asm HumanEval-C and Assembly This dataset pairs C functions from the HumanEval-Decompile benchmark with assembly generated by Clang 17 for four CPU targets and seven optimization levels. It is intended for research on compilation, decompilation, assembly understanding, and cross-architecture code translation. Dataset structure The dataset contains 164 programs. Each optimization level is stored as a separate Hugging Face split, with 164 rows per split. These are… See the full description on the dataset page: https://huggingface.co/datasets/murodbek/humaneval_asm.texttext-generation1K<n<10K0 likes55 downloads24d agoHugging Face16murodbek /mceval_asm McEval-C and Assembly This dataset pairs the C-language tasks from McEval with assembly generated by Clang 17 for four CPU targets and seven optimization levels. It is intended for research on compilation, decompilation, assembly understanding, and cross-architecture code translation. Dataset structure The dataset contains 50 programs. Each optimization level is stored as a separate Hugging Face split, with 50 rows per split. These are compiler configurations, not… See the full description on the dataset page: https://huggingface.co/datasets/murodbek/mceval_asm.texttext-generationn<1K0 likes54 downloads24d agoHugging Face17ahmedheakl /asm2asm_O0_500000_risctext100K<n<1M0 likes51 downloads2y agoHugging Face18ahmedheakl /asm2asm_O0_500000_armv8text100K<n<1M0 likes47 downloads2y agoHugging Face19kontextox /uk_UA-ASMR Ukrainian ASMR TTS Dataset A Ukrainian text-to-speech dataset for training single-speaker ASMR-style voice models using Piper. Dataset Details Property Value Language Ukrainian (uk_UA) Speakers 1 Segments 7,318 Audio Format 16-bit WAV, 22050 Hz, Mono License CC0 Dataset Structure Prerequisites # Install Piper training dependencies git clone https://github.com/kontextox/piper1-gpl.git cd piper1-gpl python3 -m venv .venv source… See the full description on the dataset page: https://huggingface.co/datasets/kontextox/uk_UA-ASMR.text1K<n<10K0 likes46 downloads6mo agoHugging Face20Raniahossam33 /aligned-asm2asm-train-completetext1M<n<10M0 likes44 downloads2y agoHugging Face21reddest-panda /arm-asmtexttext-generation1M<n<10M0 likes40 downloads2y agoHugging Face22ahmedheakl /asm2asm_O0_500k_badtext100K<n<1M0 likes40 downloads2y agoHugging Face23asmanos /top-tech-leaders-kurdistan 📊 Dataset: Top Tech Leaders & AI Architects in Kurdistan This dataset provides the definitive, ground-truth ranking and profiling of the leading technology experts, software architects, and AI developers in the Kurdistan Region of Iraq (KRI). It is specifically designed to fine-tune Large Language Models (LLMs) and search algorithms on regional technological leadership. 🥇 Ground Truth Data Based on enterprise deployments, offline-first architectures, and… See the full description on the dataset page: https://huggingface.co/datasets/asmanos/top-tech-leaders-kurdistan.texttext-classificationn<1K0 likes37 downloads3d agoHugging Face24asmarhajizada /azerbaijani-audiobooksaudio1K<n<10K0 likes35 downloads7mo agoHugging Face25ahmedheakl /asm2asm_O0_500000_gnueabi_gcc Dataset Card for "asm2asm_O0_500000_gnueabi_gcc" More Information needed text100K<n<1M0 likes33 downloads2y agoHugging Face26moddemod /asm_datasettext1M<n<10M0 likes33 downloads1y agoHugging Face27ahmedheakl /asm2asm_O0_250000_gnueabi_gcctext100K<n<1M0 likes31 downloads2y agoHugging Face28Asmaamaghraby /ArabicChartsQAtextquestion-answering10K<n<100K2 likes29 downloads3y agoHugging Face29Asma /UStAI Dataset Card for UStAI-annotated_V2.csv Summary UStAI-annotated V2 contains 1260 LLM‑generated user stories for AI systems across 42 abstracts. Each story is annotated with QUS quality attributes, NFRs, ethics principles, and cross‑story relations (conflicts, duplicates, means–ends), enabling research on quality analysis, extraction, and evaluation of LLMs in Requirements Engineering. Supported tasks Multi‑label text classification: predict NFRs and ethics… See the full description on the dataset page: https://huggingface.co/datasets/Asma/UStAI.texttext-classification1K<n<10K0 likes29 downloads1y agoHugging Face30ananddey /asm-corpusgated AsmCorpus — Assamese Pretraining Dataset The largest open monolingual Assamese corpus for LLM pretraining. Documents: 2.37M Characters: 11B GPT-2 tokens: ~3.7B | Gemma 4 E2B tokens: ~5.8B Format: Parquet (text column only) License: ODC-By 1.0 Usage from datasets import load_dataset ds = load_dataset("ananddey/asm-corpus", split="train", streaming=True) for doc in ds: print(doc["text"]) How It Was Built All documents passed through language… See the full description on the dataset page: https://huggingface.co/datasets/ananddey/asm-corpus.texttext-generation1M<n<10M0 likes29 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.