datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Iris_Database
Synthetic Iris Image Dataset
Overview
This repository contains a dataset of synthetic colored iris images generated using diffusion models based on our paper "Synthetic Iris Image Generation Using Diffusion Networks." The dataset comprises 17,695 high-quality synthetic iris images designed to be biometrically unique from the training data while maintaining realistic iris pigmentation distributions. In this repository we contain about 10000 filtered iris images with the… See the full description on the dataset page: https://huggingface.co/datasets/fatdove/Iris_Database.CTIS
Dataset Card for Chinese Traditional Instrument Sound
Original Content
The original dataset is created by [1], with no evaluation provided. The original CTIS dataset contains recordings from 287 varieties of Chinese traditional instruments, reformed Chinese musical instruments, and instruments from ethnic minority groups. Notably, some of these instruments are rarely encountered by the majority of the Chinese populace. The dataset was later utilized by [2] for Chinese… See the full description on the dataset page: https://huggingface.co/datasets/ccmusic-database/CTIS.pianos
Dataset Card for Piano Sound Quality Dataset
The original dataset is sourced from the Piano Sound Quality Dataset, which includes 12 full-range audio files in .wav/.mp3/.m4a format representing seven models of pianos: Kawai upright piano, Kawai grand piano, Young Change upright piano, Hsinghai upright piano, Grand Theatre Steinway piano, Steinway grand piano, and Pearl River upright piano. Additionally, there are 1,320 split monophonic audio files in .wav/.mp3/.m4a format, bringing… See the full description on the dataset page: https://huggingface.co/datasets/ccmusic-database/pianos.Guzheng_Tech99
Dataset Card for Guzheng Technique 99 Dataset
Original Content
This dataset is created and used by [1] for frame-level Guzheng playing technique detection. The original dataset encompasses 99 solo compositions for Guzheng, recorded by professional musicians within a studio environment. Each composition is annotated for every note, indicating the onset, offset, pitch, and playing techniques. This is different from the GZ IsoTech, which is annotated at the clip-level. Also… See the full description on the dataset page: https://huggingface.co/datasets/ccmusic-database/Guzheng_Tech99.acapella
Dataset Card for Acapella Evaluation
The original dataset, sourced from the Acapella Evaluation Dataset, comprises six Mandarin pop song segments performed by 22 singers, resulting in a total of 132 audio clips. Each segment includes both a verse and a chorus. Four judges from the China Conservatory of Music assess the singing across nine dimensions: pitch, rhythm, vocal range, timbre, pronunciation, vibrato, dynamics, breath control, and overall performance, using a 10-point scale.… See the full description on the dataset page: https://huggingface.co/datasets/ccmusic-database/acapella.music_genre
Dataset Card for Music Genre
The Default dataset comprises approximately 1,700 musical pieces in .mp3 format, sourced from the NetEase music. The lengths of these pieces range from 270 to 300 seconds. All are sampled at the rate of 22,050 Hz. As the website providing the audio music includes style labels for the downloaded music, there are no specific annotators involved. Validation is achieved concurrently with the downloading process. They are categorized into a total of 16… See the full description on the dataset page: https://huggingface.co/datasets/ccmusic-database/music_genre.bel_canto
Dataset Card for Bel Conto and Chinese Folk Song Singing Tech
Original Content
This dataset is created by the authors and encompasses two distinct singing styles: bel canto and Chinese folk singing. Bel canto is a vocal technique frequently employed in Western classical music and opera, symbolizing the zenith of vocal artistry within the broader Western musical heritage. Chinese folk singing, for which there is no official English translation, is referred to here as a… See the full description on the dataset page: https://huggingface.co/datasets/ccmusic-database/bel_canto.fragrance-database
FragDB v5.16 — Fragrance Database (Multilingual Sample)
The most comprehensive structured fragrance database available. This is a free sample of FragDB: 140,230 perfumes, 23 languages — 10-row CSV samples at root.
Full dataset: fragdb.net.
What's New in v5.16
Data updated from v5.15 → v5.16 (snapshot 2026-09-19):
Fragrances: 139,501 → 140,230 (+729)
Brands: 8,272 → 8,316 (+44)
Perfumers: 3,116 → 3,126 (+10)
Notes: 2,596 → 2,606 rows in notes.csv (+10)
Companion… See the full description on the dataset page: https://huggingface.co/datasets/FragDBnet/fragrance-database.GZ_IsoTech
Dataset Card for GZ_IsoTech Dataset
Original Content
The dataset is created and used for Guzheng playing technique detection by [1]. The original dataset comprises 2,824 variable-length audio clips showcasing various Guzheng playing techniques. Specifically, 2,328 clips were sourced from virtual sound banks, while 496 clips were performed by a professional Guzheng artist.
The clips are annotated in eight categories, with a Chinese pinyin and Chinese characters written in… See the full description on the dataset page: https://huggingface.co/datasets/ccmusic-database/GZ_IsoTech.erhu_playing_tech
Dataset Card for Erhu Playing Technique
Original Content
This dataset was created and has been utilized for Erhu playing technique detection by [1], which has not undergone peer review. The original dataset comprises 1,253 Erhu audio clips, all performed by professional Erhu players. These clips were annotated according to three levels, resulting in annotations for four, seven, and 11 categories. Part of the audio data is sourced from the CTIS dataset described earlier.… See the full description on the dataset page: https://huggingface.co/datasets/ccmusic-database/erhu_playing_tech.timbre_range
Dataset Card for Timbre and Range Dataset
Dataset Summary
The timbre dataset contains acapella singing audio of 9 singers, as well as cut single-note audio, totaling 775 clips (.wav format)
The vocal range dataset includes several up and down chromatic scales audio clips of several vocals, as well as the cut single-note audio clips (.wav format).
Supported Tasks and Leaderboards
Audio classification
Languages
Chinese, English
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ccmusic-database/timbre_range.chest_falsetto
Dataset Card for Chest voice and Falsetto Dataset
The original dataset, sourced from the Chest Voice and Falsetto Dataset, includes 1,280 monophonic singing audio files in .wav format, performed, recorded, and annotated by students majoring in Vocal Music at the China Conservatory of Music. The chest voice is tagged as "chest" and the falsetto voice as "falsetto." Additionally, the dataset encompasses the Mel spectrogram, Mel frequency cepstral coefficient (MFCC), and spectral… See the full description on the dataset page: https://huggingface.co/datasets/ccmusic-database/chest_falsetto.Affine_Transformation_Databasesatvian-databaseinstrument_timbre
Dataset Card for Chinese Musical Instruments Timbre Evaluation Database
The original dataset is sourced from the National Musical Instruments Timbre Evaluation Dataset, which includes subjective timbre evaluation scores using 16 terms such as bright, dark, raspy, etc., evaluated across 37 Chinese instruments and 24 Western instruments by Chinese participants with musical backgrounds in a subjective evaluation experiment. Additionally, it contains 10 spectrogram analysis reports for… See the full description on the dataset page: https://huggingface.co/datasets/ccmusic-database/instrument_timbre.gif-database
GIF Anatomical Atlas Database
Status: PRIVATE. This repository will be made public once data-sharing /
ethics governance for redistribution is separately confirmed.
Contents
Reference atlas database (T1, FLAIR, and GIF anatomical segmentation labels)
used by the leukoquant GIF pipeline, de-identified with MiDeFace.
db_mideface.tar.gz -- T1 images (102 subjects) + GIF anatomical labels +
GIF database manifest (db.xml, labels.xml, GroupMask.nii.gz).… See the full description on the dataset page: https://huggingface.co/datasets/stylianosc/gif-database.Traffic_Sign_Recogntion_DatabaseTSRD (Traffic Sign Recognition Database) 是一个中国交通标志数据集,包含多种交通标志类别。数据集分为训练集和测试集:
训练集:包含约4170张图像
测试集:包含约1994张图像
类别数:约58个不同的交通标志类别
数据集格式为:
图像文件名;宽;高;x1;y1;x2;y2;类别;
包含全种类数据集 / 4方向指示牌数据集
danbooru2023-metadata-database
Metadata Database for Danbooru2023
Danbooru 2023 datasets: https://huggingface.co/datasets/nyanko7/danbooru2023
The latest entry of this database is id 7,866,491. Which is newer than nyanko7's dataset.
This dataset contains a sqlite db file which have all the tags and posts metadata in it.
The Peewee ORM config file is provided too, plz check it for more information. (Especially on how I link posts and tags together)
The original data is from the official dump of the posts info.… See the full description on the dataset page: https://huggingface.co/datasets/KBlueLeaf/danbooru2023-metadata-database.ArASL_Database_Grayscale
Dataset Card for "ArASL_Database_Grayscale"
Dataset Summary
A new dataset consists of 54,049 images of ArSL alphabets performed by more than 40 people for 32 standard Arabic signs and alphabets.
The number of images per class differs from one class to another. Sample image of all Arabic Language Signs is also attached. The CSV file contains the Label of each corresponding Arabic Sign Language Image based on the image file name.
Supported Tasks and Leaderboards… See the full description on the dataset page: https://huggingface.co/datasets/pain/ArASL_Database_Grayscale.mixed_databaseCNPM
Dataset Card for Chinese National Pentatonic Mode Dataset
Original Content
The dataset is initially created by [1]. It is then expanded and used for automatic Chinese national pentatonic mode recognition by [2], to which readers can refer for more details along with a brief introduction to the modern theory of Chinese pentatonic mode. This includes the definition of "system", "tonic", "pattern", and "type," which will be included in one unified table during our… See the full description on the dataset page: https://huggingface.co/datasets/ccmusic-database/CNPM.Phone_Timings_Database
📖 TajweedAI: Quranic Phoneme Timing Benchmark (Phases 1, 2 & 3)
📌 Project Overview
TajweedAI evaluates Quranic recitation accuracy by analyzing both pronunciation (phoneme classification) and timing (rule duration evaluation).
This benchmark provides empirical, tempo-normalized duration boundaries for all 70 Quranic phonemes derived from forced alignments (MFA trained on Quranic audio) across 7 master reference reciters:
Sheikh Mahmoud Khalil Al-Husary (Gold… See the full description on the dataset page: https://huggingface.co/datasets/AhmedTamertechno1/Phone_Timings_Database.Iris_Database
Synthetic Iris Image Dataset
Overview
This repository contains a dataset of synthetic colored iris images generated using diffusion models based on our paper "Synthetic Iris Image Generation Using Diffusion Networks." The dataset comprises 17,695 high-quality synthetic iris images designed to be biometrically unique from the training data while maintaining realistic iris pigmentation distributions. In this repository we contain about 10000 filtered iris images with the… See the full description on the dataset page: https://huggingface.co/datasets/AditiGupta2004/Iris_Database.European_Soccer_Databasesong_structure
Dataset Card for Song Structure
The raw dataset comprises 300 pop songs in .mp3 format, sourced from the NetEase music, accompanied by a structure annotation file for each song in .txt format. The annotator for music structure is a professional musician and teacher from the China Conservatory of Music. For the statistics of the dataset, there are 208 Chinese songs, 87 English songs, three Korean songs and two Japanese songs. The song structures are labeled as follows: intro… See the full description on the dataset page: https://huggingface.co/datasets/ccmusic-database/song_structure.ai-tools-database-25k-aitoolbuzzfigure2data-database-v1abhishekgupta56447_anime-offline-database
Anime Offline Database
"35,000+ anime titles with cross-referenced IDs
Dataset Info
Source: Kaggle
Original Size: 9.39 MB
Kaggle Downloads: 45
Files: 1
Files
anime_database.csv
Mirrored from Kaggle
dog-image-databasev3la_eval_database
