datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CTIS
Dataset Card for Chinese Traditional Instrument Sound
Original Content
The original dataset is created by [1], with no evaluation provided. The original CTIS dataset contains recordings from 287 varieties of Chinese traditional instruments, reformed Chinese musical instruments, and instruments from ethnic minority groups. Notably, some of these instruments are rarely encountered by the majority of the Chinese populace. The dataset was later utilized by [2] for Chinese… See the full description on the dataset page: https://huggingface.co/datasets/ccmusic-database/CTIS.pianos
Dataset Card for Piano Sound Quality Dataset
The original dataset is sourced from the Piano Sound Quality Dataset, which includes 12 full-range audio files in .wav/.mp3/.m4a format representing seven models of pianos: Kawai upright piano, Kawai grand piano, Young Change upright piano, Hsinghai upright piano, Grand Theatre Steinway piano, Steinway grand piano, and Pearl River upright piano. Additionally, there are 1,320 split monophonic audio files in .wav/.mp3/.m4a format, bringing… See the full description on the dataset page: https://huggingface.co/datasets/ccmusic-database/pianos.TORGO-database
The TORGO Database: Acoustic and articulatory speech from speakers with dysarthria
Dataset Summary
This database only includes the short words and restricted sentence portion of the TORGO dataset.
For the full dataset which also includes non-words and unrestricted sentences please see: https://www.cs.toronto.edu/~complingweb/data/TORGO/torgo.html.
Transcripts have been normalized to remove punctuation but casing has been left. Few transcripts only had 'xxx' as text… See the full description on the dataset page: https://huggingface.co/datasets/abnerh/TORGO-database.Guzheng_Tech99
Dataset Card for Guzheng Technique 99 Dataset
Original Content
This dataset is created and used by [1] for frame-level Guzheng playing technique detection. The original dataset encompasses 99 solo compositions for Guzheng, recorded by professional musicians within a studio environment. Each composition is annotated for every note, indicating the onset, offset, pitch, and playing techniques. This is different from the GZ IsoTech, which is annotated at the clip-level. Also… See the full description on the dataset page: https://huggingface.co/datasets/ccmusic-database/Guzheng_Tech99.acapella
Dataset Card for Acapella Evaluation
The original dataset, sourced from the Acapella Evaluation Dataset, comprises six Mandarin pop song segments performed by 22 singers, resulting in a total of 132 audio clips. Each segment includes both a verse and a chorus. Four judges from the China Conservatory of Music assess the singing across nine dimensions: pitch, rhythm, vocal range, timbre, pronunciation, vibrato, dynamics, breath control, and overall performance, using a 10-point scale.… See the full description on the dataset page: https://huggingface.co/datasets/ccmusic-database/acapella.bel_canto
Dataset Card for Bel Conto and Chinese Folk Song Singing Tech
Original Content
This dataset is created by the authors and encompasses two distinct singing styles: bel canto and Chinese folk singing. Bel canto is a vocal technique frequently employed in Western classical music and opera, symbolizing the zenith of vocal artistry within the broader Western musical heritage. Chinese folk singing, for which there is no official English translation, is referred to here as a… See the full description on the dataset page: https://huggingface.co/datasets/ccmusic-database/bel_canto.GZ_IsoTech
Dataset Card for GZ_IsoTech Dataset
Original Content
The dataset is created and used for Guzheng playing technique detection by [1]. The original dataset comprises 2,824 variable-length audio clips showcasing various Guzheng playing techniques. Specifically, 2,328 clips were sourced from virtual sound banks, while 496 clips were performed by a professional Guzheng artist.
The clips are annotated in eight categories, with a Chinese pinyin and Chinese characters written in… See the full description on the dataset page: https://huggingface.co/datasets/ccmusic-database/GZ_IsoTech.erhu_playing_tech
Dataset Card for Erhu Playing Technique
Original Content
This dataset was created and has been utilized for Erhu playing technique detection by [1], which has not undergone peer review. The original dataset comprises 1,253 Erhu audio clips, all performed by professional Erhu players. These clips were annotated according to three levels, resulting in annotations for four, seven, and 11 categories. Part of the audio data is sourced from the CTIS dataset described earlier.… See the full description on the dataset page: https://huggingface.co/datasets/ccmusic-database/erhu_playing_tech.timbre_range
Dataset Card for Timbre and Range Dataset
Dataset Summary
The timbre dataset contains acapella singing audio of 9 singers, as well as cut single-note audio, totaling 775 clips (.wav format)
The vocal range dataset includes several up and down chromatic scales audio clips of several vocals, as well as the cut single-note audio clips (.wav format).
Supported Tasks and Leaderboards
Audio classification
Languages
Chinese, English
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/ccmusic-database/timbre_range.chest_falsetto
Dataset Card for Chest voice and Falsetto Dataset
The original dataset, sourced from the Chest Voice and Falsetto Dataset, includes 1,280 monophonic singing audio files in .wav format, performed, recorded, and annotated by students majoring in Vocal Music at the China Conservatory of Music. The chest voice is tagged as "chest" and the falsetto voice as "falsetto." Additionally, the dataset encompasses the Mel spectrogram, Mel frequency cepstral coefficient (MFCC), and spectral… See the full description on the dataset page: https://huggingface.co/datasets/ccmusic-database/chest_falsetto.SIWIS_French_Speech_Synthesis_Database
SIWIS French Speech Synthesis Database
This README provides a concise description of the dataset, including its structure, file naming conventions, and known labeling issues. Additionally, suggestions for potential improvements are outlined in the TODO section.
The dataset is distributed under the Creative Commons Attribution 4.0 International (CC BY 4.0) license, permitting its use for any purpose.
For more details about the database design and recording process, please refer… See the full description on the dataset page: https://huggingface.co/datasets/Aviv-anthonnyolime/SIWIS_French_Speech_Synthesis_Database.instrument_timbre
Dataset Card for Chinese Musical Instruments Timbre Evaluation Database
The original dataset is sourced from the National Musical Instruments Timbre Evaluation Dataset, which includes subjective timbre evaluation scores using 16 terms such as bright, dark, raspy, etc., evaluated across 37 Chinese instruments and 24 Western instruments by Chinese participants with musical backgrounds in a subjective evaluation experiment. Additionally, it contains 10 spectrogram analysis reports for… See the full description on the dataset page: https://huggingface.co/datasets/ccmusic-database/instrument_timbre.saarbruecken-voice-database-16khzCNPM
Dataset Card for Chinese National Pentatonic Mode Dataset
Original Content
The dataset is initially created by [1]. It is then expanded and used for automatic Chinese national pentatonic mode recognition by [2], to which readers can refer for more details along with a brief introduction to the modern theory of Chinese pentatonic mode. This includes the definition of "system", "tonic", "pattern", and "type," which will be included in one unified table during our… See the full description on the dataset page: https://huggingface.co/datasets/ccmusic-database/CNPM.Phone_Timings_Database
📖 TajweedAI: Quranic Phoneme Timing Benchmark (Phases 1, 2 & 3)
📌 Project Overview
TajweedAI evaluates Quranic recitation accuracy by analyzing both pronunciation (phoneme classification) and timing (rule duration evaluation).
This benchmark provides empirical, tempo-normalized duration boundaries for all 70 Quranic phonemes derived from forced alignments (MFA trained on Quranic audio) across 7 master reference reciters:
Sheikh Mahmoud Khalil Al-Husary (Gold… See the full description on the dataset page: https://huggingface.co/datasets/AhmedTamertechno1/Phone_Timings_Database.song_structure
Dataset Card for Song Structure
The raw dataset comprises 300 pop songs in .mp3 format, sourced from the NetEase music, accompanied by a structure annotation file for each song in .txt format. The annotator for music structure is a professional musician and teacher from the China Conservatory of Music. For the statistics of the dataset, there are 208 Chinese songs, 87 English songs, three Korean songs and two Japanese songs. The song structures are labeled as follows: intro… See the full description on the dataset page: https://huggingface.co/datasets/ccmusic-database/song_structure.ptsl-databaseSSL_databaseTORGO-database
The TORGO Database: Acoustic and articulatory speech from speakers with dysarthria
Dataset Summary
This database only includes the short words and restricted sentence portion of the TORGO dataset.
For the full dataset which also includes non-words and unrestricted sentences please see: https://www.cs.toronto.edu/~complingweb/data/TORGO/torgo.html.
Transcripts have been normalized to remove punctuation but casing has been left. Few transcripts only had 'xxx' as text… See the full description on the dataset page: https://huggingface.co/datasets/kingp12/TORGO-database.EE200_Project_Song_DatabasePrivate-Sound-Database
