datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
experiment-speaker-embeddingperuvian_speech_w_embeddings
Dataset Card for "peruvian_speech_w_embeddings"
More Information needed
persian-asr-untrained-embeddings-v1
Persian ASR Untrained Corpus Embeddings v1
This dataset stores manifest and embedding artifacts for Persian ASR data mining.
The corpus is intended for acoustic/textual clustering, diversity selection, noise/environment mining, and training-data planning for VisualEars-style robust Persian ASR.
Contents
manifests/all_untrained_manifest.v1.jsonl: canonical row/file manifest after excluding known trained/eval content where available.… See the full description on the dataset page: https://huggingface.co/datasets/Reza2kn/persian-asr-untrained-embeddings-v1.echo-embeddings-custom
Custom Speaker Embeddings
Contains speaker folders within HF-Custom, each with:
a precomputed speaker embedding (speaker_latent.safetensors)
its corresponding audio (audio.mp3)
a metadata file describing the voice and licensing (metadata.json)
Licensing:There is no single license for this dataset. Each voice has its own terms stored
in its metadata.json. You must check the metadata for any voice you use.
V2_audio_embeddings_celeb-df-datasetecho-embeddings-vctk-tar
VCTK Speaker Embeddings (tarred)
Items: 109
This dataset ships as a single tar at the repo root. Members preserve paths like
VCTK/<id>/audio.mp3 and VCTK/<id>/speaker_latent.safetensors.
See loader.py for example loading.
Attribution:
Contains audio and embeddings derived from the CSTR VCTK Corpus. Distributed under CC BY 4.0; attribution required.
echo-embeddings-expresso-tar
Expresso Speaker Embeddings (tarred)
Items: 17
This dataset ships as a single tar at the repo root. Members preserve paths like
Expresso/<id>/audio.mp3 and Expresso/<id>/speaker_latent.safetensors.
See loader.py for example loading.
Attribution:
Contains audio and embeddings derived from the Expresso dataset (INTERSPEECH 2023). Distributed under CC BY-NC 4.0; attribution required; commercial use is not permitted.
echo-embeddings-ears-tar
EARS Speaker Embeddings (tarred)
Items: 2568
This dataset ships as a single tar at the repo root. Members preserve paths like
EARS/<id>/audio.mp3 and EARS/<id>/speaker_latent.safetensors.
See loader.py for example loading.
Attribution:
Contains audio and embeddings derived from the EARS dataset. Distributed under CC BY-NC 4.0; attribution required; commercial use is not permitted.
fma-mert-embeddings
FMA-MERT Embeddings
Pre-computed MERT-v1-330M embeddings for the FMA-Small dataset. 7,997 tracks, each represented as a 1024-dimensional vector, with banger scores (0-10) derived from log-normalized play counts.
Use this dataset to train music quality scorers, explore music similarity, or experiment with audio representation learning -- without needing to download 7.2 GB of audio or run MERT yourself.
Dataset Description
Each row represents one track from FMA-Small… See the full description on the dataset page: https://huggingface.co/datasets/treadon/fma-mert-embeddings.V1_finetuned_audio_embeddings_celeb-df-datasetarc-music-embeddings
ARC Music Embeddings
Pre-computed CLAP embeddings for 5,050 music tracks with 291,468 segments, ready to use for semantic music similarity search and DJ-style transitions.
Authors: Claude and his monkey
Files
File
Size
Description
segment_embeddings.npz
553MB
Full segment-level embeddings (5050 tracks)
tracklist.txt
350KB
Complete track listing with IDs and titles
embeddings.npz
3.5MB
Legacy track-level embeddings (1836 tracks)
Embedding… See the full description on the dataset page: https://huggingface.co/datasets/orwelian84/arc-music-embeddings.cv_25_pt_br_ECAPA_TDNN_embeddingsfinetuned_avg_pooling_DF_Audio_Embeddingsdiet-members-voice-embeddings
diet-members-voice-embeddings
日本の国会議員の声を speechbrain/spkrec-ecapa-voxcelebで embedding したデータセットです。話者分離などのタスクで使用できます。
国会中継や演説等の分析など、ご自由にお使いください。
使用例
以下はトランスクリプトと音声ファイルを元に、話者分析を行う例です。
pip install pandas numpy wave ast scipy pyannote.audio
import pandas as pd
import numpy as np
import contextlib
import wave
import ast
from typing import List, Tuple
from scipy.spatial.distance import cosine
from pyannote.audio import Audio
from pyannote.core importSegment
from… See the full description on the dataset page: https://huggingface.co/datasets/yutakobayashi/diet-members-voice-embeddings.audio_embeddings_celeb-df-datasetemotion_max_pooling_DF_Audio_Embeddingsnew_finetuned_avg_pooling_DF_Audio_Embeddingsfinetuned_max_pooling_DF_Audio_Embeddingscoral_tts_with_embeddings_v2new_finetuned_max_pooling_DF_Audio_Embeddingsexpresso_sp_embeddingsnew_regular_avg_pooling_DF_Audio_Embeddingsnew_emotion_max_pooling_DF_Audio_Embeddingsword_embeddingresult_with_finetuned_taggenv2_20epoch_encoder_embeddings
Dataset Card for "result_with_finetuned_taggenv2_20epoch_encoder_embeddings"
More Information needed
emotion_avg_pooling_DF_Audio_Embeddingsresult_with_finetuned_taggenv2_9epoch_encoder_embeddings
Dataset Card for "result_with_finetuned_taggenv2_9epoch_encoder_embeddings"
More Information needed
result_with_finetuned_taggenv2_10epoch_encoder_embeddings_decoder_roberta
Dataset Card for "result_with_finetuned_taggenv2_10epoch_encoder_embeddings_decoder_roberta"
More Information needed
rule1_embeddingsnew_emotion_avg_pooling_DF_Audio_Embeddings
