dasheng
Datasets
All datasets matching “dasheng”audioset-dasheng-0.6b-emb
AudioSet DaSheng-0.6B embeddings
Mean-pooled, float16 embeddings of
danjacobellis/audioset_opus_24kbps
from mispeech/dasheng-0.6B.
Columns
path: source clip path (string)
label: source AudioSet label indices (list of int64)
emb: 1,280-dimensional fixed-size list of float16
Audio is decoded from the source Opus bytes, mixed to mono, and resampled to
16 kHz. The embedding is the model's documented outputdim=None output:
sigmoid applied to the mean of the final… See the full description on the dataset page: https://huggingface.co/datasets/quinnlue/audioset-dasheng-0.6b-emb.generative-sound-masking-retrieval-dasheng-v1
Generative Sound Masking — DaSheng retrieval v1
This dataset stores exhaustive top-15 retrieval from the 237,500 prompt–seed
candidate audios for 48,840 background audios. One row per background contains
15 ranked candidates, including prompt, seed, clip ID, audio SHA256 and cosine score.
Source audio stays in the two original public datasets; no audio is duplicated here.
This is an incremental run. Read progress.json before treating it as complete.
The data/train split is an… See the full description on the dataset page: https://huggingface.co/datasets/AE-W/generative-sound-masking-retrieval-dasheng-v1.dasheng-base-audioset-mae-embeddings
