CoolFace
19 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01laion /majestrino-1.00-16xk5-sae-features Majestrino 1.00 SAE — Feature Audio Samples (16x, k=5) Top-2000 activating audio samples for each feature in the Majestrino 1.00 SAE. Overview Metric Value SAE Architecture 16x expansion, k=5, d_model=768 Total Features 12,288 Alive Features 10,684 Audio per Feature Up to 2,000 highest-activating Audio Format Opus (24 kbps OGG container) Total TAR Files 1069 Source Dataset laion/majestrino-data File Structure Each TAR file… See the full description on the dataset page: https://huggingface.co/datasets/laion/majestrino-1.00-16xk5-sae-features.audioaudio-classification10M<n<100M0 likes271 downloads6mo agoHugging Face02ghanaopenai /twi-word-level-features Twi word-level feature store (word-labeled frames) sources: asr, female vocabulary: 110,489 words (all types, no frequency floor) occurrences: 5,223,750 (53,957 singletons) top words: no, na, a, sɛ, wo, ne, yɛ, ɛyɛ, nso, wɔ, mu, deɛ, me, nti, hɔ, bi, saa, so, ho, ɛ layout: words/mel/*.npy whole-utterance log-mels, words/manifest/*.npz mel-frame to word map, words/vocab.json inventory (class 0 = <sil>). audion<1K0 likes182 downloads9d agoHugging Face03cmu-mlsp /encodec_24khz-librispeech_asr-train.clean.100-features Dataset Card for "encodec_24khz-librispeech_asr-train.clean.100-features" More Information needed audio10K<n<100K0 likes173 downloads3y agoHugging Face04cmu-mlsp /encodec_24khz-opt-125m-pretrained-ft-librispeech_asr-train.clean.100-features Dataset Card for "encodec_24khz-opt-125m-pretrained-ft-librispeech_asr-train.clean.100-features" More Information needed audio10K<n<100K0 likes143 downloads3y agoHugging Face05jay-junjiewu /side-transformer-librispeech-460-features LibriSpeech train-clean-460 ASR, Mel, and retrieval-event features Public derived-feature release for the canonical LibriSpeech train-clean-100 and train-clean-360 splits. It contains Whisper ASR word/timestamp JSON, schema-v2 retrieval-event metadata, and packed 80-bin log-Mel tensors. It contains no clinical material, model weights, private paths, or source audio. Layout data/train-clean-100/ and data/train-clean-360/: WebDataset-style tar shards.… See the full description on the dataset page: https://huggingface.co/datasets/jay-junjiewu/side-transformer-librispeech-460-features.audio0 likes101 downloads11d agoHugging Face06cmu-mlsp /encodec_24khz-opt-125m-pretrained-ft-librispeech_asr_dummy-validation-features Dataset Card for "encodec_24khz-opt-125m-pretrained-ft-librispeech_asr_dummy-validation-features" More Information needed audion<1K1 likes65 downloads3y agoHugging Face07assoni2002 /jailbreak_with_features_10kaudio10K<n<100K0 likes38 downloads1y agoHugging Face08cmu-mlsp /encodec_24khz-b24.0-librispeech_asr-features Dataset Card for "encodec_24khz-b24.0-librispeech_asr-features" More Information needed audio1K<n<10K0 likes27 downloads3y agoHugging Face09cmu-mlsp /encodec_24khz-opt-125m-pretrained-ft-librispeech_asr-validation.clean-features Dataset Card for "encodec_24khz-opt-125m-pretrained-ft-librispeech_asr-validation.clean-features" More Information needed audio1K<n<10K0 likes21 downloads3y agoHugging Face10cmu-mlsp /encodec_24khz-librispeech_asr-validation.clean-features Dataset Card for "encodec_24khz-librispeech_asr-validation.clean-features" More Information needed audio1K<n<10K0 likes16 downloads3y agoHugging Face11cmu-mlsp /encodec_24khz-opt-125m-pretrained-ft-librispeech_asr-test.clean-features Dataset Card for "encodec_24khz-opt-125m-pretrained-ft-librispeech_asr-test.clean-features" More Information needed audio1K<n<10K0 likes14 downloads3y agoHugging Face12c0derish /Tess_features_extractedaudio1K<n<10K0 likes14 downloads2y agoHugging Face13ajdsouza /ravdess_whisper_hidden_featuresaudio1K<n<10K0 likes14 downloads2y agoHugging Face14cmu-mlsp /encodec_24khz-librispeech_asr-test.clean-features Dataset Card for "encodec_24khz-librispeech_asr-test.clean-features" More Information needed audio1K<n<10K0 likes13 downloads3y agoHugging Face15c0derish /Crema_features_extractedaudio1K<n<10K0 likes13 downloads2y agoHugging Face16c0derish /master_emotion_audio_features_extractedaudio10K<n<100K0 likes9 downloads2y agoHugging Face17hltcoe /microvent-featuresgated microvent-features Derived signals for the microvent core release: per-keyframe OCR text, per-chunk ASR transcripts, and an embedding zoo (keyframe-level vision, keyframe-OCR text, audio-level, video-level, omni-modal). This card covers only the features. For the source videos, audio, keyframes, and the public eval annotations, see the microvent dataset card. All artifacts here key on the same chunk_id and follow the same WebDataset shard layout, so joining feature shards back… See the full description on the dataset page: https://huggingface.co/datasets/hltcoe/microvent-features.textvideo-classification10K<n<100K0 likes7 downloads4mo agoHugging Face18c0derish /Ravdess_features_extractedaudio1K<n<10K0 likes4 downloads2y agoHugging Face19c0derish /Savee_features_extractedaudion<1K0 likes4 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.