datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
majestrino-1.00-16xk5-sae-features
Majestrino 1.00 SAE — Feature Audio Samples (16x, k=5)
Top-2000 activating audio samples for each feature in the
Majestrino 1.00 SAE.
Overview
Metric
Value
SAE Architecture
16x expansion, k=5, d_model=768
Total Features
12,288
Alive Features
10,684
Audio per Feature
Up to 2,000 highest-activating
Audio Format
Opus (24 kbps OGG container)
Total TAR Files
1069
Source Dataset
laion/majestrino-data
File Structure
Each TAR file… See the full description on the dataset page: https://huggingface.co/datasets/laion/majestrino-1.00-16xk5-sae-features.twi-word-level-features
Twi word-level feature store (word-labeled frames)
sources: asr, female
vocabulary: 110,489 words (all types, no frequency floor)
occurrences: 5,223,750 (53,957 singletons)
top words: no, na, a, sɛ, wo, ne, yɛ, ɛyɛ, nso, wɔ, mu, deɛ, me, nti, hɔ, bi, saa, so, ho, ɛ
layout: words/mel/*.npy whole-utterance log-mels, words/manifest/*.npz mel-frame to word map, words/vocab.json inventory (class 0 = <sil>).
encodec_24khz-librispeech_asr-train.clean.100-features
Dataset Card for "encodec_24khz-librispeech_asr-train.clean.100-features"
More Information needed
encodec_24khz-opt-125m-pretrained-ft-librispeech_asr-train.clean.100-features
Dataset Card for "encodec_24khz-opt-125m-pretrained-ft-librispeech_asr-train.clean.100-features"
More Information needed
side-transformer-librispeech-460-features
LibriSpeech train-clean-460 ASR, Mel, and retrieval-event features
Public derived-feature release for the canonical LibriSpeech train-clean-100 and
train-clean-360 splits. It contains Whisper ASR word/timestamp JSON, schema-v2
retrieval-event metadata, and packed 80-bin log-Mel tensors. It contains no clinical
material, model weights, private paths, or source audio.
Layout
data/train-clean-100/ and data/train-clean-360/: WebDataset-style tar shards.… See the full description on the dataset page: https://huggingface.co/datasets/jay-junjiewu/side-transformer-librispeech-460-features.encodec_24khz-opt-125m-pretrained-ft-librispeech_asr_dummy-validation-features
Dataset Card for "encodec_24khz-opt-125m-pretrained-ft-librispeech_asr_dummy-validation-features"
More Information needed
jailbreak_with_features_10kencodec_24khz-b24.0-librispeech_asr-features
Dataset Card for "encodec_24khz-b24.0-librispeech_asr-features"
More Information needed
encodec_24khz-opt-125m-pretrained-ft-librispeech_asr-validation.clean-features
Dataset Card for "encodec_24khz-opt-125m-pretrained-ft-librispeech_asr-validation.clean-features"
More Information needed
encodec_24khz-librispeech_asr-validation.clean-features
Dataset Card for "encodec_24khz-librispeech_asr-validation.clean-features"
More Information needed
encodec_24khz-opt-125m-pretrained-ft-librispeech_asr-test.clean-features
Dataset Card for "encodec_24khz-opt-125m-pretrained-ft-librispeech_asr-test.clean-features"
More Information needed
Tess_features_extractedravdess_whisper_hidden_featuresencodec_24khz-librispeech_asr-test.clean-features
Dataset Card for "encodec_24khz-librispeech_asr-test.clean-features"
More Information needed
Crema_features_extractedmaster_emotion_audio_features_extractedmicrovent-features
microvent-features
Derived signals for the microvent core release: per-keyframe OCR text,
per-chunk ASR transcripts, and an embedding zoo (keyframe-level vision,
keyframe-OCR text, audio-level, video-level, omni-modal).
This card covers only the features. For the source videos, audio,
keyframes, and the public eval annotations, see the microvent dataset
card. All artifacts here key on the same chunk_id and follow the same
WebDataset shard layout, so joining feature shards back… See the full description on the dataset page: https://huggingface.co/datasets/hltcoe/microvent-features.Ravdess_features_extractedSavee_features_extracted
