datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
majestrino-1.00-16xk5-sae-features
Majestrino 1.00 SAE — Feature Audio Samples (16x, k=5)
Top-2000 activating audio samples for each feature in the
Majestrino 1.00 SAE.
Overview
Metric
Value
SAE Architecture
16x expansion, k=5, d_model=768
Total Features
12,288
Alive Features
10,684
Audio per Feature
Up to 2,000 highest-activating
Audio Format
Opus (24 kbps OGG container)
Total TAR Files
1069
Source Dataset
laion/majestrino-data
File Structure
Each TAR file… See the full description on the dataset page: https://huggingface.co/datasets/laion/majestrino-1.00-16xk5-sae-features.microvent-features
microvent-features
Derived signals for the microvent core release: per-keyframe OCR text,
per-chunk ASR transcripts, and an embedding zoo (keyframe-level vision,
keyframe-OCR text, audio-level, video-level, omni-modal).
This card covers only the features. For the source videos, audio,
keyframes, and the public eval annotations, see the microvent dataset
card. All artifacts here key on the same chunk_id and follow the same
WebDataset shard layout, so joining feature shards back… See the full description on the dataset page: https://huggingface.co/datasets/hltcoe/microvent-features.
