datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
sada-train-wav2vec2-xls-r-300m-ar-preprocessedpashto-audio-wav2vecsimplified_google_speech_commands_wav2vec2_960hwav2vec2-test-datasetwav2vec2_basespot_data_allgoogle-speech-commands-wav2vec2-960hWav2Vec_MELD_Audiosimplified-google-speech-commands-wav2vec2-960hsada-validation-wav2vec2-xls-r-300m-ar-preprocessedWav2Vec2_ASVSpoof5_FULLwav2vec2-vd-bird-sound-classification-datasetwav2vec2_phoneme_spot_data_allenenlhet-wav2vec2-dataset
Enenlhet Wav2Vec2 Dataset
This dataset contains preprocessed audio features and tokenized text for training Wav2Vec2 models on the Enenlhet language.
Dataset Summary
Train: 3,053 examples
Test: 170 examples
Validation: 170 examples
Total: 3,393 examples
Features
input_values: Preprocessed audio features (16kHz, normalized float32 arrays)
labels: Tokenized text as integer sequences
Usage
from datasets import load_dataset
# Load the dataset… See the full description on the dataset page: https://huggingface.co/datasets/sjhuskey/enenlhet-wav2vec2-dataset.sada-test-wav2vec2-xls-r-300m-ar-preprocessedwav2vec2-jailbreak-classificationViMD_north_wav2vec2wav2vec_filter_wo_processing_spotifywav2vec2-urdu-finetuned-ASR-datasetwav2vec2_conformer_spot_data_allVietMed_labeled_wav2vec2vivos_wav2vec2ViMD_central_wav2vec2MinskGemini_wav2vec2ViMD_south_wav2vec2wpp_pav_transcrito_jonatasgrosman-wav2vec2-large-xlsr-53-portugueseMinskGemini_wav2vec2_v2wpp_pav_transcrito_wav2vec2-portuguese-wpp-checkpoint-480wav2vec2-20pct-20250624-122057ORIGINAL_wav2vec2-20pct-20250624-173821-originalwav2vec2-20pct-20250624-122057-checkpoint-18000
