datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Spoken-Tutorial
BhasaAnuvaad: A Speech Translation Dataset for 13 Indian Languages
Overview
BhasaAnuvaad, is the largest Indic-language AST dataset spanning over 44,400 hours of speech and 17M text segments for 13 of 22 scheduled Indian languages and English.
This repository consists of parallel data for Speech Translation from Spoken-Tutorial youtube channel, a subset of BhasaAnuvaad.
How to use
The datasets library allows you to load and pre-process your… See the full description on the dataset page: https://huggingface.co/datasets/ai4bharat/Spoken-Tutorial.circor-digiscope-physionet22-tutorialtutorial
Dataset Card for TEST HUGGINGFACE DATA SET
This is a test dataset built from AudioFile and Hub API
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
license: mit
dataset_info:
features:
- name: audio
dtype: audio
- name: label
dtype:
class_label:
names:
'0': beeps
'1': solid
splits:
- name: train
num_bytes: 601869
num_examples: 70
-… See the full description on the dataset page: https://huggingface.co/datasets/igsxf/tutorial.Spoken-Tutorial-Hindi-Filtered
