catala
Datasets
All datasets matching “catala”wikimedia-common-audio-catalanThis is a collection of Catalan-language audio with free licenses extracted from Wikimedia Commons.
License identifiers are normalized to cc-zero, cc-by-4.0,
cc-by-sa-3.0, cc-by-sa-4.0, GFDL, and PD-self.
This provides a richer alternative to Common Voice.
Characteristics of the dataset:
One or multiple speakers
Different accents
Different domain texts
761 audio files
We found this dataset useful for audio tasks such as:
Language detection
Evaluation of STT systems
New candidates are… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/wikimedia-common-audio-catalan.VST-Training-Data
VST Training Data
Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
📄 Paper | 🌐 Project Page | 💻 Code | 🤗 Models
This dataset contains the full training data used for Video Streaming Thinking (VST), including both supervised fine-tuning (SFT) and reinforcement learning (RL) stages.
Dataset Structure
Subset
Description
vst_sft_data
SFT data including video-text pairs from multiple sources
vst_rl_data
RL data for reinforcement… See the full description on the dataset page: https://huggingface.co/datasets/Catalan258/VST-Training-Data.catalan_commonvoice
Dataset Card for "catalan_commonvoice"
More Information needed
aya-global-exams-catalanCatalan exams for the Aya Global Exams.
Original data and file available here: link
Github Repo: link
commonvoice_benchmark_catalan_accentsA new presentation of the corpus Catalan Common Voice v17.0 - metadata annotated version with the splits redefined to benchmark ASR models with various Catalan accentcatalanqa
Dataset Card for CatalanQA
Dataset Summary
This dataset can be used to build extractive-QA and Language Models. It is an aggregation and balancing of 2 previous datasets: VilaQuAD and ViquiQuAD.
Splits have been balanced by kind of question, and unlike other datasets like SQuAD, it only contains, per record, one question and one answer for each context, although the contexts can repeat multiple times.
This dataset was developed by BSC TeMU as part of Projecte AINA, to… See the full description on the dataset page: https://huggingface.co/datasets/projecte-aina/catalanqa.
