datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
frisian-asr-cv22
Frisian ASR (Common Voice 22, filtered)
Open Standard West Frisian (fy-NL) speech for ASR, built from Mozilla Common Voice 22.0
(CC0). The validated training split is augmented with the unvalidated other bucket, which is
auto-filtered by CTC agreement with the known prompt using a Frisian-specialized wav2vec2 model.
Built to fine-tune streaming ASR models (e.g. nvidia/nemotron-3.5-asr-streaming-0.6b).
Splits
Split
Clips
Hours
Composition
train
29,929… See the full description on the dataset page: https://huggingface.co/datasets/LokaalHub/frisian-asr-cv22.Frisian-dialogue
Frisian dialogue (Project FRIS)
This dataset consists of transcribed spontaneous Frisian dialogue. The conversations are mocked consults between doctor and patient and their topics cover various medical concerns and ailments that are generally discussed in a GP practice. The dataset has been collected in a spontaneous conversational setting, resulting in realistic audio quality as encountered in real-world usage of speech-to-text models for medical scribes.
This repository contains… See the full description on the dataset page: https://huggingface.co/datasets/Juvoly/Frisian-dialogue.alpaca_frisian_tacoThis repository contains the dataset used for the TaCo paper.
The dataset follows the style outlined in the TaCo paper, as follows:
{
"instruction": "instruction in xx",
"input": "input in xx",
"output": "Instruction in English: instruction in en ,
Response in English: response in en ,
Response in xx: response in xx "
}
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca_frisian_taco.frisianalpaca-frisian-cleanedThis repository contains the dataset used for the TaCo paper.
Please refer to the paper for more details: OpenReview
If you have used our dataset, please cite it as follows:
Citation
@inproceedings{upadhayay2024taco,
title={TaCo: Enhancing Cross-Lingual Transfer for Low-Resource Languages in {LLM}s through Translation-Assisted Chain-of-Thought Processes},
author={Bibek Upadhayay and Vahid Behzadan},
booktitle={5th Workshop on practical ML for limited/low resource settings, ICLR},
year={2024}… See the full description on the dataset page: https://huggingface.co/datasets/saillab/alpaca-frisian-cleaned.
