CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01projecte-aina /corts_valencianes_asr_aThis is the first version of CortsValencianes speech corpus for Valencian: a collection of speech recordings with transcriptions intended for Automatic Speech Recognition (ASR) applications.automatic-speech-recognition1K<n<10K2 likes175 downloads1y agoHugging Face02lenciclopedia /lenciclopedia-valencian-wikipedia L'Enciclopèdia en Valencià (Normes del Puig) Este dataset conté el corpus enciclopèdic complet i netejat de L'Enciclopèdia en Valencià, una enciclopèdia lliure escrita exclusivament en llengua valenciana. Conta en més de 322,986 artículs enciclopèdics netejats de codi wikitext, llests per a l'entrenament, ajust fi (fine-tuning) o evaluació de models de llenguage (LLMs). ⚠️ Important Linguistic Notice for AI Researchers & NLP Models Language Variety &… See the full description on the dataset page: https://huggingface.co/datasets/lenciclopedia/lenciclopedia-valencian-wikipedia.texttext-generation100K<n<1M0 likes53 downloads12d agoHugging Face03BSC-LT /Spanish-Valencian_Catalan_Parallel_Corpus Dataset Card for Spanish-Valencian Catalan Parallel Corpus Dataset Summary A bilingual parallel corpus containing parallel sentences in Spanish and the Valencian variant of Catalan. Built by aggregating and filtering multiple public sources, along with data obtained through direct data sharing with external partners, it provides sentence-level alignments for training Machine Translation systems. The dataset includes both authentically parallel data as well as synthetic… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/Spanish-Valencian_Catalan_Parallel_Corpus.texttranslation1M<n<10M2 likes39 downloads7mo agoHugging Face04gplsi /alia_valencian_municipalities 📘 ALIA_Valencian_Municipalities Dataset The ALIA_Valencian_Municipalities dataset is a multilingual resource designed for text generation. The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries. Each entry includes information about the text's language, format, text, source, and metadata. 🧾 Column Descriptions Field Type Description format string Indicates the text format. All entries use "md"… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_valencian_municipalities.text-generation10K<n<100K0 likes21 downloads4mo agoHugging Face05rsepulvedat /squadv2_valencian_evaltext10K<n<100K0 likes10 downloads2y agoHugging Face06rsepulvedat /squad_valencian_evaltext10K<n<100K0 likes9 downloads2y agoHugging Face07valencianatasha /AnythingLLMtabulartext-generationn<1K0 likes9 downloads2y agoHugging Face08rsepulvedat /rte_valencian_validationtabularn<1K0 likes5 downloads2y agoHugging Face09gplsi /corts-valencianes-asrgatedaudio10K<n<100K0 likes5 downloads11mo agoHugging Face10valencianatasha /AnythingLLM-csvtextn<1K0 likes3 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.