datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
corts_valencianes_asr_aThis is the first version of CortsValencianes speech corpus for Valencian: a collection of speech recordings with transcriptions intended for Automatic Speech Recognition (ASR) applications.lenciclopedia-valencian-wikipedia
L'Enciclopèdia en Valencià (Normes del Puig)
Este dataset conté el corpus enciclopèdic complet i netejat de L'Enciclopèdia en Valencià, una enciclopèdia lliure escrita exclusivament en llengua valenciana.
Conta en més de 322,986 artículs enciclopèdics netejats de codi wikitext, llests per a l'entrenament, ajust fi (fine-tuning) o evaluació de models de llenguage (LLMs).
⚠️ Important Linguistic Notice for AI Researchers & NLP Models
Language Variety &… See the full description on the dataset page: https://huggingface.co/datasets/lenciclopedia/lenciclopedia-valencian-wikipedia.Spanish-Valencian_Catalan_Parallel_Corpus
Dataset Card for Spanish-Valencian Catalan Parallel Corpus
Dataset Summary
A bilingual parallel corpus containing parallel sentences in Spanish and the Valencian variant of Catalan. Built by aggregating and filtering multiple public sources, along with data obtained through direct data sharing with external partners, it provides sentence-level alignments for training Machine Translation systems. The dataset includes both authentically parallel data as well as synthetic… See the full description on the dataset page: https://huggingface.co/datasets/BSC-LT/Spanish-Valencian_Catalan_Parallel_Corpus.alia_valencian_municipalities
📘 ALIA_Valencian_Municipalities Dataset
The ALIA_Valencian_Municipalities dataset is a multilingual resource designed for text generation.
The dataset consists of textual documents formatted in Markdown (.md), each provided as structured JSONL entries.
Each entry includes information about the text's language, format, text, source, and metadata.
🧾 Column Descriptions
Field
Type
Description
format
string
Indicates the text format. All entries use "md"… See the full description on the dataset page: https://huggingface.co/datasets/gplsi/alia_valencian_municipalities.squadv2_valencian_evalsquad_valencian_evalAnythingLLMrte_valencian_validationcorts-valencianes-asrAnythingLLM-csv
