softcatala/ca_text_corpus
Dataset Card for ca-text-corpus Dataset Summary Public domain corpus of Catalan text. Supported Tasks and Leaderboards This dataset can be used as a small Catalan text corpus for language modeling, text generation experiments, sentence selection, and prompt sentence sourcing for speech datasets. It is not associated with a public leaderboard. Languages Catalan (ca). Dataset Structure Data Instances Each… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/ca_text_corpus.
Dataset Card for ca-text-corpus
Table of Contents
- Table of Contents
- Dataset Description
- Dataset Summary
- Supported Tasks and Leaderboards
- Languages
- Dataset Structure
- Data Instances
- Data Fields
- Data Splits
- Dataset Creation
- Curation Rationale
- Source Data
- Annotations
- Personal and Sensitive Information
- Considerations for Using the Data
- Social Impact of Dataset
- Discussion of Biases
- Other Known Limitations
- Additional Information
- Dataset Curators
- Licensing Information
- Citation Information
- Contributions
Dataset Description
- Homepage:
- Repository: https://github.com/Softcatala/ca-text-corpus
- Paper:
- Leaderboard:
- Point of Contact:
Dataset Summary
Public domain corpus of Catalan text.
Supported Tasks and Leaderboards
This dataset can be used as a small Catalan text corpus for language modeling, text generation experiments, sentence selection, and prompt sentence sourcing for speech datasets. It is not associated with a public leaderboard.
Languages
Catalan (ca).
Dataset Structure
Data Instances
Each instance is a Catalan sentence or short text fragment, for example a common sentence, a proverb, a sentence selected from a public-domain source, or a municipality name rendered as a sentence.
Data Fields
The dataset has one field:
text: Catalan sentence or short text fragment.
Data Splits
The dataset provides a single train split.
Dataset Creation
Curation Rationale
The dataset collects public-domain Catalan sentences that are useful for Catalan language technology projects, especially tasks that need short, reusable, permissively licensed text.
Source Data
Initial Data Collection and Normalization
The source repository groups sentences from several public-domain or permissively reusable sources, including common short sentences, proverbs, selected sentences from DOGC and DOGV public journals, Riurau Editors texts, Softcatala web pages, the book Programari lliure: tecnicament viable, economicament sostenible i socialment just, Common Voice sentence contributions, and municipality names from Catalan-speaking territories.
Who are the source language producers?
The source texts come from public institutions, publishers, Softcatala contributors, individual translators or authors, and public-domain geographic name lists, depending on the source file.
Annotations
Annotation process
No linguistic annotation is added. The dataset is a collection of source sentences and short text fragments.
Who are the annotators?
Not applicable.
Personal and Sensitive Information
The dataset is intended to contain public-domain or permissively reusable text. It may contain place names and references present in the source material, but it was not designed to collect personal or sensitive information.
Considerations for Using the Data
Social Impact of Dataset
The dataset supports Catalan language technology by making reusable Catalan text available for experimentation and downstream resource creation.
Discussion of Biases
The corpus is not a balanced sample of Catalan. It over-represents short sentences, proverbs, public-administration text, public-domain curated material, and geographic names.
Other Known Limitations
The dataset is small compared with modern web-scale text corpora and should not be used as a standalone representation of contemporary Catalan usage, dialectal variation, or conversational language.
Additional Information
Dataset Curators
Softcatala.
Licensing Information
Citation Information
If you use this dataset, cite the dataset repository:
Softcatala/ca-text-corpus: Public domain corpus of Catalan text. https://github.com/Softcatala/ca-text-corpus
Contributions
Thanks to @albertvillanova for adding this dataset.
