CoolFace
Datasetpublic

softcatala/ca_text_corpus

Dataset Card for ca-text-corpus Dataset Summary Public domain corpus of Catalan text. Supported Tasks and Leaderboards This dataset can be used as a small Catalan text corpus for language modeling, text generation experiments, sentence selection, and prompt sentence sourcing for speech datasets. It is not associated with a public leaderboard. Languages Catalan (ca). Dataset Structure Data Instances Each… See the full description on the dataset page: https://huggingface.co/datasets/softcatala/ca_text_corpus.

sourceHugging Facecc0-1.0updated 2mo agoView on Hugging Face
0likes27downloads
Dataset Card

Dataset Card for ca-text-corpus

Table of Contents

Dataset Description

  • Homepage:
  • Repository: https://github.com/Softcatala/ca-text-corpus
  • Paper:
  • Leaderboard:
  • Point of Contact:

Dataset Summary

Public domain corpus of Catalan text.

Supported Tasks and Leaderboards

This dataset can be used as a small Catalan text corpus for language modeling, text generation experiments, sentence selection, and prompt sentence sourcing for speech datasets. It is not associated with a public leaderboard.

Languages

Catalan (ca).

Dataset Structure

Data Instances

Each instance is a Catalan sentence or short text fragment, for example a common sentence, a proverb, a sentence selected from a public-domain source, or a municipality name rendered as a sentence.

Data Fields

The dataset has one field:

  • text: Catalan sentence or short text fragment.

Data Splits

The dataset provides a single train split.

Dataset Creation

Curation Rationale

The dataset collects public-domain Catalan sentences that are useful for Catalan language technology projects, especially tasks that need short, reusable, permissively licensed text.

Source Data

Initial Data Collection and Normalization

The source repository groups sentences from several public-domain or permissively reusable sources, including common short sentences, proverbs, selected sentences from DOGC and DOGV public journals, Riurau Editors texts, Softcatala web pages, the book Programari lliure: tecnicament viable, economicament sostenible i socialment just, Common Voice sentence contributions, and municipality names from Catalan-speaking territories.

Who are the source language producers?

The source texts come from public institutions, publishers, Softcatala contributors, individual translators or authors, and public-domain geographic name lists, depending on the source file.

Annotations

Annotation process

No linguistic annotation is added. The dataset is a collection of source sentences and short text fragments.

Who are the annotators?

Not applicable.

Personal and Sensitive Information

The dataset is intended to contain public-domain or permissively reusable text. It may contain place names and references present in the source material, but it was not designed to collect personal or sensitive information.

Considerations for Using the Data

Social Impact of Dataset

The dataset supports Catalan language technology by making reusable Catalan text available for experimentation and downstream resource creation.

Discussion of Biases

The corpus is not a balanced sample of Catalan. It over-represents short sentences, proverbs, public-administration text, public-domain curated material, and geographic names.

Other Known Limitations

The dataset is small compared with modern web-scale text corpora and should not be used as a standalone representation of contemporary Catalan usage, dialectal variation, or conversational language.

Additional Information

Dataset Curators

Softcatala.

Licensing Information

CC0 1.0 Universal.

Citation Information

If you use this dataset, cite the dataset repository:

Softcatala/ca-text-corpus: Public domain corpus of Catalan text. https://github.com/Softcatala/ca-text-corpus

Contributions

Thanks to @albertvillanova for adding this dataset.