CoolFace
Datasetpublic

fhgr/sursilvan-sprachspende-2025

Sursilvan Dataset - Sprachspende 2025 Dataset Description This dataset contains authentic conversational exchanges in Sursilvan, a Romansh idiom spoken in the Surselva region of Switzerland. The data was collected through a language donation survey ("Sprachspende") conducted in 2025 by FH Graubünden (FHGR), capturing natural language use across various everyday topics. Participants were native Sursilvan speakers who voluntarily contributed their language samples.… See the full description on the dataset page: https://huggingface.co/datasets/fhgr/sursilvan-sprachspende-2025.

sourceHugging Facecc-by-4.0updated 6mo agoView on Hugging Face
0likes28downloads
Dataset Card

Sursilvan Dataset - Sprachspende 2025

Dataset Description

This dataset contains authentic conversational exchanges in Sursilvan, a Romansh idiom spoken in the Surselva region of Switzerland. The data was collected through a language donation survey ("Sprachspende") conducted in 2025 by FH Graubünden (FHGR), capturing natural language use across various everyday topics. Participants were native Sursilvan speakers who voluntarily contributed their language samples.

Dataset Summary

  • —Language: Romansh Sursilvan
  • —Format: Conversational dialogues in ShareGPT format
  • —Size: 775 dialogues from 79 native speakers
  • —Topics: Daily activities, traditions, holidays, hobbies, work, social interactions
  • —Purpose: Language preservation, low-resource NLP, conversational AI

Supported Tasks

  • —Conversational AI: Training chatbots and dialogue systems
  • —Text Generation: Fine-tuning language models for Sursilvan
  • —Question Answering: Question-response pairs in authentic Sursilvan
  • —Cultural Documentation: Preserving expressions of Sursilvan culture and daily life

Dataset Structure

Data Format

Each dialogue follows the standard chat format with three messages:

json
{
  "conversations": [
    {
      "from": "system",
      "value": "Persuna:\n- vegliadetgna: 30-49 onns\n- schlattaina: masculin\n- liug: Bonaduz, derivonza: Lumnezia"
    },
    {
      "from": "human",
      "value": "Co festiveschas ti Nadal? Tgei fiastas enconuschas ti?"
    },
    {
      "from": "gpt",
      "value": "Cun mia famiglia cun pigniel da Nadal e schenghetgs. Jeu enconuschel pliras fiastas."
    }
  ]
}

Data Fields

Conversations Array
  • —from: One of "system", "human", or "gpt"
  • —value: The message text
System Message

Contains speaker demographics:

  • —vegliadetgna (age group): < 18 onns, 18-29 onns, 30-49 onns, 50-69 onns, > 70 onns
  • —schlattaina (gender): masculin, feminin
  • —liug (location): Place of residence and region of origin

Dataset Statistics

MetricValue
Total dialogues775
Unique speakers79
Unique questions12
Avg. dialogues per speaker9.8
Avg. question length75 characters
Avg. answer length154 characters
Answer length range1-1,008 characters

Demographic Distribution

Age Groups:

  • —30-49 onns: 280 dialogues (36%)
  • —50-69 onns: 272 dialogues (35%)
  • —\> 70 onns: 132 dialogues (17%)
  • —18-29 onns: 79 dialogues (10%)
  • —< 18 onns: 12 dialogues (2%)

Gender:

  • —feminin: 511 dialogues (66%)
  • —masculin: 264 dialogues (34%)

Topics Covered

The dataset includes natural responses about:

  1. 1.Personal celebrations: Birthdays, wishes, and gift-giving
  2. 2.Holiday traditions: Christmas and other festivals
  3. 3.Daily routines: Morning, afternoon, evening activities, weekend plans
  4. 4.Seasonal activities: Spring, summer, autumn, winter
  5. 5.Hobbies and leisure: Sports, music, free time activities
  6. 6.Work and professions: Career interests and preferences
  7. 7.Social interactions: Casual conversations and conflicts
  8. 8.Local life: Descriptions of villages, towns, and favorite places

Licensing

This dataset is released under CC-BY-4.0 (Creative Commons Attribution 4.0 International).

Additional Information

Contact

For more information on this project and the team, please visit the project website: fhgr.ch/idiomvoice

Acknowledgments

Special thanks to all participants who generously donated their language samples to support Sursilvan language preservation and research.

This project was financially supported by the Amt für Kultur Graubünden, Sprachenförderung. We are grateful for their commitment to preserving and promoting the Romansh language.