datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Irish-English-Parallel-Collection
UCCIX's English-Irish Parallel Textual Corpus
Dataset Summary
This parallel English-Irish text dataset includes data from various sources such as paracrawl.eu, ECLR.
This dataset is feed to the English-centric pre-trained LLM at the start of continual pre-training, with the hypothesis to allow the LLM to draw the connections between the two languages easier, before learning on mono Irish data.
Dataset Sources
Source
Description
Statistics
Note… See the full description on the dataset page: https://huggingface.co/datasets/ReliableAI/Irish-English-Parallel-Collection.iranian-elderly-psychospiritual-interviews
Iranian Elderly Psycho-Spiritual Interviews
A culturally grounded, fully synthetic conversational interview dataset for assessing the mental and spiritual health of Iranian older adults, generated using Large Language Models.
Dataset Summary
This dataset introduces a culturally grounded, fully synthetic conversational interview corpus designed for the assessment and analysis of mental and spiritual health among Iranian older adults. All interviews are conducted in… See the full description on the dataset page: https://huggingface.co/datasets/liamirali/iranian-elderly-psychospiritual-interviews.offences_and_penalties_in_general_2018_datasetircam-dglai-dataset
Dataset Card for DGLAi-Augmented
An augmented, quad-lingual electronic parallel dataset based on the Dictionnaire Général de la Langue Amazighe (DGLAi). The original standard source fields (Amazigh, Arabic, French) are curated by IRCAM, supplemented with automated English translations to maximize its utility for modern machine translation, cross-lingual NLP, and LLM applications.
Dataset Details
Dataset Description
The original Dictionnaire Général… See the full description on the dataset page: https://huggingface.co/datasets/abdelhaqueidali/ircam-dglai-dataset.Irish-Text-Collection
UCCIX's Irish Textual Corpus
Dataset Summary
This monolingual Irish text dataset includes data from various sources such as CulturaX, Glot500, Irish Wikipedia, providing valuable content from Irish sites and pages.
Our primary sources include CulturaX and Glot500, both of which provide important information from multilingual websites, including a subset dedicated to Irish. Additionally, we incorporate data from the Irish segment of the ga-en bitext pair of ParaCrawl v7… See the full description on the dataset page: https://huggingface.co/datasets/ReliableAI/Irish-Text-Collection.Roleplay-Irish
RolePlay-Irish
Roleplay-Irish Dataset is a dataset for roleplaying in the Irish language for Large Language Model.
The base dataset is GPTeacher role play dataset by teknium 1, which can be found under this link, released under MIT License. The dataset is then translated into respective languages. The translation process is powered by Google Translate, using cloud translation API.
For more information and other language datasets for roleplay, it can be found at this github repo.… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Roleplay-Irish.
