datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
irishmanIf you prefer MIDI or MusicXML, download IrishMAN-MIDI or IrishMAN-XML. For better use of structural info in control codes, consider ABC notation.
Dataset Summary
The Irish Massive ABC Notation (IrishMAN) dataset includes 216,284 Irish tunes in ABC notation, divided into 99% (214,122 tunes) for training and 1% (2,162 tunes) for validation. These tunes were collected from thesession.org and abcnotation.com, both renowned for sharing traditional music. To ensure uniformity in… See the full description on the dataset page: https://huggingface.co/datasets/sander-wood/irishman.oneills-irish-tunes-1850
O'Neill's Irish Tunes (1850)
1,849 traditional Irish tunes in ABC notation, transcribed from Captain
Francis O'Neill's O'Neill's Music of Ireland: 1850 Melodies (Chicago,
1903). Each row is one tune: title, type, key, meter, and the full ABC
source.
Dataset structure
Field
Type
Description
tune_id
string
Stable id, e.g. oneills1850-1 (tune number in the original book)
name
string
Tune title
tune_type
string
Rhythm/category: reel, jig, slip jig… See the full description on the dataset page: https://huggingface.co/datasets/ecairol/oneills-irish-tunes-1850.Irish-English-Parallel-Collection
UCCIX's English-Irish Parallel Textual Corpus
Dataset Summary
This parallel English-Irish text dataset includes data from various sources such as paracrawl.eu, ECLR.
This dataset is feed to the English-centric pre-trained LLM at the start of continual pre-training, with the hypothesis to allow the LLM to draw the connections between the two languages easier, before learning on mono Irish data.
Dataset Sources
Source
Description
Statistics
Note… See the full description on the dataset page: https://huggingface.co/datasets/ReliableAI/Irish-English-Parallel-Collection.Irish-Text-Collection
UCCIX's Irish Textual Corpus
Dataset Summary
This monolingual Irish text dataset includes data from various sources such as CulturaX, Glot500, Irish Wikipedia, providing valuable content from Irish sites and pages.
Our primary sources include CulturaX and Glot500, both of which provide important information from multilingual websites, including a subset dedicated to Irish. Additionally, we incorporate data from the Irish segment of the ga-en bitext pair of ParaCrawl v7… See the full description on the dataset page: https://huggingface.co/datasets/ReliableAI/Irish-Text-Collection.irish-english-dialectRoleplay-Irish
RolePlay-Irish
Roleplay-Irish Dataset is a dataset for roleplaying in the Irish language for Large Language Model.
The base dataset is GPTeacher role play dataset by teknium 1, which can be found under this link, released under MIT License. The dataset is then translated into respective languages. The translation process is powered by Google Translate, using cloud translation API.
For more information and other language datasets for roleplay, it can be found at this github repo.… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Roleplay-Irish.
