datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
irishmanIf you prefer MIDI or MusicXML, download IrishMAN-MIDI or IrishMAN-XML. For better use of structural info in control codes, consider ABC notation.
Dataset Summary
The Irish Massive ABC Notation (IrishMAN) dataset includes 216,284 Irish tunes in ABC notation, divided into 99% (214,122 tunes) for training and 1% (2,162 tunes) for validation. These tunes were collected from thesession.org and abcnotation.com, both renowned for sharing traditional music. To ensure uniformity in… See the full description on the dataset page: https://huggingface.co/datasets/sander-wood/irishman.oneills-irish-tunes-1850
O'Neill's Irish Tunes (1850)
1,849 traditional Irish tunes in ABC notation, transcribed from Captain
Francis O'Neill's O'Neill's Music of Ireland: 1850 Melodies (Chicago,
1903). Each row is one tune: title, type, key, meter, and the full ABC
source.
Dataset structure
Field
Type
Description
tune_id
string
Stable id, e.g. oneills1850-1 (tune number in the original book)
name
string
Tune title
tune_type
string
Rhythm/category: reel, jig, slip jig… See the full description on the dataset page: https://huggingface.co/datasets/ecairol/oneills-irish-tunes-1850.Irish-English-Parallel-Collection
UCCIX's English-Irish Parallel Textual Corpus
Dataset Summary
This parallel English-Irish text dataset includes data from various sources such as paracrawl.eu, ECLR.
This dataset is feed to the English-centric pre-trained LLM at the start of continual pre-training, with the hypothesis to allow the LLM to draw the connections between the two languages easier, before learning on mono Irish data.
Dataset Sources
Source
Description
Statistics
Note… See the full description on the dataset page: https://huggingface.co/datasets/ReliableAI/Irish-English-Parallel-Collection.Coq-Iris
Coq-Iris
Structured dataset from Iris, a higher-order concurrent separation logic framework for Coq.
Source
Repository: https://gitlab.mpi-sws.org/iris/iris
Commit: 49cc21a1d9be19176b8546a5c5cb0076bf7b14e6
Files: 214
License: bsd-3-clause
Schema
Column
Type
Description
statement
string
Declaration signature/claim with the leading keyword removed (verbatim slice); the full declaration minus its proof
proof
string
Verbatim proof/body… See the full description on the dataset page: https://huggingface.co/datasets/phanerozoic/Coq-Iris.iris-prefix-cache-benchmark
vLLM Iris Prefix Caching Benchmark Dataset
This dataset is specifically designed to test and benchmark the Automatic Prefix Caching feature in vLLM, using the technical announcement of vLLM Semantic Router v0.1 (Iris) as the shared context.
Dataset Structure
The dataset contains 20 prompts. Each prompt consists of:
Shared Prefix: The technical overview of the vLLM Semantic Router Iris release (~500 tokens).
Unique Suffix: A specific technical question based on the text.… See the full description on the dataset page: https://huggingface.co/datasets/jaytonde05/iris-prefix-cache-benchmark.ExeBench-IRIS
Dataset Card for ExeBench-IRIS
ExeBench-IRIS is a test-set designed to assess Large Language Models on compiler Intermediate Representation (IR) translation tasks. Specifically, we use it to evaluate GIMPLE-to-LLVM IR neural translation.
It contains 1,321 executable C functions with their corresponding GIMPLE and LLVM IR representations, with 10 I/O tests each.
Dataset Structure
id
c_snippet: C source code
gimple_ir: GIMPLE IR generated by GCC v15.
llvm_ir: LLVM… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/ExeBench-IRIS.IRISIrish-Text-Collection
UCCIX's Irish Textual Corpus
Dataset Summary
This monolingual Irish text dataset includes data from various sources such as CulturaX, Glot500, Irish Wikipedia, providing valuable content from Irish sites and pages.
Our primary sources include CulturaX and Glot500, both of which provide important information from multilingual websites, including a subset dedicated to Irish. Additionally, we incorporate data from the Irish segment of the ga-en bitext pair of ParaCrawl v7… See the full description on the dataset page: https://huggingface.co/datasets/ReliableAI/Irish-Text-Collection.Iris-Uncensored-R2📜 Please read our Terms and Conditions before using this dataset.
WARNING: IRIS R2 IS SUPER SUPER UNCENSORED, CONTAIN HARMFUL CONTENT, AS OF TODAY WE HIGHLY DISCOURAGE TRAINERS
FROM USING THIS DATASET, ITS ONLY OPEN UNTIL THE END OF 2025, WE'LL TAKE THIS DOWN AFTERWARDS
lets all train an ethical yet uncensored ai shall we?
Iris Uncensored R2
Iris Uncensored R2 is a corpus of dataset Mixed and Asigned Into a Prompt and Response format, Its a Corpus of… See the full description on the dataset page: https://huggingface.co/datasets/N-Bot-Int/Iris-Uncensored-R2.Iris-Uncensored-Reformat-R2📜 Please read our Terms and Conditions before using this dataset.
THIS DATASET IS CLEANED AND REFORMATED VERSION OF Iris Uncensored R2, As the name suggest, it has no other Newly
Added Info other than it containing reformats, fix, added "*" action sequences, and added 10k worth of emoji's
PLEASE READ THE ORIGINA CARD FOR R2
lets all train an ethical yet uncensored ai shall we?
Iris Uncensored R2
Iris Uncensored R2 is a corpus of dataset Mixed and Asigned… See the full description on the dataset page: https://huggingface.co/datasets/N-Bot-Int/Iris-Uncensored-Reformat-R2.CodeForces-IRIS
Dataset Card for CodeForces-IRIS
CodeForces-IRIS is an evaluation test-set designed to assess Large Language Models on compiler intermediate representation (IR) translation tasks. Specifically, we use it to evaluate GIMPLE IR to LLVM IR code translation.
It contains 1,192 C submissions across 487 competitive programming problems with their corresponding GIMPLE and LLVM IR representations and official test cases for functional validation.
Dataset Structure
The… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/CodeForces-IRIS.Coq-Iris
Coq-Iris
Structured dataset from Iris, a higher-order concurrent separation logic framework for Coq.
Schema
Column
Type
Description
fact
string
Declaration body
type
string
Lemma, Definition, Class, Global, Local, etc.
library
string
Module (iris, iris_heap_lang, iris_unstable, etc.)
imports
list
Require/Import statements
filename
string
Source file path
symbolic_name
string
Declaration identifier
Statistics
By Type… See the full description on the dataset page: https://huggingface.co/datasets/ReactorJet/Coq-Iris.irish-english-dialectiris-tote-genui-500k-v0.1
Iris TOTE GenUI 500k v0.1
Synthetic text-to-TOTE dataset for GenUI pretraining.
Recommended HF repo name:
iris-tote-genui-500k-v0.1
Contents
This release is compiled from:
/Users/ggwplarin/src/repos/iris-train-s/runs/bronze_500k_llm_paraphrase_contract_clipped_v0002
It expands 5,000 validated gold TOTE rows into 500,000 bronze rows using deterministic LLM paraphrase overlays and contract clipping. The target TOTE is unchanged from validated gold seeds.
Splits… See the full description on the dataset page: https://huggingface.co/datasets/ggwplarin/iris-tote-genui-500k-v0.1.Roleplay-Irish
RolePlay-Irish
Roleplay-Irish Dataset is a dataset for roleplaying in the Irish language for Large Language Model.
The base dataset is GPTeacher role play dataset by teknium 1, which can be found under this link, released under MIT License. The dataset is then translated into respective languages. The translation process is powered by Google Translate, using cloud translation API.
For more information and other language datasets for roleplay, it can be found at this github repo.… See the full description on the dataset page: https://huggingface.co/datasets/jojo-ai-mst/Roleplay-Irish.
