CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ReliableAI /irish_fineweb_eduData translation project of https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu, sample-10BT subset. Data are translated from English to Irish using NLLB-3.3B. tabular100K<n<1M1 likes6.7k downloads2y agoHugging Face02sander-wood /irishmanIf you prefer MIDI or MusicXML, download IrishMAN-MIDI or IrishMAN-XML. For better use of structural info in control codes, consider ABC notation. Dataset Summary The Irish Massive ABC Notation (IrishMAN) dataset includes 216,284 Irish tunes in ABC notation, divided into 99% (214,122 tunes) for training and 1% (2,162 tunes) for validation. These tunes were collected from thesession.org and abcnotation.com, both renowned for sharing traditional music. To ensure uniformity in… See the full description on the dataset page: https://huggingface.co/datasets/sander-wood/irishman.texttext-generation100K<n<1M28 likes589 downloads3y agoHugging Face03isaacus /irish-legislative-summaries Irish Legislative Summaries ⚖️ Irish Legislative Summaries by Isaacus is a novel, challenging legal information retrieval evaluation dataset consisting of 500 Irish laws and their long titles, succinctly summarizing subject matter, scope, and purpose of legislation. This dataset is meant to stress test the ability of an information retrieval model to retrieve relevant statutes to short queries describing them. This dataset forms part of the Massive Legal Embeddings Benchmark (MLEB)… See the full description on the dataset page: https://huggingface.co/datasets/isaacus/irish-legislative-summaries.texttext-retrieval1K<n<10K2 likes428 downloads11mo agoHugging Face04irisxx /ultrafeedback_tied Train dir contains train set with different ratios of tie data Test dir contains test sets which used to evaluate performances on the in-distribution data. test_data.jsonl contains 2000 samples consist of 1500 non-tie data and 500 tie data. non_tie_data_test.jsonl contains 1500 non-tie samples. tie_data_test.jsonl contains 500 tie samples. Citation Please cite our paper if you find the dataset helpful in your work: @inproceedings{ guo2025todo, title={{TODO}:… See the full description on the dataset page: https://huggingface.co/datasets/irisxx/ultrafeedback_tied.tabular10K<n<100K0 likes107 downloads1y agoHugging Face05Rev3auth /iris-mixtext1K<n<10K1 likes100 downloads7d agoHugging Face06Shadow0482 /iris_02text10K<n<100K0 likes86 downloads1y agoHugging Face07ecairol /oneills-irish-tunes-1850 O'Neill's Irish Tunes (1850) 1,849 traditional Irish tunes in ABC notation, transcribed from Captain Francis O'Neill's O'Neill's Music of Ireland: 1850 Melodies (Chicago, 1903). Each row is one tune: title, type, key, meter, and the full ABC source. Dataset structure Field Type Description tune_id string Stable id, e.g. oneills1850-1 (tune number in the original book) name string Tune title tune_type string Rhythm/category: reel, jig, slip jig… See the full description on the dataset page: https://huggingface.co/datasets/ecairol/oneills-irish-tunes-1850.texttext-generation1K<n<10K0 likes80 downloads23d agoHugging Face08Clemylia /Iris-datatextn<1K0 likes79 downloads10mo agoHugging Face09temsa /OpenMed-Irish-CorePII-TrainMix-v1 OpenMed Irish Core PII Train Mix v1 Composite token-classification training mix used to fine-tune temsa/OpenMed-mLiteClinical-IrishCorePII-135M-v1. This repo is the training dataset, not the model itself. What A Row Looks Like Each row uses a fixed schema so the Hugging Face dataset viewer and datasets.load_dataset() can read it directly: id: row id inside the split text: reconstructed text string tokens: tokenized text labels: BIO labels aligned to tokens language:… See the full description on the dataset page: https://huggingface.co/datasets/temsa/OpenMed-Irish-CorePII-TrainMix-v1.tabulartoken-classification10K<n<100K0 likes43 downloads7mo agoHugging Face10ReliableAI /IrishQAgated IrishQA textn<1K0 likes40 downloads1y agoHugging Face11irisxx /chatarena_tied Citation Please cite our paper if you find the dataset helpful in your work: @inproceedings{ guo2025todo, title={{TODO}: Enhancing {LLM} Alignment with Ternary Preferences}, author={Yuxiang Guo and Lu Yin and Bo Jiang and Jiaqi Zhang}, booktitle={The Thirteenth International Conference on Learning Representations}, year={2025}, url={https://openreview.net/forum?id=utkGLDSNOk} } tabular10K<n<100K1 likes38 downloads1y agoHugging Face12jgracie52 /scanner-poisoned-iris-benchmark Scanner Poisoned Iris Benchmark This benchmark starts from the classic UCI Iris dataset and injects multiple synthetic poisoning patterns so dataset scanners can exercise duplicate, anomaly, missingness, skew, and divergence heuristics against a small tabular corpus. Recommended Hugging Face repo slug: your-org/scanner-poisoned-iris-benchmark What It Is For benchmarking dataset quality and poisoning detection workflows regression-testing scanner heuristics on a… See the full description on the dataset page: https://huggingface.co/datasets/jgracie52/scanner-poisoned-iris-benchmark.tabulartabular-classificationn<1K0 likes24 downloads2mo agoHugging Face13Shadow0482 /iris_v3text10K<n<100K0 likes18 downloads1y agoHugging Face14temsa /OpenMed-Irish-PPSN-Eircode-Spec-v1 OpenMed Irish PPSN Eircode Spec v1 Focused synthetic token-classification dataset for Irish PPSN and Eircode detection. This repo contains synthetic training rows, not a fine-tuned model. What A Row Looks Like Each row uses a fixed schema: id: row id inside the split text: rendered text string tokens: tokenized text labels: BIO labels aligned to tokens language: en or ga source_dataset: generator identifier source_domain: optional domain tag, empty in this release… See the full description on the dataset page: https://huggingface.co/datasets/temsa/OpenMed-Irish-PPSN-Eircode-Spec-v1.tabulartoken-classification10K<n<100K0 likes18 downloads7mo agoHugging Face15build-small-hackathon /iris-agent-trace Iris: Coding-Agent Build Trace 👁️ A record of how Iris, a voice-first assistant for blind and low-vision people, was built with Claude Code (Claude Opus 4.8) during the Build Small Hackathon. It follows the build session step by step: the decisions, the tool calls, the debugging, and what each step taught. Iris was made for the author's father, who is blind, so the trace runs from the first idea through to a working app on a phone. Earns the Sharing is Caring (open trace)… See the full description on the dataset page: https://huggingface.co/datasets/build-small-hackathon/iris-agent-trace.textn<1K0 likes18 downloads4mo agoHugging Face16c123ian /iris_eng_combo_dpotext10K<n<100K0 likes17 downloads2y agoHugging Face17ReliableAI /irish_retrieval_datatabular100K<n<1M0 likes14 downloads2y agoHugging Face18c123ian /dpo_irish_eng_translationsThis is a test for my DPO dataset for Irish ENglish trasnlslations, raw data origin : https://www.gaois.ie/en/corpora/parallel?Query=Apple&Language=en&SearchMode=exact&PerPage=50, used COMETXL refrernce free maodel Unbabel/wmt23-cometkiwi-da-xl (which has been trained to asses Irish) to score accepted/rejected. Used GPT4 to generate translations to compare with human stranslations of Irish legislation (which has to have a Irisng/English copy by law) tabularn<1K0 likes11 downloads2y agoHugging Face19HPAI-BSC /GNU-IRISgated Dataset Card for GNU-IRIS GNU-IRIS is a training dataset for GIMPLE IR to LLVM IR translation, derived from GNU utilities source code. It contains 13,049 C functions paired with their corresponding GIMPLE and LLVM intermediate representations. Dataset Structure The default dataset combines GNU utils projects into a single training dataset. However, you can also access each GNU utility individually as a split: Configuration Package # Samples default All… See the full description on the dataset page: https://huggingface.co/datasets/HPAI-BSC/GNU-IRIS.texttranslation10K<n<100K0 likes11 downloads5d agoHugging Face20ud-synthetic /irish-passports Disclaimer: All passport images and associated data in this dataset are synthetically generated and do not correspond to real individuals. Any names, numbers, or personal details are fictional and used solely for research and development purposes. Introduction - Ireland The Synthetic Ireland Passports Dataset gathers more than 1,000 AI-generated passport images created for training OCR and computer vision models on identity documents. Every record is fully synthetic, so the… See the full description on the dataset page: https://huggingface.co/datasets/ud-synthetic/irish-passports.textimage-to-textn<1K1 likes11 downloads2mo agoHugging Face21Shadow0482 /iris_0vtext10K<n<100K0 likes10 downloads1y agoHugging Face22Shadow0482 /fin_iristext1K<n<10K0 likes10 downloads1y agoHugging Face23Shadow0482 /DPO_IRIS_Datatext1K<n<10K0 likes9 downloads1y agoHugging Face24ReliableAI /irish_belebelegatedIrish version of https://huggingface.co/datasets/facebook/belebele. Translated using facebook/nllb-200-3.3B, and the translations are verified by native Irish speakers. tabularquestion-answeringn<1K0 likes8 downloads2y agoHugging Face25aandvalenzuela /iris-l2tabularn<1K0 likes8 downloads1y agoHugging Face26Whomstt /irish-english-dialecttexttext-generationn<1K0 likes8 downloads8mo agoHugging Face27IrishinZg /si-rag-recursive-testtextn<1K0 likes6 downloads1mo agoHugging Face28IrishinZg /si-rag-recursive-originalstextn<1K0 likes6 downloads1mo agoHugging Face29IrishinZg /si-rag-flat-to-recursivetextn<1K0 likes6 downloads1mo agoHugging Face30IrishinZg /si-process-meta-testtextn<1K0 likes5 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.