datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CD-ESA
CD-ESA: Cross-Domain Error Span Annotation Dataset
This dataset contains the publicly releasable WMT23 and Emea portions of CD-ESA (Cross-Domain Error Span Annotation), introduced in our work “Who Watches the Watchmen? Humans Disagree With Translation Metrics on Unseen Domains”. CD-ESA was created to study how well reference-free machine translation evaluation metrics, i.e. quality estimation (QE) metrics, generalize to unseen domains. The release comprises 4,728 translation rows… See the full description on the dataset page: https://huggingface.co/datasets/FinnSchmidt/CD-ESA.test-tracesesab-filler-metal-recommendations-by-astm-steel-grade
ESAB recommended filler metals and suggested preheat group, by ASTM steel grade
Canonical, always-current version: https://referencesource.org/esab-filler-metal-recommendations-by-astm-steel-grade/
Machine-readable: https://referencesource.org/esab-filler-metal-recommendations-by-astm-steel-grade/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-19
Stale after: 2028-08-18 (past this date, prefer the canonical copy —
it re-verifies on a cadence this… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/esab-filler-metal-recommendations-by-astm-steel-grade.processed-turkish-sentencesmorpeheme_processed_tinystoriesThis is a modified version of a Turkish TinyStories dataset: https://huggingface.co/datasets/umarigan/tinystories_tr
The modification is that suffixes are replaced with PUA (Private Use Area) characters. Suffixes are preceded by a label that specifies whether the word is a noun, verb, or named entity.
Here is a full list of all suffixes and word labels:
[
"Root", "Noun", "Adj", "Verb", "Pron", "Adv", "Conj", "Punc", "Ques",
"Postp", "Det", "Num", "Dup", "Interj", "A1sg", "A2sg"… See the full description on the dataset page: https://huggingface.co/datasets/esat-krky/morpeheme_processed_tinystories.llama-2-papersLyriQGen.dataset
