datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
setimes-en-tr-aligned-corpus
SETimes EN-TR — Sentence-Aligned, LLM-Cleaned
A cleaned and re-aligned version of the SETimes English-Turkish parallel corpus. The original SETimes data is paragraph-style — each "pair" can contain a headline, a dateline, several body sentences, and a source citation, all glued together on one line. This version splits everything into proper sentence pairs so each row is one English sentence next to its Turkish translation.
144,064 sentence pairs, split into train (142,064)… See the full description on the dataset page: https://huggingface.co/datasets/atahanuz/setimes-en-tr-aligned-corpus.ud-conll2017-aligned
UD CoNLL-U 2017 aligned
This dataset is a processed version of the dataset made available for the UD CoNLL Shared Task 2017, entitled "Multilingual Parsing from Raw Text to Universal Dependencies." This task is based on the Universal Dependencies 2.0 dataset.
The processing aligns words and tokens along with their morphological annotations using the universal part-of-speech set for sequence-to-sequence task training.
The dataset fields are described in the table below:
Field… See the full description on the dataset page: https://huggingface.co/datasets/fdemelo/ud-conll2017-aligned.safety_aligned_datasets
Safety Aligned Datasets
A high-fidelity adversarial corpus engineered for alignment research, refusal boundary modeling, and robustness evaluation of Small Language Models.
The Problem This Solves
Fine-tuning a Small Language Model to be safe is not the same as fine-tuning it to understand safety.
Most safety datasets give models clean refusal examples on obvious prompts — and those models fail the moment an adversary wraps a harmful request in a… See the full description on the dataset page: https://huggingface.co/datasets/vvsd-charan/safety_aligned_datasets.Human_Aligned_Benchnomnaocr-alignedOdia_English_Sentences_Aligned_81k
🌐 Vaqas AI: Odia-English Sentence-Aligned Precision Corpus
Created and Maintained by Vaqas AI Creator & Owner: Vaqas Ahmed
🚀 Why This Dataset?
Standard parallel corpora often consist of large paragraphs that exceed the 4k or 8k context windows typical of LLM training, leading to truncated data and poor model performance.
The Vaqas AI Sentence-Aligned Corpus was engineered to solve this. We transformed the original vaqasai/Odia_English_News dataset by decomposing dense… See the full description on the dataset page: https://huggingface.co/datasets/vaqasai/Odia_English_Sentences_Aligned_81k.singlish-english-aligned-translation-seat
Singlish-English Aligned Translation (SEAT) Dataset
Dataset Description
The Singlish-English Aligned Translation (SEAT) Dataset is a sentence-aligned parallel corpus for translation from Singlish to English.It contains synthetic and/or real sentence pairs for NLP research and model training.
Language(s): Singlish, English
Size: X examples (train + validation + test)
Task: Translation
Dataset Structure
Each row in the dataset has the following columns:… See the full description on the dataset page: https://huggingface.co/datasets/kumarsujit474/singlish-english-aligned-translation-seat.setimes-en-tr-aligned-corpus-model-answers
SETimes EN-TR — Model Answers (Test Set)
Translation outputs from two NMT architectures (a Transformer and an RNN seq2seq) on the 1,000-sentence test split of the SETimes EN-TR aligned corpus. Both models were trained on the same data with a joint 32k BPE vocabulary, and each was run in both directions (Turkish→English and English→Turkish). Each row pairs the human reference translations with all four model hypotheses, so the file is self-contained for re-scoring.
1,000… See the full description on the dataset page: https://huggingface.co/datasets/atahanuz/setimes-en-tr-aligned-corpus-model-answers.
