datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
English_Malayalam_Translation_Human_annotated
English-Malayalam Government Parallel Corpus Synth
This dataset contains synthetic machine-translated English-Malayalam sentence pairs aligned from government and administrative text.
Machine Translation Notice
All parallel text in this dataset should be treated as synthetic machine-translated data. It is intended for research, corpus filtering, model adaptation, and experimentation. It should not be treated as human-verified gold translation without additional… See the full description on the dataset page: https://huggingface.co/datasets/icfoss/English_Malayalam_Translation_Human_annotated.Calibration-translation-human-eval
Translation Evaluation Dataset: Tower vs Calibration
This dataset compares translations generated by two models ("Tower-system" and "Calibration") along with human ratings.
clinical-preclinical-human-translation-coherence-v0.1What this repo does
This dataset tests whether preclinical results translate coherently into early human trials.
Many drug programs show strong effects in animals or cell models but fail in humans.
The failure is often visible before Phase 2.
It appears as a mismatch between the model, the biology, the exposure, and the early human signal.
This dataset trains a model to detect that translation risk.
You are given
the preclinical model used
how relevant that model is to human disease
whether… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/clinical-preclinical-human-translation-coherence-v0.1.Calibration-translation-human-eval
Translation Evaluation Dataset: Tower vs Calibration
This dataset compares translations generated by two models ("Tower-system" and "Calibration") along with human ratings.
