datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rfm-rm-as-user-dataset
RFM Reward Model As User Dataset
This dataset was generated for the NeurIPS 2025 paper titled "Capturing Individual Human Preferences with Reward Features". It is released to support the reproducibility of the experiments described in the paper, particularly those in the "Modelling groups of real users" section.
Instead of containing preferences from human raters, this dataset uses 8 publicly available reward models (RMs) as proxies for human raters. This allows for large-scale… See the full description on the dataset page: https://huggingface.co/datasets/google/rfm-rm-as-user-dataset.Copy_Dakshina_Google_research_dataset
Copy_Dakshina_Google_research_dataset
This repository is a structured, processed version of the Dakshina Dataset, originally released by Google Research. It has been reorganized into a unified Hugging Face format to support NLP research in South Asian languages, specifically focusing on sentence-level and word-level transliteration tasks.
Dataset Overview
The original Dakshina dataset is a collection of text in both Latin and native scripts for 12 South Asian… See the full description on the dataset page: https://huggingface.co/datasets/Anvesh-Lankala/Copy_Dakshina_Google_research_dataset.
