datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
image-preference-demo
Image dataset for preference aquisition demo
This dataset provides the files used to run the example that we use in this blog post to illustrate how easily
you can set up and run the annotation process to collect a huge preference dataset using Rapidata's API.
The goal is to collect human preferences based on pairwise image matchups.
The dataset contains:
Generated images: A selection of example images generated using Flux.1 and Stable Diffusion. The images are provided in a .zip… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/image-preference-demo.multilingual-llm-jokes-4o-claude-gemini
Rapidata Generated Joke Preference Dataset
We collected 1'000'000+ human opinions on the jokes generated by state-of-the-art LLMs to decide which model is the funniest. The labelers are shown a joke in their language and asked to answer 'Yes' or 'No' to the question 'Is this joke funny?'.
It took us less than 5 days to get all of the responses.
The jokes are evenly distributed across 5 languages: English, Arabic, Japanese, Vietnamese, Portuguese and across 4 model… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/multilingual-llm-jokes-4o-claude-gemini.
