mess
Datasets
All datasets matching “mess”hot-mess-data
Dataset Card for The Hot Mess of AI: How Does Misalignment Scale With Model Intelligence and Task Complexity?
This dataset contains the raw output of the experiments of our paper The Hot Mess of AI: How Does Misalignment Scale With Model Intelligence and Task Complexity?.
Structure
mcq/: Raw JSONL files of all runs with the LM Eval Harness Fork here.
mwe/: Model-Written Eval Suite, both multiple choice mcq and open-ended formats, obtained with the codebase of the… See the full description on the dataset page: https://huggingface.co/datasets/hot-mess/hot-mess-data.wan2.2_loramessage_historymessy-prompt-datasets
🎨 Messy Prompt Dataset
🎨 A mixed collection of AI image prompts (500+). A bit of everything — raw and uncurated. Truly open source: No login, no ads, no redirection. Just pure data for AI creators.
This project is a growing collection of diverse image generation prompts gathered from social platforms like Twitter/X. The entire dataset contains 500+ images, all structured into a comprehensive dataset.
Due to GitHub's limitations with large file storage, the full dataset… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/messy-prompt-datasets.messirveJuly 2025 UPDATE: We released version 1.1, adding almost 200k new queries 🎉🎉🎉.
v1.2 further adds the article titles as columns for convenience.
Use with:
country = "full" # "ar", "bo", ...
version = "1.2"
dataset = datasets.load_dataset("spanish-ir/messirve", country, revision=version)
print(dataset)
Dataset Card for MessIRve
MessIRve is a large-scale dataset for Spanish IR, designed to better capture the information needs of Spanish speakers across different countries.… See the full description on the dataset page: https://huggingface.co/datasets/spanish-ir/messirve.alpaca_messages_2k_dpo_test
