datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
research-ideation-arena-si-rm
Research Ideation Arena — Scientific Ideation RM Splits
Derived from Research Ideation Arena, revision f5704385bd66781d504e44810a9a8b56c1623b7a.
Original authors: Zhiyu Chen et al. See the paper and official code.
Splits and evaluation caveat
Train: 3,047 preference pairs. Test: 500 fixed preference pairs.
All remaining pairs from the 3,547-pair filtered pool are assigned to training.
Exact sample/pair overlap is zero, but 607 training rows share a connected… See the full description on the dataset page: https://huggingface.co/datasets/tintin1027/research-ideation-arena-si-rm.brainstorming-ideation-sft-100k
Brainstorming and Ideation SFT (100K)
100,000 ShareGPT conversations demonstrating structured, high-quality brainstorming and ideation across 22 professional domains. Each example takes a realistic context and constraint, then generates specific, actionable, well-reasoned ideas — not generic advice dressed as creativity.
Motivation
Brainstorming and ideation is one of the highest-value use cases for AI assistants, and one where models routinely underperform:… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/brainstorming-ideation-sft-100k.
