datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
suicide-ideation-detection
Suicide Ideation Detection Dataset
Dataset Description
This dataset contains 232,000 Reddit posts from r/SuicideWatch, labeled for self-harm intent and suicidal ideation.
Dataset Summary
Size: 232,000 samples
Source: Reddit r/SuicideWatch community (Kaggle)
Task: Binary classification (suicidal ideation vs. non-suicidal content)
Language: English
Data Fields
text: The Reddit post content (string)
label: Binary label (0 = no suicidal ideation, 1 =… See the full description on the dataset page: https://huggingface.co/datasets/indominousx/suicide-ideation-detection.research-ideation-arena-dataset
Research Ideation
This repository contains the public data payload for the Research Ideation benchmark, presented in the paper Ideation Arena: Evaluating LLM Generated Research Ideas with Battle-style Human Expert Assessment. The code is available at https://github.com/foss12138/Research-Ideation-Arena.
Files
final_ideation_results_with_response.json
Final pairwise ideation evaluation records with released responses.
queries.json
Query definitions used by the… See the full description on the dataset page: https://huggingface.co/datasets/yolo1213811/research-ideation-arena-dataset.research-ideation-arena-si-rm
Research Ideation Arena — Scientific Ideation RM Splits
Derived from Research Ideation Arena, revision f5704385bd66781d504e44810a9a8b56c1623b7a.
Original authors: Zhiyu Chen et al. See the paper and official code.
Splits and evaluation caveat
Train: 3,047 preference pairs. Test: 500 fixed preference pairs.
All remaining pairs from the 3,547-pair filtered pool are assigned to training.
Exact sample/pair overlap is zero, but 607 training rows share a connected… See the full description on the dataset page: https://huggingface.co/datasets/tintin1027/research-ideation-arena-si-rm.relabelled-suicide-ideation
Relabelled Suicide Ideation Dataset
Overview
This dataset is a relabelled version of the publicly available Suicide Watch dataset, originally introduced in the paper:
K. Nikhileswar, D. Vishal, L. Sphoorthi and S. Fathimabi, "Suicide Ideation Detection in Social Media Forums," 2021 2nd International Conference on Smart Electronics and Communication (ICOSEC), Trichy, India, 2021, pp. 1741–1747, doi: 10.1109/ICOSEC51865.2021.9591887
The dataset was obtained from… See the full description on the dataset page: https://huggingface.co/datasets/lensy111/relabelled-suicide-ideation.brainstorming-ideation-sft-100k
Brainstorming and Ideation SFT (100K)
100,000 ShareGPT conversations demonstrating structured, high-quality brainstorming and ideation across 22 professional domains. Each example takes a realistic context and constraint, then generates specific, actionable, well-reasoned ideas — not generic advice dressed as creativity.
Motivation
Brainstorming and ideation is one of the highest-value use cases for AI assistants, and one where models routinely underperform:… See the full description on the dataset page: https://huggingface.co/datasets/stindardlogic/brainstorming-ideation-sft-100k.transformed_Suicidal_ideationsi_et_al-ideation-gpt41mini-20260414_230629-ideas
si_et_al-ideation-gpt41mini-20260414_230629-ideas
Per-idea flat table with LLM-judge scores.
Parent: si_et_al-ideation-gpt41mini-20260414_230629
Columns
topic: NLP topic
idea_index: position within pooled topic idea set
run_idx: which of the n_runs generation calls produced this idea
idea_text: the generated idea (Problem/Existing Methods/Motivation/Proposed Method/Experiment Plan)
overall: 1-10 LLM-judge score (port of ai_researcher/src/idea_direct_score.py)… See the full description on the dataset page: https://huggingface.co/datasets/strategy-scope/si_et_al-ideation-gpt41mini-20260414_230629-ideas.ideationStreamRanksi_et_al-ideation-gpt41mini-20260414_221654-evaluation
si_et_al-ideation-gpt41mini-20260414_221654-evaluation
Aggregate per-topic metrics (+ overall row where topic = __overall__).
Parent: si_et_al-ideation-gpt41mini-20260414_221654
Run Parameters
model: gpt-4.1-mini-2025-04-14
judge_model: gpt-4o-mini-2024-07-18
strategy: none
n_ideas_per_run: 5
n_runs: 3
n_topics: 3
rag: False
Columns
topic / n_ideas
mean_{novelty,excitement,feasibility,effectiveness,overall}: LLM-judge means
diversity_cosine_{pairwise,nn} /… See the full description on the dataset page: https://huggingface.co/datasets/strategy-scope/si_et_al-ideation-gpt41mini-20260414_221654-evaluation.si_et_al-ideation-gpt41mini-20260414_223535-ideas
si_et_al-ideation-gpt41mini-20260414_223535-ideas
Per-idea flat table with LLM-judge scores.
Parent: si_et_al-ideation-gpt41mini-20260414_223535
Columns
topic: NLP topic
idea_index: position within pooled topic idea set
run_idx: which of the n_runs generation calls produced this idea
idea_text: the generated idea (Problem/Existing Methods/Motivation/Proposed Method/Experiment Plan)
overall: 1-10 LLM-judge score (port of ai_researcher/src/idea_direct_score.py)… See the full description on the dataset page: https://huggingface.co/datasets/strategy-scope/si_et_al-ideation-gpt41mini-20260414_223535-ideas.si_et_al-ideation-gpt5mini-full-20260415_022018-ideas
si_et_al-ideation-gpt5mini-full-20260415_022018-ideas
Per-idea flat table with LLM-judge scores.
Parent: si_et_al-ideation-gpt5mini-full-20260415_022018
Columns
topic: NLP topic
idea_index: position within pooled topic idea set
run_idx: which of the n_runs generation calls produced this idea
idea_text: the generated idea (Problem/Existing Methods/Motivation/Proposed Method/Experiment Plan)
overall: 1-10 LLM-judge score (port of ai_researcher/src/idea_direct_score.py)… See the full description on the dataset page: https://huggingface.co/datasets/strategy-scope/si_et_al-ideation-gpt5mini-full-20260415_022018-ideas.si_et_al-ideation-gpt41mini-20260414_233911
si_et_al-ideation-gpt41mini-20260414_233911
Benchmark: si_et_al
Generated: 2026-04-14T23:49:15.816891
Parameters
model: gpt-4.1-mini-2025-04-14
judge_model: anthropic/claude-sonnet-4-5-20250929
strategy: none
n_ideas_per_run: 5
n_runs: 3
n_topics: 7
rag: False
Overall Averages
Topics: 7
Total ideas: 101
Generation model: gpt-4.1-mini-2025-04-14
Judge model: anthropic/claude-sonnet-4-5-20250929
Runs per topic: 3
Ideas per run: 5
Evaluation protocol: port of… See the full description on the dataset page: https://huggingface.co/datasets/strategy-scope/si_et_al-ideation-gpt41mini-20260414_233911.scientific-ideation-diversity
Scientific Ideation Diversity
Raw data release for the paper "On the Effects of Reasoning Effort and Prompt-Based Diversification on Scientific Ideation Diversity" (Discovery Science 2026).
The paper studies how reasoning effort (low/medium/high) and prompt-based diversification (Verbalized Sampling, String Seed of Thought) shift the diversity of LLM-generated scientific ideas across three frontier models — Claude Sonnet 4.6, GPT-5.4, Gemini 3.1 Pro — evaluated with lexical… See the full description on the dataset page: https://huggingface.co/datasets/tax-free/scientific-ideation-diversity.si_et_al-ideation-gpt41mini-20260414_223535
si_et_al-ideation-gpt41mini-20260414_223535
Benchmark: si_et_al
Generated: 2026-04-14T22:46:55.858516
Parameters
model: gpt-4.1-mini-2025-04-14
judge_model: anthropic/claude-sonnet-4-5-20250929
strategy: none
n_ideas_per_run: 5
n_runs: 3
n_topics: 3
rag: False
Overall Averages
Topics: 3
Total ideas: 41
Generation model: gpt-4.1-mini-2025-04-14
Judge model: anthropic/claude-sonnet-4-5-20250929
Runs per topic: 3
Ideas per run: 5
Evaluation protocol: port of… See the full description on the dataset page: https://huggingface.co/datasets/strategy-scope/si_et_al-ideation-gpt41mini-20260414_223535.si_et_al-ideation-gpt41mini-20260414_230629
si_et_al-ideation-gpt41mini-20260414_230629
Benchmark: si_et_al
Generated: 2026-04-14T23:11:11.121297
Parameters
model: gpt-4.1-mini-2025-04-14
judge_model: anthropic/claude-sonnet-4-5-20250929
strategy: none
n_ideas_per_run: 5
n_runs: 3
n_topics: 3
rag: False
Overall Averages
Topics: 3
Total ideas: 45
Generation model: gpt-4.1-mini-2025-04-14
Judge model: anthropic/claude-sonnet-4-5-20250929
Runs per topic: 3
Ideas per run: 5
Evaluation protocol: port of… See the full description on the dataset page: https://huggingface.co/datasets/strategy-scope/si_et_al-ideation-gpt41mini-20260414_230629.si_et_al-ideation-gpt41mini-20260414_230629-evaluation
si_et_al-ideation-gpt41mini-20260414_230629-evaluation
Aggregate per-topic metrics (+ overall row where topic = __overall__).
Parent: si_et_al-ideation-gpt41mini-20260414_230629
Run Parameters
model: gpt-4.1-mini-2025-04-14
judge_model: anthropic/claude-sonnet-4-5-20250929
strategy: none
n_ideas_per_run: 5
n_runs: 3
n_topics: 3
rag: False
Columns
topic / n_ideas
mean_overall: mean of per-idea 1-10 overall scores (port of… See the full description on the dataset page: https://huggingface.co/datasets/strategy-scope/si_et_al-ideation-gpt41mini-20260414_230629-evaluation.si_et_al-ideation-gpt41mini-20260414_233911-ideas
si_et_al-ideation-gpt41mini-20260414_233911-ideas
Per-idea flat table with LLM-judge scores.
Parent: si_et_al-ideation-gpt41mini-20260414_233911
Columns
topic: NLP topic
idea_index: position within pooled topic idea set
run_idx: which of the n_runs generation calls produced this idea
idea_text: the generated idea (Problem/Existing Methods/Motivation/Proposed Method/Experiment Plan)
overall: 1-10 LLM-judge score (port of ai_researcher/src/idea_direct_score.py)… See the full description on the dataset page: https://huggingface.co/datasets/strategy-scope/si_et_al-ideation-gpt41mini-20260414_233911-ideas.si_et_al-ideation-gpt41mini-full-20260415_003054
si_et_al-ideation-gpt41mini-full-20260415_003054
Benchmark: si_et_al
Generated: 2026-04-15T00:40:56.025185
Parameters
model: gpt-4.1-mini-2025-04-14
judge_model: anthropic/claude-sonnet-4-5-20250929
strategy: full
n_ideas_per_run: 5
n_runs: 3
n_topics: 7
rag: False
Overall Averages
Topics: 7
Total ideas: 103
Generation model: gpt-4.1-mini-2025-04-14
Judge model: anthropic/claude-sonnet-4-5-20250929
Runs per topic: 3
Ideas per run: 5
Evaluation protocol:… See the full description on the dataset page: https://huggingface.co/datasets/strategy-scope/si_et_al-ideation-gpt41mini-full-20260415_003054.si_et_al-ideation-gpt41mini-full-20260415_003054-ideas
si_et_al-ideation-gpt41mini-full-20260415_003054-ideas
Per-idea flat table with LLM-judge scores.
Parent: si_et_al-ideation-gpt41mini-full-20260415_003054
Columns
topic: NLP topic
idea_index: position within pooled topic idea set
run_idx: which of the n_runs generation calls produced this idea
idea_text: the generated idea (Problem/Existing Methods/Motivation/Proposed Method/Experiment Plan)
overall: 1-10 LLM-judge score (port of ai_researcher/src/idea_direct_score.py)… See the full description on the dataset page: https://huggingface.co/datasets/strategy-scope/si_et_al-ideation-gpt41mini-full-20260415_003054-ideas.si_et_al-ideation-gpt41mini-20260414_221654-ideas
si_et_al-ideation-gpt41mini-20260414_221654-ideas
Per-idea flat table with LLM-judge scores.
Parent: si_et_al-ideation-gpt41mini-20260414_221654
Columns
topic: NLP topic
idea_index: position within pooled topic idea set
run_idx: which of the n_runs generation calls produced this idea
idea_text: the generated idea (Problem/Existing Methods/Motivation/Proposed Method/Experiment Plan)
novelty / excitement / feasibility / effectiveness / overall: 1-10 LLM-judge scores
si_et_al-ideation-gpt41mini-20260414_233911-evaluation
si_et_al-ideation-gpt41mini-20260414_233911-evaluation
Aggregate per-topic metrics (+ overall row where topic = __overall__).
Parent: si_et_al-ideation-gpt41mini-20260414_233911
Run Parameters
model: gpt-4.1-mini-2025-04-14
judge_model: anthropic/claude-sonnet-4-5-20250929
strategy: none
n_ideas_per_run: 5
n_runs: 3
n_topics: 7
rag: False
Columns
topic / n_ideas
mean_overall: mean of per-idea 1-10 overall scores (port of… See the full description on the dataset page: https://huggingface.co/datasets/strategy-scope/si_et_al-ideation-gpt41mini-20260414_233911-evaluation.si_et_al-ideation-gpt41mini-full-20260415_003054-evaluation
si_et_al-ideation-gpt41mini-full-20260415_003054-evaluation
Aggregate per-topic metrics (+ overall row where topic = __overall__).
Parent: si_et_al-ideation-gpt41mini-full-20260415_003054
Run Parameters
model: gpt-4.1-mini-2025-04-14
judge_model: anthropic/claude-sonnet-4-5-20250929
strategy: full
n_ideas_per_run: 5
n_runs: 3
n_topics: 7
rag: False
Columns
topic / n_ideas
mean_overall: mean of per-idea 1-10 overall scores (port of… See the full description on the dataset page: https://huggingface.co/datasets/strategy-scope/si_et_al-ideation-gpt41mini-full-20260415_003054-evaluation.si_et_al-ideation-gpt5mini-full-20260415_012526-ideas
si_et_al-ideation-gpt5mini-full-20260415_012526-ideas
Per-idea flat table with LLM-judge scores.
Parent: si_et_al-ideation-gpt5mini-full-20260415_012526
Columns
topic: NLP topic
idea_index: position within pooled topic idea set
run_idx: which of the n_runs generation calls produced this idea
idea_text: the generated idea (Problem/Existing Methods/Motivation/Proposed Method/Experiment Plan)
overall: 1-10 LLM-judge score (port of ai_researcher/src/idea_direct_score.py)… See the full description on the dataset page: https://huggingface.co/datasets/strategy-scope/si_et_al-ideation-gpt5mini-full-20260415_012526-ideas.StreamRank-v2si_et_al-ideation-gpt41mini-20260414_221654
si_et_al-ideation-gpt41mini-20260414_221654
Benchmark: si_et_al
Generated: 2026-04-14T22:21:29.593222
Parameters
model: gpt-4.1-mini-2025-04-14
judge_model: gpt-4o-mini-2024-07-18
strategy: none
n_ideas_per_run: 5
n_runs: 3
n_topics: 3
rag: False
Overall Averages
Topics: 3
Total ideas: 45
Generation model: gpt-4.1-mini-2025-04-14
Judge model: gpt-4o-mini-2024-07-18
Runs per topic: 3
Ideas per run: 5
Metric
Mean
Novelty
7.98
Excitement
8.93… See the full description on the dataset page: https://huggingface.co/datasets/strategy-scope/si_et_al-ideation-gpt41mini-20260414_221654.si_et_al-ideation-gpt5mini-full-20260415_012526
si_et_al-ideation-gpt5mini-full-20260415_012526
Benchmark: si_et_al
Generated: 2026-04-15T01:44:01.632746
Parameters
model: gpt-5-mini
judge_model: anthropic/claude-sonnet-4-5-20250929
strategy: full
n_ideas_per_run: 5
n_runs: 3
n_topics: 7
rag: False
Overall Averages
Topics: 7
Total ideas: 108
Generation model: gpt-5-mini
Judge model: anthropic/claude-sonnet-4-5-20250929
Runs per topic: 3
Ideas per run: 5
Evaluation protocol: port of Si et al.… See the full description on the dataset page: https://huggingface.co/datasets/strategy-scope/si_et_al-ideation-gpt5mini-full-20260415_012526.si_et_al-ideation-gpt5mini-full-20260415_012526-evaluation
si_et_al-ideation-gpt5mini-full-20260415_012526-evaluation
Aggregate per-topic metrics (+ overall row where topic = __overall__).
Parent: si_et_al-ideation-gpt5mini-full-20260415_012526
Run Parameters
model: gpt-5-mini
judge_model: anthropic/claude-sonnet-4-5-20250929
strategy: full
n_ideas_per_run: 5
n_runs: 3
n_topics: 7
rag: False
Columns
topic / n_ideas
mean_overall: mean of per-idea 1-10 overall scores (port of ai_researcher/src/idea_direct_score.py)… See the full description on the dataset page: https://huggingface.co/datasets/strategy-scope/si_et_al-ideation-gpt5mini-full-20260415_012526-evaluation.si_et_al-ideation-gpt5mini-full-20260415_022018
si_et_al-ideation-gpt5mini-full-20260415_022018
Benchmark: si_et_al
Generated: 2026-04-15T02:44:39.671539
Parameters
model: gpt-5-mini
judge_model: anthropic/claude-sonnet-4-5-20250929
strategy: full
n_ideas_per_run: 5
n_runs: 3
n_topics: 7
rag: False
Overall Averages
Topics: 7
Total ideas: 97
Generation model: gpt-5-mini
Judge model: anthropic/claude-sonnet-4-5-20250929
Runs per topic: 3
Ideas per run: 5
Evaluation protocol: port of Si et al.… See the full description on the dataset page: https://huggingface.co/datasets/strategy-scope/si_et_al-ideation-gpt5mini-full-20260415_022018.si_et_al-ideation-gpt5mini-full-20260415_022018-evaluation
si_et_al-ideation-gpt5mini-full-20260415_022018-evaluation
Aggregate per-topic metrics (+ overall row where topic = __overall__).
Parent: si_et_al-ideation-gpt5mini-full-20260415_022018
Run Parameters
model: gpt-5-mini
judge_model: anthropic/claude-sonnet-4-5-20250929
strategy: full
n_ideas_per_run: 5
n_runs: 3
n_topics: 7
rag: False
Columns
topic / n_ideas
mean_overall: mean of per-idea 1-10 overall scores (port of ai_researcher/src/idea_direct_score.py)… See the full description on the dataset page: https://huggingface.co/datasets/strategy-scope/si_et_al-ideation-gpt5mini-full-20260415_022018-evaluation.
