datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SciPredict
SciPredict: Can LLMs Predict the Outcomes of Research Experiments?
Paper: SciPredict: Can LLMs Predict the Outcomes of Research Experiments in Natural Sciences?
Overview
SciPredict is a benchmark evaluating whether AI systems can predict experimental outcomes in physics, biology, and chemistry. The dataset comprises 405 questions derived from recently published empirical studies (post-March 2025), spanning 33 subdomains.
Dataset Structure
Total Questions: 405… See the full description on the dataset page: https://huggingface.co/datasets/ScaleAI/SciPredict.2026.RA.Frontier-and-Scale-Cells
Rational-Agent Frontier, Scale, and Framing Cells
This public dataset is a sibling of siddharthmb/2026.RA.Negotiation-Campaigns (the frozen P1-P4 experimental record for the ii_mats/experiments/rational_agents negotiation program) and follows the same conventions: raw per-episode JSON, per-turn oracle annotations, Markdown/HTML transcripts, run manifests, analysis tables, and an integrity manifest over every uploaded file. It packages eight later campaigns that were run against… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.Frontier-and-Scale-Cells.agent-sandbox-negotiation-benchmark
Agent Sandbox Negotiation Benchmark v1
Overview
A dataset of simulated multi-agent negotiations generated using the open-source Agent Sandbox framework.
This dataset captures the final negotiation outcomes, turn depths, strategy alignments, and agreed prices of local LLMs (Llama-3 and Mistral) engaged in intense, adversarial price negotiations at massive scale.
Dataset Statistics
Simulations: 24,122
Strategies: 4 (Balanced, Aggressive, Conservative, Adaptive)… See the full description on the dataset page: https://huggingface.co/datasets/ScareRezume/agent-sandbox-negotiation-benchmark.linical_inference_debt_scanner_v0.1Clinical Inference Debt Scanner
PurposeDetect when a clinical plan relies on stacked assumptions rather than evidence.
You receive:
evidence_signals
a narrative_chain
a planned_action
You output:
inference_debt_level0 to 3
debt_itemthe single most dangerous leap
paydown_stepthe corrective step that restores evidence grounding
Debt scale0 none1 minor2 moderate3 severe
Scoring
debt_level_scoregraded by distance from gold
debt_item_similaritytoken overlap similarity… See the full description on the dataset page: https://huggingface.co/datasets/ClarusC64/linical_inference_debt_scanner_v0.1.scandi-reddit-filtered
Dataset Card for ScandiRedditFiltered
Dataset Summary
ScandiRedditFiltered is manually filtered and post-processed corpus consisting of comments from ScandiReddit.
The intended use of the filtered sentences is for Text-To-Speech (TTS) models.
Supported Tasks and Leaderboards
Training language models is the intended task for this dataset. No leaderboard is active at this point.
Languages
The dataset is available in Danish (da).
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/scandi-reddit-filtered.
