datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BixBench
BixBench Dataset
Contains the dataset file BixBench.jsonl and each corresponding data capsule as a .zip file. Capsules are named CapsuleFolder-{uuid}.zip
IMPORTANT UPDATE 2025/09/23:
In ongoing work, we found that many questions in the original dataset had insufficient detail to be answerable, especially in the preferred open-answer setting. To address this, we've extensively re-reviewed and revised a substantial portion of the benchmark. We have also updated the format of the… See the full description on the dataset page: https://huggingface.co/datasets/futurehouse/BixBench.FutureOmni
FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs
Predicting the future requires listening as well as seeing.
📖 Dataset Summary
Although Multimodal Large Language Models (MLLMs) demonstrate strong omni-modal perception, their ability to forecast future events from audio–visual cues remains largely unexplored, as existing benchmarks focus mainly on retrospective understanding.
FutureOmni is the first benchmark designed… See the full description on the dataset page: https://huggingface.co/datasets/OpenMOSS-Team/FutureOmni.DramaBench
DramaBench: Drama Script Continuation Dataset
Dataset Summary
DramaBench is a comprehensive benchmark dataset for evaluating drama script continuation capabilities of large language models.
Current Release: v3.0 Full (1,103 samples) - The complete DramaBench collection is now openly available, with context-continuation pairs designed to assess models across six independent evaluation dimensions.
Release Roadmap
Version
Samples
Status… See the full description on the dataset page: https://huggingface.co/datasets/FutureMa/DramaBench.FutureWorldv2
Offline Forecasting Benchmark (v5)
2067 resolved forecasting questions settling between 2026-04-01 and 2026-08-29.
Every row has a ground truth. Answers are known, so this is an offline benchmark:
a system is given an information cutoff and must not use evidence published after it.
This dataset is updated. Question count, taxonomy and per-class counts change
between releases. Pin the revision you evaluated on and report it beside any score;
a number without a revision is not… See the full description on the dataset page: https://huggingface.co/datasets/weichy2023/FutureWorldv2.Future_Education_Nepal_MBBS_FAQ_Dataset
Future Education Nepal — MBBS FAQ Dataset (Nepali)
Overview
This dataset (future_education_nepal_mbbs_faq_nepali_sharegpt.jsonl) is a collection of 38 question-answering conversation pairs in Nepali, covering frequently asked questions about studying MBBS (medicine) in Nepal — primarily aimed at Indian students considering Nepal as a study destination. Each record is a single-turn human↔gpt exchange in ShareGPT-style format: a Nepali-language question about MBBS… See the full description on the dataset page: https://huggingface.co/datasets/sabin1234/Future_Education_Nepal_MBBS_FAQ_Dataset.bharatbbq-argilla-200VQA-stage3Mobility_Future
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/Matsakitkat/Mobility_Future.MovieStoryGen
MovieStoryGen: Movie-Inspired Creative Writing Dataset
Dataset Description
MovieStoryGen is a high-quality dataset for evaluating and fine-tuning large language models on creative story generation. The dataset contains structured writing prompts paired with detailed story responses that are inspired by IMDb's top 250 movies. Each entry contains a creative writing prompt and a corresponding well-crafted story that reimagines the essence of a classic film in a new context.… See the full description on the dataset page: https://huggingface.co/datasets/FutureMa/MovieStoryGen.ai-energy-futures-2026resultsVQA-stage2-streamingfuture_time_referencesFutureSimtoy_maze_2d_hard_allstep_thinking_future_rollout_cot_500k
ToyMaze2D Hard All-Step Future-Rollout COT
This dataset is generated from the local VisGym ToyMaze2D maze_2d/hard environment.
Rows:
train/: 500,000 gzip-compressed JSONL rows.
test/: 100 gzip-compressed JSONL rows.
Each row is a full trajectory conversation. Every user turn stores a prompt and
one JPEG image item with image_prev, image, and image_next; image_prev == image
is validated for every step, and the final step has image_next == image.
Two non-stop move steps per… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/toy_maze_2d_hard_allstep_thinking_future_rollout_cot_500k.FutureVisiontoy_maze_2d_easy_allstep_thinking_future_rollout_cot_500k
ToyMaze2D Easy All-Step Future-Rollout COT
This dataset is generated from the local VisGym ToyMaze2D maze_2d/easy environment.
Rows:
train/: 500,000 gzip-compressed JSONL rows.
test/: 100 gzip-compressed JSONL rows.
Each row is a full trajectory conversation. Every user turn stores a prompt and
one JPEG image item with image_prev, image, and image_next; image_prev == image
is validated for every step, and the final step has image_next == image.
Two non-stop move steps per… See the full description on the dataset page: https://huggingface.co/datasets/novastar112/toy_maze_2d_easy_allstep_thinking_future_rollout_cot_500k.FutureSimwb-.cheifteen_future_of_work_ai_impact_chat.json1FutureWorktoolsmith-sft-data
ToolSmith SFT data
A unified, de-duplicated function/tool-calling supervised fine-tuning corpus in
the native Qwen3 tool format, used to
train the Future-Labs/ToolSmith-* models.
Format
JSON Lines, one conversation per line:
{
"tools": [{"type": "function", "function": {"name": "...", "description": "...", "parameters": {...}}}],
"messages": [
{"role": "user", "content": "..."},
{"role": "assistant", "content": "", "tool_calls": [{"type": "function"… See the full description on the dataset page: https://huggingface.co/datasets/Future-Labs/toolsmith-sft-data.
