datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GMAI-VL-5.5M
GMAI-VL-5.5M Dataset
GMAI-VL-5.5M is a comprehensive, large-scale medical General Medical AI Vision-Language (GMAI-VL) dataset built specifically for training multimodal foundation models in the medical domain. It contains an extraordinary scale of high-quality instructions encompassing over 5.5 million multimodal question-answering pairs, carefully constructed based on hundreds of medical classification, segmentation, and detection datasets.
This repository… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-VL-5.5M.reward-projection-goal-generalisation-vlmseedance_general_all_dance_scm_latent_lmdb
Seedance General-All + Dance SCM Latent LMDB
This dataset stores precomputed SCM latents used for TurboT2AV training.
Source mapping: seedance_general_all_dance_mapping.csv
Successful latent samples: 44,305
Shards: 8 LMDB shards under scm_latent_lmdb/shard_00000 ... shard_00007
Video latent shape per sample: (1, 16, 128, 16, 24)
Audio latent shape per sample: (1, 127, 128)
The source mapping combines Seedance general-all data with a dance subset. The mapping contains 44,504… See the full description on the dataset page: https://huggingface.co/datasets/luyu1021/seedance_general_all_dance_scm_latent_lmdb.GeneralAgentBench
GeneralAgentBench
GeneralAgentBench is a 1,400+ task benchmark for evaluating whether general-purpose AI agents have genuinely completed a task, spanning Mobile / Browser / Desktop environments. It is the evaluation resource accompanying an anonymous NeurIPS 2026 Evaluations & Datasets Track submission (Submission 173, AgentJudge). This release is fully anonymized for double-blind review.
Each task provides a natural-language instruction plus a list of verification
checkpoints.… See the full description on the dataset page: https://huggingface.co/datasets/agentjudge-anon/GeneralAgentBench.repro-flat-minima-and-generalization-insights-from-stochastic-convex-optimization-traces
Agent traces
Agent sessions published from a Trackio Logbook.
Marsouuu__general3Bv2-ECE-PRYMMAL-Martial-details
Dataset Card for Evaluation run of Marsouuu/general3Bv2-ECE-PRYMMAL-Martial
Dataset automatically created during the evaluation run of model Marsouuu/general3Bv2-ECE-PRYMMAL-Martial
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Marsouuu__general3Bv2-ECE-PRYMMAL-Martial-details.wish-engine-toolcall-next-v3-strict-general
wish-engine-toolcall-next-v3-strict-general
Wish-engine implementor next-step tool-calling dataset (v3 strict generalization subset, dynamic aliases).
Splits
train.jsonl: 8081980 bytes
validation.jsonl: 1008543 bytes
test.jsonl: 997791 bytes
Schema
Rows are JSONL with at least:
id
messages (chat format with assistant tool_calls)
tool_name
metadata fields (mode, status, trajectory_*)
Notes
Tool names are dynamically aliased per sample.
A tool… See the full description on the dataset page: https://huggingface.co/datasets/sahilmob/wish-engine-toolcall-next-v3-strict-general.GenerAlign
Dataset Card
GenerAlign is collected to help construct well-aligned LLMs in general domains, such as harmlessness, helpfulness, and honesty. It contains 31398 prompts from existed datasets, including:
FLAN
HH-RLHF
FalseQA
UltraChat
ShareGPT
Similar to UltraFeedback, we complete each prompt with responses from different LLMs, including:
Llama-3.1-Nemotron-70B-Instruct-HF
Llama-3.2-3B-Instructgemma-2-27b-it
All responses are annotated by ArmoRM-Llama3-8B-v0.1.
This dataset has… See the full description on the dataset page: https://huggingface.co/datasets/songff/GenerAlign.general-reasoner-fineweb-filter200exp-primacy-generalization
Experiment E: Primacy Effect Generalization Across LLM Elicitation Formats
Dataset Summary
This dataset tests whether the serial position (primacy) effect found in JSON-formatted LLM elicitation generalizes to other response formats (natural language, Likert, ranking). A methodological contribution applicable to all LLM-as-respondent research. Records 2,400 calls (2,351 valid, 98.0%) across 4 response formats x 8 Latin-square orderings x 5 focal brands x 5 LLM… See the full description on the dataset page: https://huggingface.co/datasets/spectralbranding/exp-primacy-generalization.general_trajectorygeneral-reasoner-fineweb-filter400Dans-Reasoningmaxx-GeneralReasoningMarsouuu__general3B-ECE-PRYMMAL-Martial-details
Dataset Card for Evaluation run of Marsouuu/general3B-ECE-PRYMMAL-Martial
Dataset automatically created during the evaluation run of model Marsouuu/general3B-ECE-PRYMMAL-Martial
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Marsouuu__general3B-ECE-PRYMMAL-Martial-details.wish-engine-toolcall-next-v3-general
wish-engine-toolcall-next-v3-general
Wish-engine implementor next-step tool-calling dataset (v3 generalization, dynamic tool aliases).
Splits
train.jsonl: 19689428 bytes
validation.jsonl: 2711092 bytes
test.jsonl: 2444482 bytes
Schema
Rows are JSONL with at least:
id
messages (chat format with assistant tool_calls)
tool_name
metadata fields (mode, status, trajectory_*)
Notes
Tool names are dynamically aliased per sample.
A tool roster is… See the full description on the dataset page: https://huggingface.co/datasets/sahilmob/wish-engine-toolcall-next-v3-general.task_data_general-math_DeepSeek-R1GeneralMicrobiology
