datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
brittleness-results
Adapters copied (2026-09-08). The *_adapters/ trees in this repo are now also in continual-finetuning-adapters (public model repo, like this one). Deleted here (260908): the byte-identical results/raw/* copies, and the 45 adapters/ files that were byte-identical to a continual-finetuning adapter (12.3 GB); both lists are in MIGRATION_260908.md of any new repo. Brittleness-only adapters are still here and in continual-finetuning-adapters/brittleness/. Please prefer the new repo for loading.… See the full description on the dataset page: https://huggingface.co/datasets/false-facts-finetuning/brittleness-results.tiiuae__Falcon3-7B-Instruct-details
Dataset Card for Evaluation run of tiiuae/Falcon3-7B-Instruct
Dataset automatically created during the evaluation run of model tiiuae/Falcon3-7B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tiiuae__Falcon3-7B-Instruct-details.falsifyrl-source
FalsifyRL Reward-Hacking Falsification
FalsifyRL is a synthetic, executable benchmark for identifying and repairing proxy-reward failures
in embodied multi-agent reinforcement learning.
Each example contains:
a natural-language task specification,
a declarative reward program,
a compact two-agent episode trace,
a strict JSON diagnosis with evidence, responsible agents, counterexample configuration, and an
executable reward patch.
Dataset design
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/KuanKuanKuan/falsifyrl-source.tiiuae__falcon-7b-details
Dataset Card for Evaluation run of tiiuae/falcon-7b
Dataset automatically created during the evaluation run of model tiiuae/falcon-7b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tiiuae__falcon-7b-details.tiiuae__falcon-40b-details
Dataset Card for Evaluation run of tiiuae/falcon-40b
Dataset automatically created during the evaluation run of model tiiuae/falcon-40b
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tiiuae__falcon-40b-details.icml-2026-mds-falsification-trace
Codex trace: Minimum Distance Summaries falsification audit
This public trace documents the focused claim-complete upgrade of
Minimum Distance Summaries for Robust Neural Posterior Estimation.
It covers primary-source inspection, pinned author-code review, real 32×32
HSP90 cryo-EM simulation, a 71,077-parameter Gaussian NPE, 480 adaptations,
the OC-SVM counterexample, two byte-identical executions, static-bundle
verification, and exact-SHA judge discovery.… See the full description on the dataset page: https://huggingface.co/datasets/ProCreations/icml-2026-mds-falsification-trace.fcos_dbss_falsificationneopolita__jessi-v0.5-falcon3-7b-instruct-details
Dataset Card for Evaluation run of neopolita/jessi-v0.5-falcon3-7b-instruct
Dataset automatically created during the evaluation run of model neopolita/jessi-v0.5-falcon3-7b-instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/neopolita__jessi-v0.5-falcon3-7b-instruct-details.tiiuae__falcon-11B-details
Dataset Card for Evaluation run of tiiuae/falcon-11B
Dataset automatically created during the evaluation run of model tiiuae/falcon-11B
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tiiuae__falcon-11B-details.tiiuae__Falcon3-10B-Instruct-details
Dataset Card for Evaluation run of tiiuae/Falcon3-10B-Instruct
Dataset automatically created during the evaluation run of model tiiuae/Falcon3-10B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tiiuae__Falcon3-10B-Instruct-details.Locutusque__CollectiveLM-Falcon-3-7B-details
Dataset Card for Evaluation run of Locutusque/CollectiveLM-Falcon-3-7B
Dataset automatically created during the evaluation run of model Locutusque/CollectiveLM-Falcon-3-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Locutusque__CollectiveLM-Falcon-3-7B-details.tiiuae__Falcon3-3B-Instruct-details
Dataset Card for Evaluation run of tiiuae/Falcon3-3B-Instruct
Dataset automatically created during the evaluation run of model tiiuae/Falcon3-3B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tiiuae__Falcon3-3B-Instruct-details.tiiuae__falcon-40b-instruct-details
Dataset Card for Evaluation run of tiiuae/falcon-40b-instruct
Dataset automatically created during the evaluation run of model tiiuae/falcon-40b-instruct
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tiiuae__falcon-40b-instruct-details.tiiuae__falcon-mamba-7b-details
Dataset Card for Evaluation run of tiiuae/falcon-mamba-7b
Dataset automatically created during the evaluation run of model tiiuae/falcon-mamba-7b
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tiiuae__falcon-mamba-7b-details.tiiuae__Falcon3-10B-Base-details
Dataset Card for Evaluation run of tiiuae/Falcon3-10B-Base
Dataset automatically created during the evaluation run of model tiiuae/Falcon3-10B-Base
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tiiuae__Falcon3-10B-Base-details.tiiuae__Falcon3-1B-Instruct-details
Dataset Card for Evaluation run of tiiuae/Falcon3-1B-Instruct
Dataset automatically created during the evaluation run of model tiiuae/Falcon3-1B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tiiuae__Falcon3-1B-Instruct-details.tiiuae__falcon-7b-instruct-details
Dataset Card for Evaluation run of tiiuae/falcon-7b-instruct
Dataset automatically created during the evaluation run of model tiiuae/falcon-7b-instruct
The dataset is composed of 44 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tiiuae__falcon-7b-instruct-details.human_fall_flat_recordings_01
人类一败涂地 raw recordings
This dataset contains raw game recordings managed by Game Data Platform. Access requests require manual approval.
Game ID: game_8b46e9573d7ed0a5597c311de5d10d53
Collection: general (泛数据)
Recordings: 22
Layout: recordings/<recording_id>/<raw component>
tiiuae__Falcon3-Mamba-7B-Instruct-details
Dataset Card for Evaluation run of tiiuae/Falcon3-Mamba-7B-Instruct
Dataset automatically created during the evaluation run of model tiiuae/Falcon3-Mamba-7B-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tiiuae__Falcon3-Mamba-7B-Instruct-details.tiiuae__Falcon3-Mamba-7B-Base-details
Dataset Card for Evaluation run of tiiuae/Falcon3-Mamba-7B-Base
Dataset automatically created during the evaluation run of model tiiuae/Falcon3-Mamba-7B-Base
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tiiuae__Falcon3-Mamba-7B-Base-details.tensopolis__falcon3-10b-tensopolis-v2-details
Dataset Card for Evaluation run of tensopolis/falcon3-10b-tensopolis-v2
Dataset automatically created during the evaluation run of model tensopolis/falcon3-10b-tensopolis-v2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tensopolis__falcon3-10b-tensopolis-v2-details.tensopolis__falcon3-10b-tensopolis-v1-details
Dataset Card for Evaluation run of tensopolis/falcon3-10b-tensopolis-v1
Dataset automatically created during the evaluation run of model tensopolis/falcon3-10b-tensopolis-v1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 11 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tensopolis__falcon3-10b-tensopolis-v1-details.tiiuae__Falcon3-3B-Base-details
Dataset Card for Evaluation run of tiiuae/Falcon3-3B-Base
Dataset automatically created during the evaluation run of model tiiuae/Falcon3-3B-Base
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tiiuae__Falcon3-3B-Base-details.tiiuae__Falcon3-1B-Base-details
Dataset Card for Evaluation run of tiiuae/Falcon3-1B-Base
Dataset automatically created during the evaluation run of model tiiuae/Falcon3-1B-Base
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tiiuae__Falcon3-1B-Base-details.gol-rl-wiki-fallback-37237020
GoL Wiki Fallback Run 37237020
Wiki-enabled Game of Life collection and GRPO run.
Collection Metrics
{
"line_counts": {
"grpo_training_data": 500,
"rl_trajectories_with_cot": 500
},
"reward": {
"min": -2.9,
"max": 6.3,
"mean": -1.1316,
"positive_count": 128
},
"tool_errors_count": 112,
"positive_tool_errors_count": 0,
"time0_positive_count": 0,
"format_failures": 2,
"empty_thoughts": 0,
"unique_action_signatures": 200… See the full description on the dataset page: https://huggingface.co/datasets/brysgo/gol-rl-wiki-fallback-37237020.tiiuae__Falcon3-7B-Base-details
Dataset Card for Evaluation run of tiiuae/Falcon3-7B-Base
Dataset automatically created during the evaluation run of model tiiuae/Falcon3-7B-Base
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/tiiuae__Falcon3-7B-Base-details.qingy2024__Falcon3-2x10B-MoE-Instruct-details
Dataset Card for Evaluation run of qingy2024/Falcon3-2x10B-MoE-Instruct
Dataset automatically created during the evaluation run of model qingy2024/Falcon3-2x10B-MoE-Instruct
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/qingy2024__Falcon3-2x10B-MoE-Instruct-details.hotmailuser__FalconSlerp3-10B-details
Dataset Card for Evaluation run of hotmailuser/FalconSlerp3-10B
Dataset automatically created during the evaluation run of model hotmailuser/FalconSlerp3-10B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/hotmailuser__FalconSlerp3-10B-details.hotmailuser__FalconSlerp6-7B-details
Dataset Card for Evaluation run of hotmailuser/FalconSlerp6-7B
Dataset automatically created during the evaluation run of model hotmailuser/FalconSlerp6-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/hotmailuser__FalconSlerp6-7B-details.how2sign_rajvi
