nips-submission
Anonymous-nips-submissions
MiraBench: Evaluating Action-Conditioned Reliability in Robotic World Models
This dataset is associated with a NeurIPS 2026 submission. It contains evaluation data for assessing action-conditioned reliability in world models for robotics applications.
Dataset Structure
MiraBench_dataset/
├── action_following_fidelity/ # Action following fidelity tests
├── optimism_bias_detection/ # Optimism bias detection samples
├── physical_consistency/ # Physical… See the full description on the dataset page: https://huggingface.co/datasets/Anonymous-nips-submissions/Anonymous-nips-submissions.mlaire-xquad
MLAIRE-XQUAD
XQuAD reformatted for language-aware retrieval evaluation. Each passage appears once per language; relevance is encoded by group_id matching.
This repository is part of the MLAIRE benchmark, submitted anonymously
to the NeurIPS 2026 Evaluations & Datasets Track. Authors and affiliations
are withheld for double-blind review.
Default top-k
Reported metrics in the paper use top-20.
Layout
corpus/test-*.parquet _id, text, title, language, group_id… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-submission-nips/mlaire-xquad.mlaire-belebele
MLAIRE-BELEBELE
Belebele reformatted for language-aware retrieval evaluation. 488 underlying passages, each available in 122 languages (joined globally by the original link field). Relevance is encoded by group_id matching.
This repository is part of the MLAIRE benchmark, submitted anonymously
to the NeurIPS 2026 Evaluations & Datasets Track. Authors and affiliations
are withheld for double-blind review.
Default top-k
Reported metrics in the paper use top-200.… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-submission-nips/mlaire-belebele.mlaire-mlqa
MLAIRE-MLQA
MLQA reformatted for language-aware retrieval evaluation. Passages are deduplicated at the context level via union-find on the original MLQA ids. Relevance is encoded by group_id matching.
This repository is part of the MLAIRE benchmark, submitted anonymously
to the NeurIPS 2026 Evaluations & Datasets Track. Authors and affiliations
are withheld for double-blind review.
Default top-k
Reported metrics in the paper use top-20.
Layout… See the full description on the dataset page: https://huggingface.co/datasets/anonymous-submission-nips/mlaire-mlqa.NIPS_Submission_929coin-primitive-gr00t-nips-submission
