datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
metrollm-bench
MetroLLM-Bench
The 955 cases of MetroLLM-Bench, a benchmark for language models as the policy layer of a
transit kiosk. The model receives kiosk events, calls structured tools (route planner, fare
calculator, station info, disruption feed, knowledge base) and submits a terminal state: outcome,
fare quote where applicable, kiosk action.
Paper: MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes; HF paper page.
Code, harness and ground-truth generator:… See the full description on the dataset page: https://huggingface.co/datasets/continker/metrollm-bench.quantum-sensing-and-metrology
Neura Parse — Quantum Sensing & Metrology: Fisher Information, the Heisenberg Limit & Entanglement-Enhanced Sensors
A physics- and estimation-theory-deep vertical on the second quantum revolution's sensing pillar: how quantum Fisher information and the Cramer-Rao bound set ultimate precision, how entanglement and squeezing push sensors from the standard quantum limit toward the Heisenberg limit, and how these ideas are realized in optical atomic clocks, NV-center… See the full description on the dataset page: https://huggingface.co/datasets/Neura-parse/quantum-sensing-and-metrology.MetroSurv
MetroSurv-Bench
MetroSurv-Bench is a multimodal benchmark for intelligent traffic surveillance. It is designed to evaluate whether modern MLLMs can move beyond generic video understanding and handle surveillance-specific capabilities such as traffic element perception, dynamic event understanding, temporal grounding, and cross-camera reasoning across related road sections.
The benchmark combines three complementary task settings:
multiple-choice QA on single-view surveillance… See the full description on the dataset page: https://huggingface.co/datasets/anonyuser1/MetroSurv.MetroTalk-Core
MetroTalk Core
MetroTalk Core contains normalized public-transit support dialogues for response-quality research.
Processing
Records retain the original turn order while standardizing speaker labels and whitespace.
Intended use
Use this card for dialogue-response evaluation and controlled prompt experiments.
Verified upstream provenance
Verified upstream dataset: toolathlonEval/MetroTalk-Archive-2023
Release marker:… See the full description on the dataset page: https://huggingface.co/datasets/toolathlonEval/MetroTalk-Core.
