datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
metrollm-bench
MetroLLM-Bench
The 955 cases of MetroLLM-Bench, a benchmark for language models as the policy layer of a
transit kiosk. The model receives kiosk events, calls structured tools (route planner, fare
calculator, station info, disruption feed, knowledge base) and submits a terminal state: outcome,
fare quote where applicable, kiosk action.
Paper: MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes; HF paper page.
Code, harness and ground-truth generator:… See the full description on the dataset page: https://huggingface.co/datasets/continker/metrollm-bench.legal-metrology-accuracy-classes
Accuracy classes and maximum permissible errors for trade measuring instruments (EU MID)
Canonical, always-current version: https://referencesource.org/legal-metrology-accuracy-classes/
Machine-readable: https://referencesource.org/legal-metrology-accuracy-classes/data.json — this mirror is a point-in-time copy.
Last verified: 2026-08-05
Stale after: 2027-08-05 (past this date, prefer the canonical copy —
it re-verifies on a cadence this snapshot does not)
Records: 70
Accuracy… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/legal-metrology-accuracy-classes.MetroPulse-TLC-Validation-Evidence
MetroPulse TLC Validation Evidence
This repository contains reproducibility and validation metadata for the
MetroPulse NYC Data Reasoning Lab.
The underlying real dataset is the NYC TLC Yellow Taxi Trip Records.
Raw TLC trip data is not redistributed here.
D1 validation run
Run ID: RUN-20260913-TLC-SMOKE-001
Dataset: NYC TLC Yellow TaxiPartition: 2022-01Tier: Smoke validation
Source evidence
Input rows: 2,463,931
Source bytes: 38,139,949
SHA-256:… See the full description on the dataset page: https://huggingface.co/datasets/sanjay042211/MetroPulse-TLC-Validation-Evidence.metro_stores
🏬 METRO France – Réseau de magasins (dataset)
Description
Ce dataset recense l’ensemble des magasins METRO en France, avec leurs informations
de localisation, de contact et d’horaires d’ouverture.
Il est destiné à des usages de :
cartographie et géolocalisation,
analyse territoriale et commerciale,
enrichissement de bases de données,
systèmes RAG et applications IA,
projets open data.
Les données sont structurées et disponibles en CSV et JSON optimisé pour les LLMs.… See the full description on the dataset page: https://huggingface.co/datasets/METRO-France/metro_stores.MetroSurv
MetroSurv-Bench
MetroSurv-Bench is a multimodal benchmark for intelligent traffic surveillance. It is designed to evaluate whether modern MLLMs can move beyond generic video understanding and handle surveillance-specific capabilities such as traffic element perception, dynamic event understanding, temporal grounding, and cross-camera reasoning across related road sections.
The benchmark combines three complementary task settings:
multiple-choice QA on single-view surveillance… See the full description on the dataset page: https://huggingface.co/datasets/anonyuser1/MetroSurv.fha-denial-rates-by-metro
FHA Denial Rates by U.S. Metro Area — 2025 (320 MSAs)
To our knowledge, the first free, FHA-specific metro-level denial-rate dataset: all 320 metropolitan areas/divisions with ≥500 decisioned FHA applications in the complete 2025 CFPB HMDA record, with small-vs-large loan splits.
Headline figures: among metros with 1,000+ applications, denial rates ranged from 8.9% (Amarillo, TX) to 31.8% (Bridgeport-Stamford-Norwalk, CT). The small-loan penalty peaks at 3.61× in El Paso, TX… See the full description on the dataset page: https://huggingface.co/datasets/FinanceRateCalc/fha-denial-rates-by-metro.techwriterstyleguidestalker-metro-text-classification
databricks-custom-datasettw_sg_rulesmetrolist-updatesmetro1
