datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WorldVQA
WorldVQA
WorldVQA: Measuring Atomic World Knowledge in Multimodal Large Language Models
HomePage |
Dataset |
Paper |
Code
Abstract
We introduce WorldVQA, a benchmark designed to evaluate the atomic vision-centric world knowledge of Multimodal Large Language Models (MLLMs). Current evaluations often conflate visual knowledge retrieval with reasoning. In contrast, WorldVQA decouples these capabilities to strictly measure "what the model… See the full description on the dataset page: https://huggingface.co/datasets/moonshotai/WorldVQA.TallyQA-VLMEvalKitmoonlighter-2-wiki-data
Moonlighter 2 relic prices and weapons dataset
An open, versioned export from www.moonlighter2.wiki, an independent and unofficial Moonlighter 2 wiki published as The Endless Ledger.
This release contains two small, research-friendly tables:
Relic prices: 159 published relics, including 154 rows with sourced prices and 5 rows whose unknown prices are intentionally left blank.
Weapons: 30 published weapons with weapon type, effects, upgrade information, verification state and… See the full description on the dataset page: https://huggingface.co/datasets/zyzw10086/moonlighter-2-wiki-data.constitution_of_indiaklac_legal_aid_counseling
Dataset Description
법률구조공단의 법률구조상담 웹페이지를 크롤링하여 구축한 데이터셋 입니다.
CountBenchQA-VLMEvalKitatcosim-speaker-disjoint-splits
ATCOSIM speaker-disjoint splits (metadata only)
This dataset contains no audio and no transcripts. It is a split definition:
one row per ATCOSIM utterance, giving its speaker, its recording session, its
duration, and which half of a speaker-disjoint evaluation it belongs to.
The audio and transcriptions are not here because they cannot be redistributed.
The ATCOSIM corpus
manual §5.2 states that the corpus is "provided free of charge" and "permitted
to use ... for research and… See the full description on the dataset page: https://huggingface.co/datasets/moonshine-ai/atcosim-speaker-disjoint-splits.MoonRide-LLM-Index-v7
MoonRide LLM Index v7
Results of testing wide range of LLMs against my private benchmark (v7). More information in this short article.
MotoRisk-CARLA
MotoRisk-CARLA
MotoRisk-CARLA is a CARLA-based motorcycle riding dataset for accident detection and unsafe riding behavior research. It contains labeled session CSV files for normal riding, reckless/aggressive riding, zigzag/weaving, and accident scenarios collected with simulated motorcycle sensor data.
The dataset is organized into four folders: Accident, Normal_Riding, Reckless_Aggressive, and Zigzag_Weaving.
uwb-atcc-session-disjoint-splits
UWB-ATCC session-disjoint splits
IDs only, no audio. Train drops the one session shared with the published test split (TWR-34720N) so an in-domain number is domain adaptation, not session leakage. Audio stays on Jzuluaga/uwb_atcc (CC BY-NC-SA 4.0).
startups_nbv1MoonlightShiftPatterns
MoonlightShiftPatterns
tags: predictive, employment, time-series
Note: This is an AI-generated dataset so its content may be inaccurate or false
Dataset Description:
The 'MoonlightShiftPatterns' dataset captures the patterns of individuals engaged in moonlighting jobs across different industries and regions. It includes time-series data of their employment periods and activities. The dataset is structured to facilitate predictive analysis on the duration and frequency of… See the full description on the dataset page: https://huggingface.co/datasets/infinite-dataset-hub/MoonlightShiftPatterns.moonlit-datacomplie
