datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VlaserClinSeek-Bench
ClinSeek-Bench
ClinSeek-Bench is the evaluation suite introduced in
ClinSeekAgent: Automating Multimodal Evidence Seeking for Agentic Clinical
Reasoning. It evaluates clinical reasoning
under two paired settings with the same task definitions and answer labels:
Curated Input: the model answers from the evidence package provided by
the source benchmark.
Automated Evidence-Seeking: the curated context is removed, and the model
must retrieve evidence from raw clinical data using… See the full description on the dataset page: https://huggingface.co/datasets/UCSC-VLAA/ClinSeek-Bench.thinkflow-vla-features-b2tsiolkovsky-papers
Tsiolkovsky Papers: the complete personal archive as a machine-readable corpus
Machine transcriptions of all 51,008 sheets of fond 555 of the Archive of the
Russian Academy of Sciences — the personal archive of Konstantin Tsiolkovsky
(1857–1935), who derived the rocket equation and described the multistage rocket
decades before anyone could test either.
The archive had been scanned and put online, but without a catalogue you could
query, full-text search, or a dataset. This is… See the full description on the dataset page: https://huggingface.co/datasets/vladimirbesk/tsiolkovsky-papers.robotrace-vla-robustness-traces
RoboTrace Evidence Bundle
This dataset repository contains the public evidence bundle for RoboTrace, a low-cost deployment-stress evaluation scaffold for robot-learning and VLA-style inference pipelines.
The current release evaluates lerobot/pusht and includes reports, metrics, plots, summaries, and release manifests from a complete staged run.
What this bundle is for
Use this repository to inspect evidence from RoboTrace:
action-trace stability metrics
visual… See the full description on the dataset page: https://huggingface.co/datasets/i-am-shaurya05/robotrace-vla-robustness-traces.vlabench_primitive_ft_lerobot_video
VLABench Primitive Tasks — LeRobot v3.0 (TsFile)
Apache TsFile version of VLABench/vlabench_primitive_ft_lerobot_video.
Overview
This dataset is organized in the LeRobot v3.0 format and is used for integrating VLABench into the LeRobot framework officially. Compared with the v2.0 and the RLDS versions, this release stores the visual observations in a video-compressed format rather than as individual image files, giving better storage efficiency and data-loading… See the full description on the dataset page: https://huggingface.co/datasets/THULab/vlabench_primitive_ft_lerobot_video.urban-vla-expert-v1
Urban VLA Expert v1
Urban VLA Expert v1 is a simulator dataset for language-conditioned urban driving. Each frame pairs a 256 x 256 front-camera image with ego state, a natural-language instruction, and continuous driving controls.
This is a small research dataset, not evidence that a policy is ready for a real vehicle. The expert is a deterministic simulator controller, and the language prompts are curated paraphrases rather than speech collected from drivers.
What… See the full description on the dataset page: https://huggingface.co/datasets/Mayank022/urban-vla-expert-v1.opensubs-collocations
OpenSubtitles Collocations
NPMI-scored bigram collocations extracted from the OpenSubtitles parallel corpus. Three languages, three relation types, ~43K bigrams total.
Languages & Corpus Size
Language
Code
Corpus lines
Bigrams
English
en
~100M
15,000
Dutch
nl
~105M
15,000
Serbian
sr
~50M
13,586
Relation Types
ADJ+NOUN — adjective-noun pairs: "slim contract", "kreditan kartica"
VERB+ADP — phrasal verbs / verb-preposition: "come on", "houden… See the full description on the dataset page: https://huggingface.co/datasets/vladvlasov256/opensubs-collocations.ViLReward-73KProcess Reward Data for ViLBench: A Suite for Vision-Language Process Reward Modeling
Paper | Project Page
There are 73K vision-language process reward data sourcing from five training sets.
action-evidence-vla-phase-state-cache
