datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
airlens-live
AirLens Live Data
Live data layer for AirLens, an open
air-quality monitoring platform. Updated by scheduled GitHub Actions pipelines.
Layout mirrors the former Supabase Storage buckets:
Path
Content
Cadence
aq-data/current-*-grid.json
Global pollutant grids (PM2.5/PM10/O3/NO2/CO)
hourly
aq-data/timeline/
GEFS-Aerosols PM2.5 frames, -24h..+24h, 3h step
every 3h
aq-data/predictions/grid_latest.json
AOD→PM2.5 model predictions (p10-p90 + DQSS)
every 3h… See the full description on the dataset page: https://huggingface.co/datasets/Robeedau/airlens-live.math
Dataset Card for "livebench/math"
LiveBench is a benchmark for LLMs designed with test set contamination and objective evaluation in mind. It has the following properties:
LiveBench is designed to limit potential contamination by releasing new questions monthly, as well as having questions based on recently-released datasets, arXiv papers, news articles, and IMDb movie synopses.
Each question has verifiable, objective ground-truth answers, allowing hard questions to be scored… See the full description on the dataset page: https://huggingface.co/datasets/livebench/math.model_judgment
Dataset Card for "livebench/model_judgment"
LiveBench is a benchmark for LLMs designed with test set contamination and objective evaluation in mind. It has the following properties:
LiveBench is designed to limit potential contamination by releasing new questions monthly, as well as having questions based on recently-released datasets, arXiv papers, news articles, and IMDb movie synopses.
Each question has verifiable, objective ground-truth answers, allowing hard questions to be… See the full description on the dataset page: https://huggingface.co/datasets/livebench/model_judgment.LiveSports-3K
LiveSports-3K Benchmark
News
[2025.05.12] We released the ASR transcripts for the CC track. See LiveSports-3K-CC.json for details.
Overview
LiveSports‑3K is a comprehensive benchmark for evaluating streaming video understanding capabilities of large language
and multimodal models. It consists of two evaluation tracks:
Closed Captions (CC) Track: Measures models’ ability to generate real‑time commentary aligned with the
ground‑truth ASR transcripts.
Question… See the full description on the dataset page: https://huggingface.co/datasets/stdKonjac/LiveSports-3K.LiveCodeBench-Proexecution-v2LiveBenchResultsLiveHouse-TS-test
LiveHouse-TS
LiveHouse-TS is a prospective benchmark for univariate time-series forecasting.
Models receive only observations that were publicly available at the forecast
cutoff. Their predictions are frozen before the target window begins and scored
after the complete target becomes available.
Website: https://huggingface.co/spaces/Saxon0520/LiveHouse-TS-test ·
Public data: https://huggingface.co/datasets/Saxon0520/LiveHouse-TS-test ·
Source:… See the full description on the dataset page: https://huggingface.co/datasets/Saxon0520/LiveHouse-TS-test.claw-eval-live
Claw-Eval-Live
A live benchmark for workflow agents: 105 controlled tasks with fixtures,
mock services, sandboxed workspaces, task-specific graders, and recorded
execution evidence. The release is a time-stamped snapshot built from public
workflow-demand signals, and the signal-to-task pipeline is designed to be
rerun as demand and models evolve.
This dataset accompanies an anonymous submission to the NeurIPS 2026
Evaluations and Datasets Track.
Quick facts
105… See the full description on the dataset page: https://huggingface.co/datasets/claw-eval-live/claw-eval-live.dreamlake-hand-object
Hand-object 3D reconstruction demo (TACO top-5)
Five egocentric hand-object manipulation clips with full 3D reconstructions,
exported from the public DreamLake staging annotation
yancy/hand-object-recon3d-top5.
Source recordings are from the TACO dataset
(research use). Video, 2D hand keypoints, MANO hand meshes and object poses
all come from the same clip, so the 3D always matches the pixels.
episode
task
duration
measure-ruler-toy__20231102_058
measure a toy with a… See the full description on the dataset page: https://huggingface.co/datasets/live9080/dreamlake-hand-object.LiveStreamingRiskControl
Live or Lie — Live Streaming Room Risk Assessment (May/June 2025)
This dataset contains live-streaming room interaction logs for room-level risk assessment under weak supervision. It is the official dataset for the research presented in the papers:
Outsmarting the Chameleon: Counterfactual Decoupling for Tactical OOD Shifts in Live Streaming Risk Assessment (Hugging Face Papers)
Live or Lie: Action-Aware Capsule Multiple Instance Learning for Risk Assessment in Live Streaming… See the full description on the dataset page: https://huggingface.co/datasets/ByteDance/LiveStreamingRiskControl.clawhub-security-signals-live
ClawHub Security Signals Live
This dataset is the refreshed ClawHub security-signals corpus for scanner testing, prompt regression checks, and operational research against recent public ClawHub skills.
It is a moving dataset, not the fixed paper benchmark. main is expected to change when the ClawHub security dataset snapshot workflow publishes a new sanitized export. Pin a Hugging Face revision or commit when you need reproducibility.
For the frozen research-paper snapshot, use… See the full description on the dataset page: https://huggingface.co/datasets/OpenClaw/clawhub-security-signals-live.droid_100This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "unknown",
"total_episodes": 100,
"total_frames": 32212,
"total_tasks": 47,
"total_videos": 300,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 15,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/live9080/droid_100.so100_tic_tac_toe_we_do_it_liveThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "so100",
"total_episodes": 50,
"total_frames": 29756,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/bnarin/so100_tic_tac_toe_we_do_it_live.robogate-failure-dictionary
RoboGate Failure Dictionary
50,000+ Physics-Validated Pick & Place Failure Patterns across 4 Robots (Franka Panda, UR5e, UR3e, UR10e)
A structured database of robot AI failure patterns collected from NVIDIA Isaac Sim physical simulations using Two-Stage Adaptive Sampling. Each experiment records the exact conditions under which a robot succeeded or failed at Pick & Place tasks.
Quick Stats
Franka Uniform
Franka Boundary
UR5e
UR3e
UR10e
Combined… See the full description on the dataset page: https://huggingface.co/datasets/liveplex/robogate-failure-dictionary.livebench_eval_collections数据来源:https://huggingface.co/livebench
PlayTicker-Live
🏀 PlayTicker Pro: The AI Betting Advisor
A Real-Time Decision Support System (DSS) detecting market inefficiencies in live sports betting.
⚡ Executive Summary
Live betting markets are volatile and inefficient. PlayTicker Pro is an algorithmic "Copilot" designed for the Economics and Entrepreneurship landscape, identifying where the market misprices momentum based on linguistic cues.
🥊 The Edge: Why use an AI Advisor?
Most betting apps act as the… See the full description on the dataset page: https://huggingface.co/datasets/meirnm13/PlayTicker-Live.drug_induced_liver_injury
Dataset Details
Dataset Description
Drug-induced liver injury (DILI) is fatal liver disease caused by drugs
and it has been the single most frequent cause of safety-related drug marketing
withdrawals for the past 50 years (e.g. iproniazid, ticrynafen, benoxaprofen).
This dataset is aggregated from U.S. FDA 2019s National Center for Toxicological
Research.
Curated by:
License: CC BY 4.0
Dataset Sources
corresponding publication
Data source… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/drug_induced_liver_injury.africa-emissions-from-livestock-manure-left-on-pasture-n-content
Emissions from Livestock — Manure left on pasture (N content) | Africa (FAOSTAT) | Africa (Electric Sheep Africa metadata inventory)
Size category: 10K<n<100K - Formats: parquet - Sector: agriculture_food - Engineered by Electric Sheep Africa
TL;DR
This dataset is part of the Electric Sheep Africa catalog on Hugging Face. It is indexed for African data discovery with standardized metadata, loading guidance, provenance notes, and analyst-oriented context.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-emissions-from-livestock-manure-left-on-pasture-n-content.evo_llm_trajectories
What Makes an LLM a Good Optimizer? A Trajectory Analysis of LLM-Guided Evolutionary Search
This dataset contains optimization trajectories for 15 Large Language Models (LLMs) across 8 different optimization tasks, as presented in the paper What Makes an LLM a Good Optimizer? A Trajectory Analysis of LLM-Guided Evolutionary Search.
The data was collected using the LLMEvo framework to study how various LLMs behave when orchestrating evolutionary and agentic optimization systems. The… See the full description on the dataset page: https://huggingface.co/datasets/LivevreXH/evo_llm_trajectories.train_800_sparse__bbox__separate_channel__sim__all_cameras__liveThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_sparse__bbox__separate_channel__sim__all_cameras__live.EU-Retail-UX-Feedback-Live
EU-Retail-UX-Feedback-Live
Real-time, GDPR-anonymised user-experience (UX) feedback collected from the
e-commerce website, mobile app and customer-service platform of a large European
retail company (1000+ employees) operating in the United Kingdom, France and
Germany. The dataset is refreshed every 30 minutes, and every published
version is immutable, checksummed and taggable so it is fully auditable and
rollback-able.
Quick summary
Attribute
Value… See the full description on the dataset page: https://huggingface.co/datasets/toolathon123/EU-Retail-UX-Feedback-Live.LiveGamingBenchmarkfederal-register-live-graphrag-research-20260810
Federal Register live GraphRAG (research)
Local LCR-071 live pipeline output for the 2026-08-10 cutoff (11,784 documents,
CUDA thenlper/gte-small). This Hub copy is a research snapshot.
It is not a current-bundle and does not replace
justicedao/ipfs_federal_register. LCR-084 remains open. Official Federal
Register publications remain the authority.
Hub git directories may contain at most 10,000 files. Document bodies beyond
that cap are stored under corpus/bodies-part2/ rather… See the full description on the dataset page: https://huggingface.co/datasets/justicedao/federal-register-live-graphrag-research-20260810.Benchmark
Description
The document describes the LiveRAG benchmark.
For more details regarding Q&A generation see [1,2].
The LiveRAG benchmark includes 895 questions:
500 questions from Session 1, 500 questions from Session 2, with 105 shared questions from both Sessions
A total of 895 unique questions
Benchmark Fields
Field name
Description
Type
Remarks
Index
Benchmark index
int64 [0,1,...,894]
Question
DataMorgana question
String
Answer
DataMorgana ground… See the full description on the dataset page: https://huggingface.co/datasets/LiveRAG/Benchmark.train_800_sparse__mask__overlay_a75__sim__all_cameras__live__ur5eThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_sparse__mask__overlay_a75__sim__all_cameras__live__ur5e.livestock-health-disease-ssa-synthetic
Dataset Card: Livestock Health & Disease Surveillance (Synthetic Data)
Dataset Summary
This synthetic dataset represents 1,000,000 African smallholder households with livestock systems, capturing livestock health, disease surveillance, veterinary access, and herd management practices across Sub-Saharan Africa. It combines baseline farm characteristics (Dataset 1) with 15 livestock-specific variables to create a comprehensive picture of livestock production systems and… See the full description on the dataset page: https://huggingface.co/datasets/Venesa123/livestock-health-disease-ssa-synthetic.pushtThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "unknown",
"total_episodes": 206,
"total_frames": 25650,
"total_tasks": 1,
"total_videos": 206,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:206"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/live9080/pusht.train_800_dense__bbox__blackout_a50__sim__all_cameras__liveThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_dense__bbox__blackout_a50__sim__all_cameras__live.train_800_complex__mask__blackout_a50__sim__all_cameras__liveThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 20,
"features": {
"observation.state": {
"dtype": "float32",
"shape": [
9
],
"names": [
"x",
"y",
"z",
"qx",
"qy",
"qz",
"qw",
"g1",
"g2"… See the full description on the dataset page: https://huggingface.co/datasets/mim-chess-vlas/train_800_complex__mask__blackout_a50__sim__all_cameras__live.
