datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Workspace-Bench
Workspace-Bench
Workspace-Bench is a benchmark for evaluating AI agents on realistic workspace tasks with large-scale file dependencies. It is designed to measure whether an agent can discover, interpret, and use the right files inside a noisy professional workspace, rather than solving tasks from isolated inputs or fully pre-packaged evidence.
The benchmark is introduced in the paper "Workspace-Bench 1.0: Benchmarking AI Agents on Workspace Tasks with Large-Scale File… See the full description on the dataset page: https://huggingface.co/datasets/Workspace-Bench/Workspace-Bench.Workspace-Bench-Lite
Workspace-Bench-Lite
A Lightweight Subset of Workspace-Bench for Fast and Cost-Efficient Evaluation
Overview •
LeaderBoard •
Distribution •
Quick Start •
Changelog •
Citation
Overview
Workspace-Bench-Lite is the Lite split of Workspace-Bench 1.0, designed for fast iteration and lower-cost benchmarking while preserving the core evaluation setting of the full benchmark.
It contains 100 tasks selected from the full Workspace-Bench and is intended to… See the full description on the dataset page: https://huggingface.co/datasets/Workspace-Bench/Workspace-Bench-Lite.kaggle-workspace
BDMCPF
BD motorcycle ride POV footage, collected to experiment with road condition analysis from vehicle POV footage and road condition mapping.
Dataset Description
This dataset contains first-person (POV) video recordings of motorcycle rides in Bangladesh. The footage is intended for research on analyzing road conditions directly from vehicle POV video.
Use Cases
Road condition classification and analysis
Road surface quality mapping… See the full description on the dataset page: https://huggingface.co/datasets/arissassina/kaggle-workspace.
