datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SuperMemory-VQA
SuperMemoryVQA
SuperMemory-VQA is an egocentric visual question answering benchmark for
evaluating long-horizon memory in augmented reality assistant settings. The
dataset is designed around practical questions a person might ask a wearable
memory assistant, such as where an object was left, what someone said earlier,
whether a planned step was completed, or what happened next in a longer event.
The benchmark contains 4,853 human-verified question-answer pairs grounded in
52.9… See the full description on the dataset page: https://huggingface.co/datasets/OSU-AIoT-MLSys-Lab/SuperMemory-VQA.TravelPlanner
TravelPlanner Dataset
TravelPlanner is a benchmark crafted for evaluating language agents in tool-use and complex planning within multiple constraints. (See our paper for more details.)
Introduction
In TravelPlanner, for a given query, language agents are expected to formulate a comprehensive plan that includes transportation, daily meals, attractions, and accommodation for each day.
TravelPlanner comprises 1,225 queries in total. The number of days and hard constraints… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/TravelPlanner.UGround-V1-Data
Updates
[May 1, 2025] Bounding Box Data: We have added bounding box version of Web-Hybrid. For everyone's convenience, no conversation template is applied to this version of data. All the coordinates (x1, y1, x2, y2) are as always normalized to [0,999].
Notes for Requests
If you have applied for access to this dataset but have not received approval, please contact us via email (Boyu Gou) with your name, institution, and research purpose.
Typically, requests will be… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/UGround-V1-Data.osu
osu-dataset-builder Schema
This document describes the parquet file schemas generated by osu-dataset-builder.
Overview
The dataset consists of 14 parquet files organized into logical groups:
Group
Files
Description
Core
beatmaps, hit_objects, timing_points
Main beatmap data
Sliders
slider_control_points, slider_data
Slider curve details
Storyboard
storyboard_elements, storyboard_commands, storyboard_loops, storyboard_triggers
Storyboard animations
Events… See the full description on the dataset page: https://huggingface.co/datasets/lekdan/osu.AgentCL
AgentCL: Evaluation Framework for Continual Learning in Agents
Datasets
agentboard_babyai, agentboard_scienceworld, and mmlu_pro are included as ready-to-use subsets. They are direct subsets from existing public datasets, with no modifications to the original data.
Our constructed codeeval-pro and browsecomp_plus streams are provided.
The metadata in source datasets only indicates the authorship for the source datasets, and does not imply the authorship of this… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/AgentCL.lerobotDISLtestThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 50,
"total_frames": 36435,
"total_tasks": 1,
"total_videos": 50,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/haijunsu-osu/lerobotDISLtest.lerobotDISLtest_2camera_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 50,
"total_frames": 34269,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/haijunsu-osu/lerobotDISLtest_2camera_2.lerobotDISLtest_2camera724This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 50,
"total_frames": 34534,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/haijunsu-osu/lerobotDISLtest_2camera724.lerobotDISLtest_2cameraThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 50,
"total_frames": 36288,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/haijunsu-osu/lerobotDISLtest_2camera.lerobotDISLtest_2camera725This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 100,
"total_frames": 69095,
"total_tasks": 1,
"total_videos": 200,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:100"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/haijunsu-osu/lerobotDISLtest_2camera725.lerobotDISLtest_2camera_blue_bowlThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 50,
"total_frames": 34412,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/haijunsu-osu/lerobotDISLtest_2camera_blue_bowl.record-testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 2,
"total_frames": 3589,
"total_tasks": 1,
"total_videos": 2,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:2"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/haijunsu-osu/record-test.lerobotDISLtest_2camera_one_blue_bowlThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 50,
"total_frames": 21645,
"total_tasks": 1,
"total_videos": 100,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:50"},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/haijunsu-osu/lerobotDISLtest_2camera_one_blue_bowl.osu-beatmap-tags
osu! Beatmap Tags
Community-voted user tags for osu! beatmaps, scraped from the osu! API v2. The dataset covers all four game modes (osu!, osu!taiko, osu!catch, osu!mania) and includes beatmap metadata alongside tag vote counts.
Dataset Contents
The dataset consists of two CSV files:
tags.csv
One row per beatmap. Contains the vote count for each of the 122 user tags. Beatmaps without tags are not included. Tags are sorted from the most common to the least… See the full description on the dataset page: https://huggingface.co/datasets/project-riz/osu-beatmap-tags.UGround-V1-Data-Box
Updates
[May 1, 2025] Bounding Box Data: We have added bounding box version of Web-Hybrid. For everyone's convenience, no conversation template is applied to this version of data. All the coordinates (x1, y1, x2, y2) are as always normalized to [0,999]. The data has also been filtered (757k datapoints after content moderation).
Notes for Requests
If you have applied for access to this dataset but have not received approval, please contact us via email (Boyu Gou)… See the full description on the dataset page: https://huggingface.co/datasets/osunlp/UGround-V1-Data-Box.ur5_pybullet_grasp_largeur5_pickandplace_pybullet_4objectsso101_testso101_testdisl
