datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
afrolm_active_learning_dataset
AfroLM: A Self-Active Learning-based Multilingual Pretrained Language Model for 23 African Languages
GitHub Repository of the Paper
This repository contains the dataset for our paper AfroLM: A Self-Active Learning-based Multilingual Pretrained Language Model for 23 African Languages which will appear at the third Simple and Efficient Natural Language Processing, at EMNLP 2022.
Our self-active learning framework
Languages Covered
AfroLM has been… See the full description on the dataset page: https://huggingface.co/datasets/bonadossou/afrolm_active_learning_dataset.eval2_all_promptsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/robot-learning-group47/eval2_all_prompts.curriculum_learningVHR-10The VHR-10 dataset mirrored from https://github.com/chaozhong2010/VHR-10_dataset_coco
NWPU VHR-10 data set is a challenging ten-class geospatial object detection data set. This dataset contains a total of 800 VHR optical remote sensing images, where 715 color images were acquired from Google Earth with the spatial resolution ranging from 0.5 to 2 m, and 85 pansharpened color infrared images were acquired from Vaihingen data with a spatial resolution of 0.08 m. The data set is divided into two… See the full description on the dataset page: https://huggingface.co/datasets/satellite-image-deep-learning/VHR-10.Learning
SAGE-3D InteriorGS USDZ: USDZ-Format 3D Gaussian Scenes for Isaac Sim
Paper | Project Page | Code
InteriorGS dataset converted to USDZ format for seamless integration with NVIDIA Omniverse and Isaac Sim platforms.
USDZ format InteriorGS data captured on Issac Sim 5.0.
📢 News
2025-12-15: Released SAGE-3D InteriorGS USDZ dataset with 1,000 converted scenes.
📋 Overview
While the original InteriorGS dataset provides high-quality 3D Gaussian… See the full description on the dataset page: https://huggingface.co/datasets/Zikrihakim66/Learning.SODA-ASODA-A comprises 2513 high-resolution images of aerial scenes, which has 872069 instances annotated with oriented rectangle box annotations over 9 classes.
Website
CLIP-Prompt-learning-sun397eval2_all_prompts_fixed_266This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/robot-learning-group47/eval2_all_prompts_fixed_266.deep_learning_2025_vision_jointThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy",
"total_episodes": 50,
"total_frames": 10128,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Jeongeun/deep_learning_2025_vision_joint.hakka-learning-lesson-imagescmr-learning-lab-assetsWalking-Tours-Semantic
Walking Tours Semantic
Walking Tours Semantic (WT-Sem), introduced in PooDLe, provides semantic segmentation masks for videos in the Walking Tours dataset, as well as three additional videos for validation.
Frames are sampled every 2 seconds from each video and a top-of-the-line semantic segmentation model, OpenSeed, is used to generate the masks.
Specifically, the Swin-L variant of OpenSeed, pretrained on COCO and Objects365 and finetuned on ADE20K, is used.
The 3 new walkaround… See the full description on the dataset page: https://huggingface.co/datasets/agentic-learning-ai-lab/Walking-Tours-Semantic.NFA_OCR_reinforcement_learning_format_TEST5reinforcement-learningin-context-learning-cosmos3-output
Physical-ICL × Cosmos3 — generated outputs
Video-generation outputs from NVIDIA Cosmos3-Nano (Diffusers Cosmos3OmniPipeline,
image-to-video) on the Physical-ICL dataset (Vincwng/Physical-ICL, subset
physiq_prelim, 66 query samples). This studies physical in-context learning: does
showing a demonstration change how the model continues a query scene?
Total generated: 247 videos across 66 query tasks, in 6 configurations.
Configurations
Every configuration uses the… See the full description on the dataset page: https://huggingface.co/datasets/yqi19/in-context-learning-cosmos3-output.NFA_OCR_reinforcement_learning_format_TEST6learning-when-to-look-SFTdatasetmobile-robot-immitation-learningNFA_OCR_reinforcement_learning_format_TEST4unlabeled_samples
Dataset Card for "unlabeled_samples"
More Information needed
learning-xu-ly-video
Nghiên cứu xử lý video: phát hiện chuyển động, benchmark WiseNET và đếm giao thông Hà Nội
Dạ thưa, đây là repository tổng hợp ba giai đoạn nghiên cứu xử lý video của con. Phần đầu nghiên cứu các thuật toán phát hiện vùng chuyển động cơ bản; phần tiếp theo benchmark trên bộ dữ liệu WiseNET; và phần gần đây nhất xây dựng hệ thống phát hiện, theo vết và đếm phương tiện giao thông tại Hà Nội.
1. Tháng 4 — Nghiên cứu nền tảng: ba thuật toán phát hiện chuyển động
Mục… See the full description on the dataset page: https://huggingface.co/datasets/danny2507/learning-xu-ly-video.act_learnings_dataset_350_colored_bkp4Deep_LearningFew-Shot-Class-Incremental-LearningPH_Gr11-Learning_Materialsdeep_learning_2025This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "omy",
"total_episodes": 50,
"total_frames": 10128,
"total_tasks": 1,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 20,
"splits": {
"train": "0:50"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Jeongeun/deep_learning_2025.act_learnings_dataset_350_colored_bkp2motoman-up6-cq-lambda-reinforcement-learning_v1.0
CQ(λ) Bag-Shaking Dataset: Human-in-the-Loop Reinforcement Learning
Dataset Description
This dataset contains synthetic training data comparing standard Q-learning with eligibility traces [Q(λ)] against Cooperative Q-learning [CQ(λ)], a human-in-the-loop reinforcement learning algorithm. The data simulates a robotic "bag-shaking" task where an agent must extract knotted objects from a bag through strategic shaking motions.
Dataset Summary
Task: Bag-shaking… See the full description on the dataset page: https://huggingface.co/datasets/DBbun/motoman-up6-cq-lambda-reinforcement-learning_v1.0.orena-segment-annotations
ORena FOCUS 2026 — SEGMENT supplementary annotations (public half)
Supplementary VQA annotations produced by MLO-Lab for the ORena FOCUS 2026 SEGMENT
track. This is the openly releasable half; the LapChole-FOCUS half is withheld under that
dataset's usage agreement until the organisers publish it.
rows
vqa/heico_derived.jsonl — HeiCo-FOCUS
1,300
vqa/hernia_mesh.jsonl — hernia videos
140
total
1,440
Also included: raw/hernia_mesh_annotations/ (12 frame-level… See the full description on the dataset page: https://huggingface.co/datasets/Machine-Learning-Oncology/orena-segment-annotations.orena-frame-annotations
ORena FOCUS 2026 — FRAME supplementary annotations (public half)
Supplementary VQA annotations produced by MLO-Lab for the ORena FOCUS 2026 FRAME track.
This is the openly releasable half; the LapChole-FOCUS half is withheld under that dataset's
usage agreement until the organisers publish it.
rows
vqa/heico_derived.jsonl — HeiCo-FOCUS
7,608
vqa/hernia_mesh.jsonl — hernia videos
1,350
total
8,958
Also included: raw/hernia_mesh_annotations/ (12 frame-level… See the full description on the dataset page: https://huggingface.co/datasets/Machine-Learning-Oncology/orena-frame-annotations.
