datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
leaague-of-legends-decoded-replay-packets-s12-unorganized
Disclaimer
This work isn’t endorsed by Riot Games and doesn’t reflect the views or opinions of Riot Games or anyone officially involved in producing or managing League of Legends. League of Legends and Riot Games are trademarks or registered trademarks of Riot Games, Inc.
Citation
If you use this dataset in your research, please cite:
@dataset{league_of_legends_decoded_replay_packets_2025,
title={League of Legends Decoded Replay Packets Dataset},
author={maknee}… See the full description on the dataset page: https://huggingface.co/datasets/maknee/leaague-of-legends-decoded-replay-packets-s12-unorganized.psg-audio-v3-unofficial-mirror
PSG-Audio v3 — Unofficial Complete Mirror
Unofficial complete mirror of the publicly released PSG-Audio Version 3 dataset.
This repository preserves the original files without modification and provides a reliable, high-speed mirror through the Hugging Face Hub for the research community.
Overview
PSG-Audio v3 is one of the largest publicly available multimodal sleep datasets, combining overnight clinical polysomnography (PSG) with synchronized environmental… See the full description on the dataset page: https://huggingface.co/datasets/Samuelsantos777/psg-audio-v3-unofficial-mirror.UnorthoDOS
UnorthoDOS - Unorthorectified Dataset for On board Satellite methane detection
This repository contains two ML-ready datasets, comprising orthorectified and unorthorectified hyperspectral images from the Earth Surface Mineral Dust Source Investigation (EMIT) sensor, for methane detection. This dataset supports the research presented in the paper Towards Methane Detection Onboard Satellites.
Paper: Towards Methane Detection Onboard Satellites
Code:… See the full description on the dataset page: https://huggingface.co/datasets/SpaceML/UnorthoDOS.UNO-Bench UNO-Bench: A Unified Benchmark for Exploring the Compositional Law Between Uni-modal and Omni-modal in Omni Models
🔔News
🔥[2025/12/04] We have released the evaluation scripts uno-eval, a unified evaluation framework for omni-modal benchmarks. More benchmarks will be supported in the future.
🔥[2025/12/04] We have released the scoring model UNO-Scorer-Qwen3-14B. Feel free to use it!
👀 UNO-Bench Overview
Multimodal Large Languages models have… See the full description on the dataset page: https://huggingface.co/datasets/meituan-longcat/UNO-Bench.G1edu-u3_pullBowl_storage_bread_unordered_a
G1edu-u3_pullBowl_storage_bread_unordered_a
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: Unitree_G1_Dex3_phecda
| Codebase Version: v2.1
End-Effector Type: five_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset Statistics
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/G1edu-u3_pullBowl_storage_bread_unordered_a.G1edu-u3_pullBowl_storage_bread_unordered_b
G1edu-u3_pullBowl_storage_bread_unordered_b
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: Unitree_G1_Dex3_phecda
| Codebase Version: v2.1
End-Effector Type: five_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
kitchen
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
receive
place
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/G1edu-u3_pullBowl_storage_bread_unordered_b.G1edu-u3_pullBowl_storage_bread_unordered_C
G1edu-u3_pullBowl_storage_bread_unordered_C
📋 Overview
This dataset uses an extended format based on LeRobot and is fully compatible with LeRobot.
Robot Type: Unitree_G1_Dex3_phecda
| Codebase Version: v2.1
End-Effector Type: five_finger_hand
🏠 Scene Types
This dataset covers the following scene types:
home
kitchen
🤖 Atomic Actions
This dataset includes the following atomic actions:
grasp
pick
place
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/RoboCOIN/G1edu-u3_pullBowl_storage_bread_unordered_C.UnoBench
UnoBench
UnoBench is a benchmark for target-centric obstruction reasoning in robotic grasping under cluttered scenes. Given a target object, a method must identify the objects that block or constrain access to that target before grasping.
UnoBench is built upon MetaGraspNetV2 and extends the initial idea of FreeGraspData.
Resources
Resource
Link
Description
UnoGrasp code
GitHub main branch
Method code, checkpoints, inference, and evaluation.
Challenge… See the full description on the dataset page: https://huggingface.co/datasets/FBK-TeV/UnoBench.uno-deck
Uno Deck
This dataset is part of the Roboflow 100 benchmark, a diverse collection of 100 object detection datasets spanning 7 imagery domains.
Dataset Statistics
Split
Images
Train
6,295
Validation
1,798
Test
899
Total
8,992
Classes (15)
0
1
10
11
12
13
14
2
3
4
5
6
7
8
9
Usage
With LibreYOLO
from libreyolo import LIBREYOLO
# Load a model
model = LIBREYOLO(model_path="libreyoloXnano.pt")
# Train on this dataset… See the full description on the dataset page: https://huggingface.co/datasets/LibreYOLO/uno-deck.UNO-1M
Less-to-More Generalization: Unlocking More Controllability by In-Context Generation
Overview
UNO-1M is a large dataset (~1M paired images) constructed by the in-context generation pipeline introduced in the UNO paper. Its advantages include highly diverse categories (>365 categories), high-resolution images (around 1024x1024), variable resolutions (different aspect ratios), high quality (produced by state-of-the-art text-to-image models), and high subject… See the full description on the dataset page: https://huggingface.co/datasets/bytedance-research/UNO-1M.movie-scenes-captionedunofficial-pyedu
About This Dataset
The HuggingFaceTB team has released an impressive series of models called smollm (V1/V2) (paper: https://arxiv.org/abs/2502.02737).
According to their documentation, they used Stack-Edu as the code field corpus for pretraining and published https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus.
However for some reason, only a Python-Edu subset is accessible and there's no content/text field in it.
The full dataset is stored on AWS S3; downloading it… See the full description on the dataset page: https://huggingface.co/datasets/Leon-Leee/unofficial-pyedu.UNO1m-filtered-splitArabic-news-daily
Arabic News Daily 🗞️
A daily-updated, multi-domain Arabic news dataset collected automatically from 15 curated sources.
Unlike other Arabic datasets that are static snapshots, this dataset grows every day — making it ideal for research requiring fresh, current Arabic text across diverse domains.
Sources
Source
Domain
Variety
Al Jazeera Arabic
Politics
MSA
BBC Arabic
Politics
MSA
RT Arabic
Politics
MSA
Al Arabiya
Politics
MSA
AITNews
Tech & AI… See the full description on the dataset page: https://huggingface.co/datasets/unohamza/Arabic-news-daily.movie-scenesBG-20kSAM-LLAVA-55kUNOPose_dataThis repository contains the data presented in UNOPose: Unseen Object Pose Estimation with an Unposed RGB-D Reference Image.
Code: https://github.com/shanice-l/UNOPose
so101_pickplace_unoThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 7255,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ankithreddy/so101_pickplace_uno.synth-bgremove-v4-genbg-dssynth-bg-remove-v4-genbgSAM-LLAVA-20ksynth-bg-removed-v1-denoisedsynth-bg-remove-v14-smallstocks-UNOMINDA-1D-candlestbench-easy-unofficial-codex-gpt5-trialsDoc: https://docs.google.com/document/d/1qJ2DB7EvUCc-_cDReeem_kL0z4QswkbEclow_zH1e3E/edit?tab=t.7od5usjmry7e
Script: https://github.com/mlfoundations/dc-agent/pull/73
synth-bg-remove-v18-512pxUno-Curriculum
Uno-Curriculum
Training corpus for a hierarchical-delegation router: a small language
model that decomposes a task into subtasks and routes each subtask to a
(worker model, skill) pair.
Every row comes from a real public HuggingFace dataset — the
question and gold_answer are sampled verbatim from the dataset
identified by the source field. Every row then goes through the
same three-stage pipeline (router probe → teacher trajectory →
noise removal) to obtain the multi-turn trajectory… See the full description on the dataset page: https://huggingface.co/datasets/tinaxie/Uno-Curriculum.synth-bg-remove-v9stock-images-bg-removed-10k-v4
