100k
Datasets
All datasets matching “100k”FastUMI_100k_lerobot
FastUMI-100K: Advancing Data-Driven Robotic Manipulation with a Large-Scale UMI-Style Dataset
[paper] [dataset]
## Overview
FastUMI-100K is a large-scale, high-quality UMI-style dataset designed for data-driven robotic manipulation learning. Featuring over **100K+ demonstration trajectories** across **54 diverse tasks** and hundreds of object types, the dataset provides multi-view wrist-mounted fisheye images and high-frequency end-effector states. To… See the full description on the dataset page: https://huggingface.co/datasets/IPEC-COMMUNITY/FastUMI_100k_lerobot.Egocentric-100K
Egocentric-100K is the largest dataset of manual labor. You can visualize the dataset here.
Egocentric-100K is state-of-the-art in hand visibility and active manipulation density compared to previous in-the-wild egocentric datasets. The complete 30,000 frame evaluation set is available at Egocentric-100K-Evaluation.
Dataset Statistics
Attribute
Value
Total Hours
100,405
Total Frames
10.8 billion
Video Clips
2,010,759
Median Clip Length
180.0 seconds
Mean… See the full description on the dataset page: https://huggingface.co/datasets/builddotai/Egocentric-100K.FineTome-100k
FineTome-100k
The FineTome dataset is a subset of arcee-ai/The-Tome (without arcee-ai/qwen2-72b-magpie-en), re-filtered using HuggingFaceFW/fineweb-edu-classifier.
It was made for my article "Fine-tune Llama 3.1 Ultra-Efficiently with Unsloth".
TimeLens-100K
TimeLens-100K
📑 Paper | 💻 Code | 🏠 Project Page | 🤗 Model & Data
✨ Dataset Description
TimeLens-100K is a large-scale, diverse, and high-quality training dataset for video temporal grounding. It was proposed in our paper TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs and used for training TimeLens models. The annotation process was conducted using an automated pipeline powered by Gemini-2.5-Pro.
📊 Dataset Statistics
Total Videos:… See the full description on the dataset page: https://huggingface.co/datasets/TencentARC/TimeLens-100K.Complete_Data_Source_100K_HOURS
Multi-Language Audio Collection (100K Hours)
This repository is physically reorganized for Absolute 100% Data Visibility.
🏗️ Global Consolidator
Select your language subset to listen to high-quality waveform audio. All shards from legacy and modern pipelines are automatically routed here.
arxiv-ai-ml-100k-papers
license: other
tags:
- arxiv
- ocr
- machine-learning
---
# obswork/arxiv-ai-ml-100k
A 99,999-paper stratified subset of
[`Rendra8631/arxiv-papers`](https://huggingface.co/datasets/Rendra8631/arxiv-papers)
at revision `a2c6afb51332d2744b46308df6917697582f8cd4`, filtered to the primary
subjects `cs.AI`, `cs.CV`, `cs.LG`, and `stat.ML`. Only papers submitted in `2023`-`2025` are included.
This dataset is a build artifact of the OCR… See the full description on the dataset page: https://huggingface.co/datasets/obswork/arxiv-ai-ml-100k-papers.
