datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
weibo-dataset
Weibo Dataset
This repository is part of the StevenZhou0825/weibo-dataset personal Weibo archival and research corpus. It contains immutable, uncompressed TAR media shards associated with sent posts, favorites, repost chains, articles, comments, author avatars, and Weibo expressions.
Access and rights
Access requests are reviewed manually. Media remains in the representation downloaded from its source; TAR is only a container and does not grant additional rights.… See the full description on the dataset page: https://huggingface.co/datasets/StevenZhou0825/weibo-dataset.datwaste-garbage-management-dataset
Garbage Classification Dataset
Dataset Summary
This dataset contains images of garbage items categorized into 10 classes, designed for machine learning and computer vision projects focusing on recycling and waste management.
It is ideal for building classification or object detection models, or developing AI-powered solutions for sustainable waste disposal.
Total Images: 19,762
Number of Classes: 10
Class Distribution
Metal: 1020
Glass: 3061… See the full description on the dataset page: https://huggingface.co/datasets/steveharianto/waste-garbage-management-dataset.STEVE-1-datasetctx
ctx
Find the cheapest AI coding setup that actually works on your repo.
CTX Fit analyzes your repository, tests promising AI coding configurations
against real tasks in it, and produces the winning configuration as a
reviewable change — in your working tree with --apply, or as a pull request
with --pr. It picks the cheapest setup that reliably works — reliability
is a requirement, not a tie-break — and if nothing beats what you already have,
it says so.
The winner is chosen… See the full description on the dataset page: https://huggingface.co/datasets/Stevesolun/ctx.safesynth-hard-hat
SafeSynth Hard-Hat Synthetic Data
SafeSynth is a controlled synthetic-data ablation for hard-hat detection. This
release contains two equal-sized COCO annotation sets drawn from the same
14,000-image candidate pool:
Release view
Images
Annotations
Meaning
annotations_filtered.json
3,500
25,278
Images that passed every pre-registered geometry, photometry, and quality rule
annotations_unfiltered.json
3,500
29,998
A deterministic size-matched sample from the full pool… See the full description on the dataset page: https://huggingface.co/datasets/steven0226/safesynth-hard-hat.HEA-3D-maghubble_trailssteve1-training-data
STEVE-1 Training Data With MineCLIP Embeddings
This dataset contains the MineCLIP-embedded training data used for the MultiSTEVE-1s model zoo. It supports reproducing STEVE-1-style fine-tuning without regenerating MineCLIP embeddings.
Contents
Top-level directories:
dataset_contractor/: OpenAI Contractor Dataset episodes converted for STEVE-1 training.
dataset_mixed_agents/: VPT-generated Minecraft trajectories collected for STEVE-1-style training.
Each episode… See the full description on the dataset page: https://huggingface.co/datasets/randomhuggingfaceuser1273823147/steve1-training-data.motionxDAPL-datasetSci-Fi-Books-gutenberg
Gutenberg Sci-Fi Book Dataset
This dataset contains information about science fiction books. It’s designed for training AI models, research, or any other purpose related to natural language processing.
Data Format
The dataset is provided in CSV format. Each record represents a book and includes the following fields:
ID: A unique identifier for the book.
Title: The title of the book.
Author: The author(s) of the book.
Text: The text content of the book (e.g., summary… See the full description on the dataset page: https://huggingface.co/datasets/stevez80/Sci-Fi-Books-gutenberg.chinese_exam_train_datafineweb-edu-2013-qwen2-7b
FineWeb-Edu 2013 with Qwen2-7B token counts
Every 2013 FineWeb-Edu document, prepared for continued pretraining, with token
counts computed by a pinned Qwen2-7B tokenizer.
The pipeline is year-agnostic: the year, source revision, tokenizer contract,
and selection rule all come from a config file. 2013 uses
processing_config.json. The 2017 companion dataset, which is large enough to
require shuffling and a token budget rather than retaining everything, is at… See the full description on the dataset page: https://huggingface.co/datasets/stevenyuan666/fineweb-edu-2013-qwen2-7b.herb-ai-vaultcloudspotting_mars_optimizationchandra_xray_eventfilesSteve_Jobs_Interviews
Steve Jobs Interviews Database
Support this project on Ko-fi
Project Overview
This project contains multiple interviews of Steve Jobs during his time before and after Apple.
Goal
The primary goal of this dataset was to fine-tune a language model to output Steve Jobs views and thoughts.
Performance
The performance of this small dataset is very noteworthy. Do to the nature of the database being interview question and answer pairs the replies of the… See the full description on the dataset page: https://huggingface.co/datasets/Hypersniper/Steve_Jobs_Interviews.webnovel-chinese
简介
搜集网络上的网文小说,清洗,分割后,用于训练大语言模型,共计9000本左右,大约9B左右token。
使用
格式说明
采用jsonl格式存储,分为三个字段:
title :小说名称
chapter:章节
text:正文内容
示例:
{"title": "斗破苍穹", "chapter": " 第一章 陨落的天才", "text": "“斗之力,三段!”\n望着测验魔石碑上面闪亮得甚至有些刺眼的五个大字,少年面无表情,唇角有着一抹自嘲,紧握的手掌,因为大力,而导致略微尖锐的指甲深深的刺进了掌心之中,带来一阵阵钻心的疼痛……\n“萧炎,斗之力,三段!级别:低级!”测验魔石碑之旁,一位中年男子,看了一眼碑上所显示出来的信息,语气漠然的将之公布了出来……\n"}
nus8-datasetwhisper-leaderboard-evalsImageNet-21K_Fall_2011fineweb-edu-2017-qwen2-7b
FineWeb-Edu 2017 (~100B-token subset) with Qwen2-7B token counts
A ~100B-token subset of FineWeb-Edu 2017, prepared for continued pretraining,
with token counts computed by a pinned Qwen2-7B tokenizer.
This dataset is a selected subset, not the complete 2017 crawl year. 2017
contains about 168B Qwen2-7B tokens, above the 100B target, so it was shuffled
and subsetted: data/train/ holds 101,840,059 documents and 100,000,020,347
tokens, which is 59.29% of the 171,755,787 documents… See the full description on the dataset page: https://huggingface.co/datasets/stevenyuan666/fineweb-edu-2017-qwen2-7b.yodas-granary-it-neucodec-10s-20s
Granary Italian VoxPopuli NeuCodec
NeuCodec-tokenized rows for NeuTTS fine-tuning. The source rows were streamed from espnet/yodas-granary / Italian and uploaded as resumable Parquet shards.
{
"source_dataset": "espnet/yodas-granary",
"source_config": "Italian",
"source_splits": [
"ast"
],
"validation_source_split": null,
"codec_checkpoint": "neuphonic/neucodec",
"columns": [
"text",
"codes",
"duration",
"source_dataset",
"source_config"… See the full description on the dataset page: https://huggingface.co/datasets/Steveeeeeeen/yodas-granary-it-neucodec-10s-20s.steven_card_comp_playThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "dexbot_sharpa",
"total_episodes": 5,
"total_frames": 4819,
"total_tasks": 1,
"total_videos": 5,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/YingYuan0414/steven_card_comp_play.steven_nut_compThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "dexbot_sharpa",
"total_episodes": 8,
"total_frames": 14330,
"total_tasks": 1,
"total_videos": 8,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:8"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/YingYuan0414/steven_nut_comp.steve-jobs-speech-corpus
🍏 Steve Jobs Lifetime Keynotes, Speeches & Interviews Corpus (1976–2011)
Historical Eras Distribution
Early Apple Era (1976–1985): 14 keynotes & speeches
NeXT & Pixar Wilderness Era (1985–1996): 16 product launches & oral histories
Apple Renaissance & Mac OS X Era (1997–2006): 51 landmark keynotes & interviews
The Mobile & Cloud Revolution (2007–2011): 20 revolutionary product introductions & final discourses
Key Landmark Ingests
1980 McKenna… See the full description on the dataset page: https://huggingface.co/datasets/sanjeevafk/steve-jobs-speech-corpus.so100_usb_1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so100",
"total_episodes": 15,
"total_frames": 6112,
"total_tasks": 1,
"total_videos": 15,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:15"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/stevenoh2003/so100_usb_1.edacc_testgrabette_pick3_cartesian_480
grabette_pick3_cartesian_480
Raw teleoperation recording from the grabette handheld device: SLAM-tracked end-effector motion plus the two gripper joint angles, as demonstrated. Nothing here is post-processed into a policy-specific representation.
At a glance
episodes
554
frames
91123
fps
50
duration
~30.4 min
camera
observation.images.cam0 at 480x360
codebase_version
v3.0
Channels
action (11D)
channels
meaning… See the full description on the dataset page: https://huggingface.co/datasets/SteveNguyen/grabette_pick3_cartesian_480.
