datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
STEVE-1-datasetsteve1-training-data
STEVE-1 Training Data With MineCLIP Embeddings
This dataset contains the MineCLIP-embedded training data used for the MultiSTEVE-1s model zoo. It supports reproducing STEVE-1-style fine-tuning without regenerating MineCLIP embeddings.
Contents
Top-level directories:
dataset_contractor/: OpenAI Contractor Dataset episodes converted for STEVE-1 training.
dataset_mixed_agents/: VPT-generated Minecraft trajectories collected for STEVE-1-style training.
Each episode… See the full description on the dataset page: https://huggingface.co/datasets/randomhuggingfaceuser1273823147/steve1-training-data.Sci-Fi-Books-gutenberg
Gutenberg Sci-Fi Book Dataset
This dataset contains information about science fiction books. It’s designed for training AI models, research, or any other purpose related to natural language processing.
Data Format
The dataset is provided in CSV format. Each record represents a book and includes the following fields:
ID: A unique identifier for the book.
Title: The title of the book.
Author: The author(s) of the book.
Text: The text content of the book (e.g., summary… See the full description on the dataset page: https://huggingface.co/datasets/stevez80/Sci-Fi-Books-gutenberg.fineweb-edu-2013-qwen2-7b
FineWeb-Edu 2013 with Qwen2-7B token counts
Every 2013 FineWeb-Edu document, prepared for continued pretraining, with token
counts computed by a pinned Qwen2-7B tokenizer.
The pipeline is year-agnostic: the year, source revision, tokenizer contract,
and selection rule all come from a config file. 2013 uses
processing_config.json. The 2017 companion dataset, which is large enough to
require shuffling and a token budget rather than retaining everything, is at… See the full description on the dataset page: https://huggingface.co/datasets/stevenyuan666/fineweb-edu-2013-qwen2-7b.chinese_exam_train_dataSteve_Jobs_Interviews
Steve Jobs Interviews Database
Support this project on Ko-fi
Project Overview
This project contains multiple interviews of Steve Jobs during his time before and after Apple.
Goal
The primary goal of this dataset was to fine-tune a language model to output Steve Jobs views and thoughts.
Performance
The performance of this small dataset is very noteworthy. Do to the nature of the database being interview question and answer pairs the replies of the… See the full description on the dataset page: https://huggingface.co/datasets/Hypersniper/Steve_Jobs_Interviews.webnovel-chinese
简介
搜集网络上的网文小说,清洗,分割后,用于训练大语言模型,共计9000本左右,大约9B左右token。
使用
格式说明
采用jsonl格式存储,分为三个字段:
title :小说名称
chapter:章节
text:正文内容
示例:
{"title": "斗破苍穹", "chapter": " 第一章 陨落的天才", "text": "“斗之力,三段!”\n望着测验魔石碑上面闪亮得甚至有些刺眼的五个大字,少年面无表情,唇角有着一抹自嘲,紧握的手掌,因为大力,而导致略微尖锐的指甲深深的刺进了掌心之中,带来一阵阵钻心的疼痛……\n“萧炎,斗之力,三段!级别:低级!”测验魔石碑之旁,一位中年男子,看了一眼碑上所显示出来的信息,语气漠然的将之公布了出来……\n"}
fineweb-edu-2017-qwen2-7b
FineWeb-Edu 2017 (~100B-token subset) with Qwen2-7B token counts
A ~100B-token subset of FineWeb-Edu 2017, prepared for continued pretraining,
with token counts computed by a pinned Qwen2-7B tokenizer.
This dataset is a selected subset, not the complete 2017 crawl year. 2017
contains about 168B Qwen2-7B tokens, above the 100B target, so it was shuffled
and subsetted: data/train/ holds 101,840,059 documents and 100,000,020,347
tokens, which is 59.29% of the 171,755,787 documents… See the full description on the dataset page: https://huggingface.co/datasets/stevenyuan666/fineweb-edu-2017-qwen2-7b.yodas-granary-it-neucodec-10s-20s
Granary Italian VoxPopuli NeuCodec
NeuCodec-tokenized rows for NeuTTS fine-tuning. The source rows were streamed from espnet/yodas-granary / Italian and uploaded as resumable Parquet shards.
{
"source_dataset": "espnet/yodas-granary",
"source_config": "Italian",
"source_splits": [
"ast"
],
"validation_source_split": null,
"codec_checkpoint": "neuphonic/neucodec",
"columns": [
"text",
"codes",
"duration",
"source_dataset",
"source_config"… See the full description on the dataset page: https://huggingface.co/datasets/Steveeeeeeen/yodas-granary-it-neucodec-10s-20s.steve-jobs-speech-corpus
🍏 Steve Jobs Lifetime Keynotes, Speeches & Interviews Corpus (1976–2011)
Historical Eras Distribution
Early Apple Era (1976–1985): 14 keynotes & speeches
NeXT & Pixar Wilderness Era (1985–1996): 16 product launches & oral histories
Apple Renaissance & Mac OS X Era (1997–2006): 51 landmark keynotes & interviews
The Mobile & Cloud Revolution (2007–2011): 20 revolutionary product introductions & final discourses
Key Landmark Ingests
1980 McKenna… See the full description on the dataset page: https://huggingface.co/datasets/sanjeevafk/steve-jobs-speech-corpus.waa_steve_trajectories
Trajectory of STEVE-R1 model on WindowsAgentArena Benchmark
This dataset contains the evaluation trajectories of the computer-use agent STEVE-R1, as described in the paper STEVE-R1: Towards Long Reasoning Computer-use Agents.
It also contains links to the STEVE-R1-7B-SFT model and to the Github repository.
There are 16 zip files in total and each .zip file contains an experiment on 154 tasks from WindowsAgentArena. There are 37K action steps in total.
Description
A… See the full description on the dataset page: https://huggingface.co/datasets/Fanbin/waa_steve_trajectories.sports_15_AUGaction_1_AUGpokemon_card_image_for_authenticity_classification
Pokemon Card Image for Authenticity Classification
This dataset contains front/back images of Pokemon cards for authenticity experiments.
Dataset structure
Images/: all image files (.jpeg)
Images/metadata.jsonl: metadata used by Hugging Face imagefolder
labels.csv: flat label file with the same rows as metadata
Columns
image: image object loaded from file
id: image filename (unique id)
side: card side (0 = front, 1 = back)
labels: authenticity label (1 =… See the full description on the dataset page: https://huggingface.co/datasets/stevelohwc/pokemon_card_image_for_authenticity_classification.naruto-reddit-commentsThis dataset is extracted from Reddit datasets for research purpose
edacc_testedacc_test_cleanmonash_uea_ucr_tser
Dataset Card for Time Series Extrinsic Regression
Dataset Summary
A collection of datasets from Monash, UEA, and UCR supporting research into Time Series Extrinsic Regression (TSER),
a regression task of which the aim is to learn the relationship between a time series and a continuous scalar variable.
This task is closely related to time series classification, where a single categorical variable is learned.
Please read the paper for more.
If you use the results or code… See the full description on the dataset page: https://huggingface.co/datasets/foxy-steve/monash_uea_ucr_tser.adventure_15_AUGspeculative-decoding-bench-rtx4090
Speculative Decoding Benchmark — RTX 4090
TL;DR: 4,576 benchmark runs measuring speculative decoding speedup / acceptance rate
across llama.cpp and LM Studio, Qwen3 (8B/14B) and Llama-3.1-8B target models, on a
single consumer RTX 4090 (24GB). Best observed case: the draft-free ngram-mod
self-speculative mode on structured tasks (JSON extraction 2.81x, code 2.76x,
global-median aggregation at temp=0). Open-ended tasks (creative writing, translation)
with a traditional draft… See the full description on the dataset page: https://huggingface.co/datasets/steven0226/speculative-decoding-bench-rtx4090.open3dsg-repro-2026-07
Open3DSG 复现 + DiffVSGG→3D 研究备份
私有备份。2026-07-31 快照。
内容
Open3DSG-repro/ — Open3DSG (CVPR 2024, arXiv 2402.12259) 复现
FIDELITY.md — 忠实度台账:论文↔官方代码逐项对照、结构性偏离、论文对齐训练命令
ISSUES.html — 问题清单(按 P0/P1/P2 分级,八个板块)
PATCHES.md / PAPER_ALIGNMENT.md — 早期版本,行号已过期,忠实度结论以 FIDELITY.md 为准
SETUP.md / README.md — 部署与数据准备
code_patches/ — 与官方仓库的完整偏差
open3dsg_vs_official_a568358.diff — 对 boschresearch/Open3DSG@a568358 的逐字节 diff(code/ 本身是该仓库的 git clone)… See the full description on the dataset page: https://huggingface.co/datasets/Steven668866/open3dsg-repro-2026-07.maze_15_AUGdrcd-zhtw-extractive-qa-sft
steven0226/drcd-zhtw-extractive-qa-sft
繁體中文抽取式閱讀理解 SFT 資料集,衍生自 DRCD(Delta Reading Comprehension Dataset)。
來源與授權(重要)
原始資料:DRCD(Delta Research Center / 台達電子),
授權 CC BY-SA 3.0,內容改編自繁體中文維基百科。
論文引用:Shao et al., "DRCD: a Chinese Machine Reading Comprehension Dataset", arXiv:1806.00920.
本資料集是 DRCD 的 Adaptation(改編作品),依 CC BY-SA 授權鏈條,以 CC BY-SA 4.0 釋出。
所做的修改
將原始 SQuAD 風格 JSON 重新格式化為 chat SFT 格式(system/user/assistant 三則訊息,assistant 輸出固定 JSON schema)
從… See the full description on the dataset page: https://huggingface.co/datasets/steven0226/drcd-zhtw-extractive-qa-sft.preprocessed_gigaspeech_s_subsetair-quality
Taiwan Air Quality — Derived Daily and Monthly Aggregates
This bundle contains the public aggregate layers from the air-quality
reanalysis: 759,320 station-month rows and 13,589,139
station-day rows across the documented measurands.
Configurations
monthly: station-month mean plus n_days.
daily: station-day mean plus n_valid hours.
A null mean is never filled or interpolated. A zero count means that no
qualifying observations were present; a positive count beside… See the full description on the dataset page: https://huggingface.co/datasets/steven0226/air-quality.yodas-granary-it-neucodec-150k
Granary Italian VoxPopuli NeuCodec
NeuCodec-tokenized rows for NeuTTS fine-tuning. The source rows were streamed from espnet/yodas-granary / Italian and uploaded as resumable Parquet shards.
{
"source_dataset": "espnet/yodas-granary",
"source_config": "Italian",
"source_splits": [
"ast"
],
"validation_source_split": null,
"codec_checkpoint": "neuphonic/neucodec",
"columns": [
"text",
"codes",
"duration",
"source_dataset",
"source_config"… See the full description on the dataset page: https://huggingface.co/datasets/Steveeeeeeen/yodas-granary-it-neucodec-150k.multilingual_evalsmaze_1_AUGyodas-granary-it-neucodec-300k-5s30s
Granary Italian VoxPopuli NeuCodec
NeuCodec-tokenized rows for NeuTTS fine-tuning. The source rows were streamed from espnet/yodas-granary / Italian and uploaded as resumable Parquet shards.
{
"source_dataset": "espnet/yodas-granary",
"source_config": "Italian",
"source_splits": [
"ast",
"asr"
],
"validation_source_split": null,
"codec_checkpoint": "neuphonic/neucodec",
"columns": [
"text",
"codes",
"duration",
"source_dataset"… See the full description on the dataset page: https://huggingface.co/datasets/Steveeeeeeen/yodas-granary-it-neucodec-300k-5s30s.adventure_1_AUG
