CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01george-adams1 /STEVE-1-datasettext1K<n<10K0 likes2.3k downloads9mo agoHugging Face02randomhuggingfaceuser1273823147 /steve1-training-data STEVE-1 Training Data With MineCLIP Embeddings This dataset contains the MineCLIP-embedded training data used for the MultiSTEVE-1s model zoo. It supports reproducing STEVE-1-style fine-tuning without regenerating MineCLIP embeddings. Contents Top-level directories: dataset_contractor/: OpenAI Contractor Dataset episodes converted for STEVE-1 training. dataset_mixed_agents/: VPT-generated Minecraft trajectories collected for STEVE-1-style training. Each episode… See the full description on the dataset page: https://huggingface.co/datasets/randomhuggingfaceuser1273823147/steve1-training-data.text0 likes698 downloads5mo agoHugging Face03stevez80 /Sci-Fi-Books-gutenberg Gutenberg Sci-Fi Book Dataset This dataset contains information about science fiction books. It’s designed for training AI models, research, or any other purpose related to natural language processing. Data Format The dataset is provided in CSV format. Each record represents a book and includes the following fields: ID: A unique identifier for the book. Title: The title of the book. Author: The author(s) of the book. Text: The text content of the book (e.g., summary… See the full description on the dataset page: https://huggingface.co/datasets/stevez80/Sci-Fi-Books-gutenberg.texttext-generation1K<n<10K12 likes590 downloads3y agoHugging Face04stevenyuan666 /fineweb-edu-2013-qwen2-7b FineWeb-Edu 2013 with Qwen2-7B token counts Every 2013 FineWeb-Edu document, prepared for continued pretraining, with token counts computed by a pinned Qwen2-7B tokenizer. The pipeline is year-agnostic: the year, source revision, tokenizer contract, and selection rule all come from a config file. 2013 uses processing_config.json. The 2017 companion dataset, which is large enough to require shuffling and a token budget rather than retaining everything, is at… See the full description on the dataset page: https://huggingface.co/datasets/stevenyuan666/fineweb-edu-2013-qwen2-7b.tabulartext-generation10M<n<100M0 likes538 downloads13d agoHugging Face05stevenhsu123 /chinese_exam_train_datatext1K<n<10K0 likes531 downloads3y agoHugging Face06Hypersniper /Steve_Jobs_Interviews Steve Jobs Interviews Database Support this project on Ko-fi Project Overview This project contains multiple interviews of Steve Jobs during his time before and after Apple. Goal The primary goal of this dataset was to fine-tune a language model to output Steve Jobs views and thoughts. Performance The performance of this small dataset is very noteworthy. Do to the nature of the database being interview question and answer pairs the replies of the… See the full description on the dataset page: https://huggingface.co/datasets/Hypersniper/Steve_Jobs_Interviews.texttext-generationn<1K6 likes314 downloads3y agoHugging Face07steven-fei /webnovel-chinese 简介 搜集网络上的网文小说,清洗,分割后,用于训练大语言模型,共计9000本左右,大约9B左右token。 使用 格式说明 采用jsonl格式存储,分为三个字段: title :小说名称 chapter:章节 text:正文内容 示例: {"title": "斗破苍穹", "chapter": " 第一章 陨落的天才", "text": "“斗之力,三段!”\n望着测验魔石碑上面闪亮得甚至有些刺眼的五个大字,少年面无表情,唇角有着一抹自嘲,紧握的手掌,因为大力,而导致略微尖锐的指甲深深的刺进了掌心之中,带来一阵阵钻心的疼痛……\n“萧炎,斗之力,三段!级别:低级!”测验魔石碑之旁,一位中年男子,看了一眼碑上所显示出来的信息,语气漠然的将之公布了出来……\n"} texttext-generation1M<n<10M0 likes262 downloads6mo agoHugging Face08stevenyuan666 /fineweb-edu-2017-qwen2-7b FineWeb-Edu 2017 (~100B-token subset) with Qwen2-7B token counts A ~100B-token subset of FineWeb-Edu 2017, prepared for continued pretraining, with token counts computed by a pinned Qwen2-7B tokenizer. This dataset is a selected subset, not the complete 2017 crawl year. 2017 contains about 168B Qwen2-7B tokens, above the 100B target, so it was shuffled and subsetted: data/train/ holds 101,840,059 documents and 100,000,020,347 tokens, which is 59.29% of the 171,755,787 documents… See the full description on the dataset page: https://huggingface.co/datasets/stevenyuan666/fineweb-edu-2017-qwen2-7b.tabulartext-generation100M<n<1B0 likes193 downloads13d agoHugging Face09Steveeeeeeen /yodas-granary-it-neucodec-10s-20s Granary Italian VoxPopuli NeuCodec NeuCodec-tokenized rows for NeuTTS fine-tuning. The source rows were streamed from espnet/yodas-granary / Italian and uploaded as resumable Parquet shards. { "source_dataset": "espnet/yodas-granary", "source_config": "Italian", "source_splits": [ "ast" ], "validation_source_split": null, "codec_checkpoint": "neuphonic/neucodec", "columns": [ "text", "codes", "duration", "source_dataset", "source_config"… See the full description on the dataset page: https://huggingface.co/datasets/Steveeeeeeen/yodas-granary-it-neucodec-10s-20s.tabular100K<n<1M0 likes184 downloads4mo agoHugging Face10sanjeevafk /steve-jobs-speech-corpus 🍏 Steve Jobs Lifetime Keynotes, Speeches & Interviews Corpus (1976–2011) Historical Eras Distribution Early Apple Era (1976–1985): 14 keynotes & speeches NeXT & Pixar Wilderness Era (1985–1996): 16 product launches & oral histories Apple Renaissance & Mac OS X Era (1997–2006): 51 landmark keynotes & interviews The Mobile & Cloud Revolution (2007–2011): 20 revolutionary product introductions & final discourses Key Landmark Ingests 1980 McKenna… See the full description on the dataset page: https://huggingface.co/datasets/sanjeevafk/steve-jobs-speech-corpus.texttext-classificationn<1K0 likes163 downloads27d agoHugging Face11Fanbin /waa_steve_trajectories Trajectory of STEVE-R1 model on WindowsAgentArena Benchmark This dataset contains the evaluation trajectories of the computer-use agent STEVE-R1, as described in the paper STEVE-R1: Towards Long Reasoning Computer-use Agents. It also contains links to the STEVE-R1-7B-SFT model and to the Github repository. There are 16 zip files in total and each .zip file contains an experiment on 154 tasks from WindowsAgentArena. There are 37K action steps in total. Description A… See the full description on the dataset page: https://huggingface.co/datasets/Fanbin/waa_steve_trajectories.imagerobotics1 likes152 downloads2y agoHugging Face12StevenKaseiyo /sports_15_AUGimage10K<n<100K0 likes145 downloads1y agoHugging Face13StevenKaseiyo /action_1_AUGimage10K<n<100K0 likes133 downloads1y agoHugging Face14stevelohwc /pokemon_card_image_for_authenticity_classification Pokemon Card Image for Authenticity Classification This dataset contains front/back images of Pokemon cards for authenticity experiments. Dataset structure Images/: all image files (.jpeg) Images/metadata.jsonl: metadata used by Hugging Face imagefolder labels.csv: flat label file with the same rows as metadata Columns image: image object loaded from file id: image filename (unique id) side: card side (0 = front, 1 = back) labels: authenticity label (1 =… See the full description on the dataset page: https://huggingface.co/datasets/stevelohwc/pokemon_card_image_for_authenticity_classification.imageimage-classificationn<1K0 likes126 downloads7mo agoHugging Face15SteveTran /naruto-reddit-commentsThis dataset is extracted from Reddit datasets for research purpose image100K<n<1M0 likes122 downloads2y agoHugging Face16Steveeeeeeen /edacc_testaudioautomatic-speech-recognition1K<n<10K0 likes117 downloads2y agoHugging Face17Steveeeeeeen /edacc_test_cleanaudioautomatic-speech-recognition1K<n<10K0 likes114 downloads2y agoHugging Face18foxy-steve /monash_uea_ucr_tser Dataset Card for Time Series Extrinsic Regression Dataset Summary A collection of datasets from Monash, UEA, and UCR supporting research into Time Series Extrinsic Regression (TSER), a regression task of which the aim is to learn the relationship between a time series and a continuous scalar variable. This task is closely related to time series classification, where a single categorical variable is learned. Please read the paper for more. If you use the results or code… See the full description on the dataset page: https://huggingface.co/datasets/foxy-steve/monash_uea_ucr_tser.tabulartime-series-forecastingn<1K0 likes112 downloads3y agoHugging Face19StevenKaseiyo /adventure_15_AUGimage10K<n<100K0 likes109 downloads1y agoHugging Face20steven0226 /speculative-decoding-bench-rtx4090 Speculative Decoding Benchmark — RTX 4090 TL;DR: 4,576 benchmark runs measuring speculative decoding speedup / acceptance rate across llama.cpp and LM Studio, Qwen3 (8B/14B) and Llama-3.1-8B target models, on a single consumer RTX 4090 (24GB). Best observed case: the draft-free ngram-mod self-speculative mode on structured tasks (JSON extraction 2.81x, code 2.76x, global-median aggregation at temp=0). Open-ended tasks (creative writing, translation) with a traditional draft… See the full description on the dataset page: https://huggingface.co/datasets/steven0226/speculative-decoding-bench-rtx4090.tabular1K<n<10K3 likes108 downloads2mo agoHugging Face21Steven668866 /open3dsg-repro-2026-07 Open3DSG 复现 + DiffVSGG→3D 研究备份 私有备份。2026-07-31 快照。 内容 Open3DSG-repro/ — Open3DSG (CVPR 2024, arXiv 2402.12259) 复现 FIDELITY.md — 忠实度台账:论文↔官方代码逐项对照、结构性偏离、论文对齐训练命令 ISSUES.html — 问题清单(按 P0/P1/P2 分级,八个板块) PATCHES.md / PAPER_ALIGNMENT.md — 早期版本,行号已过期,忠实度结论以 FIDELITY.md 为准 SETUP.md / README.md — 部署与数据准备 code_patches/ — 与官方仓库的完整偏差 open3dsg_vs_official_a568358.diff — 对 boschresearch/Open3DSG@a568358 的逐字节 diff(code/ 本身是该仓库的 git clone)… See the full description on the dataset page: https://huggingface.co/datasets/Steven668866/open3dsg-repro-2026-07.tabularn<1K0 likes108 downloads2mo agoHugging Face22StevenKaseiyo /maze_15_AUGimage10K<n<100K0 likes105 downloads1y agoHugging Face23steven0226 /drcd-zhtw-extractive-qa-sft steven0226/drcd-zhtw-extractive-qa-sft 繁體中文抽取式閱讀理解 SFT 資料集,衍生自 DRCD(Delta Reading Comprehension Dataset)。 來源與授權(重要) 原始資料:DRCD(Delta Research Center / 台達電子), 授權 CC BY-SA 3.0,內容改編自繁體中文維基百科。 論文引用:Shao et al., "DRCD: a Chinese Machine Reading Comprehension Dataset", arXiv:1806.00920. 本資料集是 DRCD 的 Adaptation(改編作品),依 CC BY-SA 授權鏈條,以 CC BY-SA 4.0 釋出。 所做的修改 將原始 SQuAD 風格 JSON 重新格式化為 chat SFT 格式(system/user/assistant 三則訊息,assistant 輸出固定 JSON schema) 從… See the full description on the dataset page: https://huggingface.co/datasets/steven0226/drcd-zhtw-extractive-qa-sft.textquestion-answering10K<n<100K0 likes105 downloads1mo agoHugging Face24stevenmaschan /preprocessed_gigaspeech_s_subsettext100K<n<1M0 likes97 downloads1y agoHugging Face25steven0226 /air-quality Taiwan Air Quality — Derived Daily and Monthly Aggregates This bundle contains the public aggregate layers from the air-quality reanalysis: 759,320 station-month rows and 13,589,139 station-day rows across the documented measurands. Configurations monthly: station-month mean plus n_days. daily: station-day mean plus n_valid hours. A null mean is never filled or interpolated. A zero count means that no qualifying observations were present; a positive count beside… See the full description on the dataset page: https://huggingface.co/datasets/steven0226/air-quality.tabular10M<n<100M0 likes94 downloads21d agoHugging Face26Steveeeeeeen /yodas-granary-it-neucodec-150k Granary Italian VoxPopuli NeuCodec NeuCodec-tokenized rows for NeuTTS fine-tuning. The source rows were streamed from espnet/yodas-granary / Italian and uploaded as resumable Parquet shards. { "source_dataset": "espnet/yodas-granary", "source_config": "Italian", "source_splits": [ "ast" ], "validation_source_split": null, "codec_checkpoint": "neuphonic/neucodec", "columns": [ "text", "codes", "duration", "source_dataset", "source_config"… See the full description on the dataset page: https://huggingface.co/datasets/Steveeeeeeen/yodas-granary-it-neucodec-150k.tabular100K<n<1M0 likes93 downloads5mo agoHugging Face27Steveeeeeeen /multilingual_evalstabularn<1K0 likes89 downloads4mo agoHugging Face28StevenKaseiyo /maze_1_AUGimage10K<n<100K0 likes88 downloads1y agoHugging Face29Steveeeeeeen /yodas-granary-it-neucodec-300k-5s30s Granary Italian VoxPopuli NeuCodec NeuCodec-tokenized rows for NeuTTS fine-tuning. The source rows were streamed from espnet/yodas-granary / Italian and uploaded as resumable Parquet shards. { "source_dataset": "espnet/yodas-granary", "source_config": "Italian", "source_splits": [ "ast", "asr" ], "validation_source_split": null, "codec_checkpoint": "neuphonic/neucodec", "columns": [ "text", "codes", "duration", "source_dataset"… See the full description on the dataset page: https://huggingface.co/datasets/Steveeeeeeen/yodas-granary-it-neucodec-300k-5s30s.tabular100K<n<1M0 likes82 downloads4mo agoHugging Face30StevenKaseiyo /adventure_1_AUGimage10K<n<100K0 likes80 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.