datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TinyStoriesDataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary.
Described in the following paper: https://arxiv.org/abs/2305.07759.
The models referred to in the paper were trained on TinyStories-train.txt (the file tinystories-valid.txt can be used for validation loss). These models can be found on Huggingface, at roneneldan/TinyStories-1M/3M/8M/28M/33M/1Layer-21M.
Additional resources:
tinystories_all_data.tar.gz - contains a superset of… See the full description on the dataset page: https://huggingface.co/datasets/roneneldan/TinyStories.3dfront_render_views3dfront-render-viewsbanned-historical-archives
和谐历史档案馆数据集 - Banned Historical Archives Datasets
和谐历史档案馆数据集包含已录入 https://banned-historical-archives.github.io 和暂未未录入的原始文件。
目录结构
banned-historical-archives.github.io # 已录入该网站的原始数据,不定期从 github 仓库中同步
raw # 原始文件
config # 配置文件
todo # 存放暂未录入网站的文件
部分报纸和图片资料存放在单独的仓库:
名称
地址
状态
参考消息
https://huggingface.co/datasets/banned-historical-archives/ckxx
未录入
人民日报
https://huggingface.co/datasets/banned-historical-archives/rmrb
已精选重要的文章录入
文汇报… See the full description on the dataset page: https://huggingface.co/datasets/RonaldoDD/banned-historical-archives.3dfront-render-diffusegithub-top-code
GitHub Top Developer Source Code
A curated dataset of 1.3M+ source code files from GitHub's top ranked developers (2015-2025).
This dataset is based on the top ranked developers from this dataset: https://huggingface.co/datasets/ronantakizawa/github-top-developers
Dataset Summary
1.3M+ source code files from repositories across ~4,700 unique developers
80+ programming languages included (Python, JavaScript, TypeScript, Rust, Go, C/C++, Java, and more)
Source code only —… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-top-code.3dfront_renderaudio_pretraining_sono_syntheticUNICBench\
UNICBench is the first general, cross-modal, multi-level counting performance benchmark designed for
multimodal large models (MLLM). It extends the evaluation scope to three core modalities: image, text,
and audio, aiming to rigorously assess models' numerical perception and logical reasoning abilities in
complex real-world scenarios.github-codereview
Code Review Dataset
A large-scale dataset of the best human-written code reviews from top GitHub repositories.
Each row captures a moment where a human code reviewer left an inline comment on a pull request, and the author subsequently modified the code in response.
The dataset also includes negative examples — code from the same PRs that passed review without comments — to help models learn when code is acceptable.
This provides a natural signal for training models to:
Generate… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-codereview.DeepSeek-v4-Pro-AgentThis dataset was generated using teich by TeichAI
Prepare these datasets for supervised fine-tuning in just a few lines of code — see the Conversion section below.
DeepSeek v4 Pro Agent Traces
This directory contains raw agent trace files generated by teich.
All assistant responses were generated by deepseek/deepseek-v4-pro.
JSONL files: 4006
Training-ready tools
A complete configured tools schema snapshot is embedded in the collapsed section at the bottom of… See the full description on the dataset page: https://huggingface.co/datasets/ronaldcmz/DeepSeek-v4-Pro-Agent.webui
WebUI
A large-scale dataset pairing real-world UI screenshots with their original HTML, CSS, and JavaScript source code, per-viewport bounding boxes for every visible DOM element, and GPT-4.1 vision descriptions. Every sample is rendered at three responsive breakpoints. Built from public design systems, component libraries, open-source projects, and community code — not synthetically generated.
Overview
Stat
Value
Total rows
36,807
Unique UI samples
12… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/webui.rlwrld_ICLRTinyStoriesInstructwan2.2lriddles_evolved
Riddles turned into conversations using mistralai/Mistral-7B-Instruct-v0.2
Seeded with Hypersniper's riddles_v1, buy him Ko-fi
Structure: each sample = conversation with two turns: Q/A/Q/A
Process: use Mistral to 1) expand riddles 2) answer riddle 3) formulate human follow-up question 4) answer follow-up question
Code: GitHub
Note: This is an unfiltered dataset, it for sure contains very bad answers.
mv-mesh-40kwhenho3d_fullmedical_fine_tuning_12MIN1k256-AR-buckets-bfl16latents_dc-ae-f32c32-sana-1.0_recapPD12M-256px_dc-ae-f32c32-sana-1.0soccer-dialoguesrontgen
Citation
If you use the ROCOv2 dataset in your research, please cite the following paper:
Pelka, O., Menze, B. H., & Rexhausen, S. E. (2023). Radiology Objects in COntext version 2 (ROCOv2): A multimodal dataset for medical image analysis.
arXiv preprint arXiv:2405.10004.
@misc {ronan_l.m._2024,
author = { {Ronan L.M.} },
title = { ROCOv2-radiology (Revision 5d66908) },
year = 2024,
url = {… See the full description on the dataset page: https://huggingface.co/datasets/akahana/rontgen.pose6daug
pose6daug
Real-world Franka manipulation episodes with object-swap and action augmentation
artifacts. 120 training episodes over 4 objects (blue_cup, green_pear, kanu,
white_spray), dual ZED cameras (exo static + ego wrist-mounted).
Layout
Per-frame PNGs are packed into uncompressed tars per episode — the dataset has
~427k mask/plate frames and loose files hit Hugging Face's per-repo file
recommendation and API rate limits hard.
data/<object>/<NNNN>/
masks.tar… See the full description on the dataset page: https://huggingface.co/datasets/Ronaldo-GOAT/pose6daug.ai-price-index
AI Price Index
An open, dated, first-party-sourced record of AI model API prices over time.
Provider pricing pages change quietly with no changelog. This dataset is the changelog. Every price carries the date it became valid and the date it was last verified against the official source, so you can price historical token usage point-in-time instead of extrapolating from today's rate.
491 price records, 122 models, 11 providers: Anthropic, OpenAI, Google, Mistral, xAI, DeepSeek… See the full description on the dataset page: https://huggingface.co/datasets/RoninForge/ai-price-index.GRASPimagenet-1k-validation-subsetsho3d_transfer
HO-3D v3 (merged transfer)
This repository bundles the HO-3D v3 dataset together with its rendered
hand/object segmentations, kept as the original .zip archives plus the
upstream README.txt.
Contents
File
Size
Description
HO3D_v3.zip
~32 GB
Main dataset: train/, evaluation/, calibration/, manual_annotations/, train.txt, evaluation.txt
HO3D_v3_segmentations_rendered.zip
~91 MB
Rendered hand/object segmentations: train/<SEQ>/seg/*.png (1/4 RGB… See the full description on the dataset page: https://huggingface.co/datasets/Ronaldo-GOAT/ho3d_transfer.sampleflip-kits
