datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
TinyStoriesDataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary.
Described in the following paper: https://arxiv.org/abs/2305.07759.
The models referred to in the paper were trained on TinyStories-train.txt (the file tinystories-valid.txt can be used for validation loss). These models can be found on Huggingface, at roneneldan/TinyStories-1M/3M/8M/28M/33M/1Layer-21M.
Additional resources:
tinystories_all_data.tar.gz - contains a superset of… See the full description on the dataset page: https://huggingface.co/datasets/roneneldan/TinyStories.github-top-code
GitHub Top Developer Source Code
A curated dataset of 1.3M+ source code files from GitHub's top ranked developers (2015-2025).
This dataset is based on the top ranked developers from this dataset: https://huggingface.co/datasets/ronantakizawa/github-top-developers
Dataset Summary
1.3M+ source code files from repositories across ~4,700 unique developers
80+ programming languages included (Python, JavaScript, TypeScript, Rust, Go, C/C++, Java, and more)
Source code only —… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-top-code.github-codereview
Code Review Dataset
A large-scale dataset of the best human-written code reviews from top GitHub repositories.
Each row captures a moment where a human code reviewer left an inline comment on a pull request, and the author subsequently modified the code in response.
The dataset also includes negative examples — code from the same PRs that passed review without comments — to help models learn when code is acceptable.
This provides a natural signal for training models to:
Generate… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-codereview.webui
WebUI
A large-scale dataset pairing real-world UI screenshots with their original HTML, CSS, and JavaScript source code, per-viewport bounding boxes for every visible DOM element, and GPT-4.1 vision descriptions. Every sample is rendered at three responsive breakpoints. Built from public design systems, component libraries, open-source projects, and community code — not synthetically generated.
Overview
Stat
Value
Total rows
36,807
Unique UI samples
12… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/webui.riddles_evolved
Riddles turned into conversations using mistralai/Mistral-7B-Instruct-v0.2
Seeded with Hypersniper's riddles_v1, buy him Ko-fi
Structure: each sample = conversation with two turns: Q/A/Q/A
Process: use Mistral to 1) expand riddles 2) answer riddle 3) formulate human follow-up question 4) answer follow-up question
Code: GitHub
Note: This is an unfiltered dataset, it for sure contains very bad answers.
IN1k256-AR-buckets-bfl16latents_dc-ae-f32c32-sana-1.0_recapPD12M-256px_dc-ae-f32c32-sana-1.0imagenet-1k-validation-subsetsrontgen
Citation
If you use the ROCOv2 dataset in your research, please cite the following paper:
Pelka, O., Menze, B. H., & Rexhausen, S. E. (2023). Radiology Objects in COntext version 2 (ROCOv2): A multimodal dataset for medical image analysis.
arXiv preprint arXiv:2405.10004.
@misc {ronan_l.m._2024,
author = { {Ronan L.M.} },
title = { ROCOv2-radiology (Revision 5d66908) },
year = 2024,
url = {… See the full description on the dataset page: https://huggingface.co/datasets/akahana/rontgen.PutRubbishInBin_500_episodes_frames_3camsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "franka",
"total_episodes": 500,
"total_frames": 77242,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 10,
"splits": {
"train": "0:500"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/RonPlusSign/PutRubbishInBin_500_episodes_frames_3cams.Animated-World-150M-v1-UnifiedLegiSubject-Br-Summaries
🇧🇷 Brazilian Legislative Bills – Summary Dataset
This dataset contains summaries (ementas) of legislative bills proposed in the Brazilian Chamber of Deputies (BCoD) from 1991 to 2022.It is intended for multi-label classification, where each bill may be associated with one or more subject categories (temas).
🔀 This is the summary version of the dataset.If you are looking for the keywords version, see:👉 ronunes/LegiSubject-Br-Keywords
📁 Dataset Structure
The… See the full description on the dataset page: https://huggingface.co/datasets/ronunes/LegiSubject-Br-Summaries.roneneldan-TinyStories-tokenizer-distilgpt22ronec
Dataset Card for RONEC
Dataset Summary
RONEC, at version 2.0, holds 12330 sentences with over 0.5M tokens, annotated with 15 classes, to a total of 80.283 distinctly annotated entities.
The corpus has the following classes and distribution in the train/valid/test splits:
| Classes | Total | Train | | Valid | | Test | |
|------------- |:------: |:------: |:-------: |:------: |:-------: |:------: |:-------: |
| |… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/ronec.roneneldan-TinyStories-tokenizer-gpt2IN1k256-AR-buckets-latents_dc-ae-f32c32-sana-1.0NIH-Chest-X-ray-dataset_resized300pxSth_comnexus-core-data-v2IN1k256-bfl16latents_shape_dc-ae-f32c32-sana-1.0python-code-instructions-japanese
Python Code Instructions - Japanese (18K)
Dataset Description
This dataset contains 18,612 Python programming instruction-response pairs translated to Japanese. It's designed for training language models to understand and generate Python code based on Japanese instructions.
Key Features
18,612 entries covering diverse Python programming tasks
Japanese instructions and prompts for code generation
Original English text preserved for reference
Python code… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/python-code-instructions-japanese.CC12M_IN21K-256px_dc-ae-f32c32-sana-1.0PutRubbishInBin_500_episodes_frames_5camsThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "franka",
"total_episodes": 500,
"total_frames": 77242,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 10,
"splits": {
"train": "0:500"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/RonPlusSign/PutRubbishInBin_500_episodes_frames_5cams.kaggle_llm_science_examHF dataset of Kaggle's LLM Science Exam
CC12M_IN21K-256px-splits_dc-ae-f32c32-sana-1.0lerobot-so101-elevator-datasetThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "so101_follower",
"total_episodes": 51,
"total_frames": 13947,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:51"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/RonLiao/lerobot-so101-elevator-dataset.IN1k256-AR-buckets-bfl16latents_dc-ae-f32c32-sana-1.0MM-RIS
MM-RIS: Multimodal Referring Image Segmentation Dataset
The MM-RIS dataset was introduced in the paper RIS-FUSION: Rethinking Text-Driven Infrared and Visible Image Fusion from the Perspective of Referring Image Segmentation.
This large-scale benchmark supports the multimodal referring image segmentation (RIS) task by providing a goal-aligned approach to supervise and evaluate how effectively natural language contributes to infrared and visible image fusion outcomes.… See the full description on the dataset page: https://huggingface.co/datasets/ronniejiangC/MM-RIS.LegiSubject-Br-Keywords
🇧🇷 Brazilian Legislative Bills – Keyword Dataset
This dataset contains keywords of legislative bills proposed in the Brazilian Chamber of Deputies (BCoD) from 1991 to 2022.It is intended for multi-label classification, where each bill may be associated with one or more subject categories (temas).
🔀 This is the keywords version of the dataset.If you are looking for the summaries version, see:👉 ronunes/LegiSubject-Br-Keywords
📁 Dataset Structure
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/ronunes/LegiSubject-Br-Keywords.Imagenet-256-latents_dc-ae-f32c32-sana-1.0
