CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01roneneldan /TinyStoriesDataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary. Described in the following paper: https://arxiv.org/abs/2305.07759. The models referred to in the paper were trained on TinyStories-train.txt (the file tinystories-valid.txt can be used for validation loss). These models can be found on Huggingface, at roneneldan/TinyStories-1M/3M/8M/28M/33M/1Layer-21M. Additional resources: tinystories_all_data.tar.gz - contains a superset of… See the full description on the dataset page: https://huggingface.co/datasets/roneneldan/TinyStories.texttext-generation1M<n<10M1.2k likes95k downloads2y agoHugging Face02ronantakizawa /github-top-code GitHub Top Developer Source Code A curated dataset of 1.3M+ source code files from GitHub's top ranked developers (2015-2025). This dataset is based on the top ranked developers from this dataset: https://huggingface.co/datasets/ronantakizawa/github-top-developers Dataset Summary 1.3M+ source code files from repositories across ~4,700 unique developers 80+ programming languages included (Python, JavaScript, TypeScript, Rust, Go, C/C++, Java, and more) Source code only —… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-top-code.texttext-generation1M<n<10M125 likes3.7k downloads7mo agoHugging Face03ronantakizawa /github-codereview Code Review Dataset A large-scale dataset of the best human-written code reviews from top GitHub repositories. Each row captures a moment where a human code reviewer left an inline comment on a pull request, and the author subsequently modified the code in response. The dataset also includes negative examples — code from the same PRs that passed review without comments — to help models learn when code is acceptable. This provides a natural signal for training models to: Generate… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/github-codereview.tabulartext-generation100K<n<1M62 likes1.4k downloads7mo agoHugging Face04ronantakizawa /webui WebUI A large-scale dataset pairing real-world UI screenshots with their original HTML, CSS, and JavaScript source code, per-viewport bounding boxes for every visible DOM element, and GPT-4.1 vision descriptions. Every sample is rendered at three responsive breakpoints. Built from public design systems, component libraries, open-source projects, and community code — not synthetically generated. Overview Stat Value Total rows 36,807 Unique UI samples 12… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/webui.imageimage-to-text10K<n<100K26 likes1.2k downloads7mo agoHugging Face05g-ronimo /riddles_evolved Riddles turned into conversations using mistralai/Mistral-7B-Instruct-v0.2 Seeded with Hypersniper's riddles_v1, buy him Ko-fi Structure: each sample = conversation with two turns: Q/A/Q/A Process: use Mistral to 1) expand riddles 2) answer riddle 3) formulate human follow-up question 4) answer follow-up question Code: GitHub Note: This is an unfiltered dataset, it for sure contains very bad answers. text1K<n<10K5 likes815 downloads3y agoHugging Face06g-ronimo /IN1k256-AR-buckets-bfl16latents_dc-ae-f32c32-sana-1.0_recaptext1M<n<10M0 likes442 downloads2y agoHugging Face07g-ronimo /PD12M-256px_dc-ae-f32c32-sana-1.0text10M<n<100M0 likes408 downloads1y agoHugging Face08ronanhansel /imagenet-1k-validation-subsetsimage100K<n<1M0 likes388 downloads10mo agoHugging Face09akahana /rontgen Citation If you use the ROCOv2 dataset in your research, please cite the following paper: Pelka, O., Menze, B. H., & Rexhausen, S. E. (2023). Radiology Objects in COntext version 2 (ROCOv2): A multimodal dataset for medical image analysis. arXiv preprint arXiv:2405.10004. @misc {ronan_l.m._2024, author = { {Ronan L.M.} }, title = { ROCOv2-radiology (Revision 5d66908) }, year = 2024, url = {… See the full description on the dataset page: https://huggingface.co/datasets/akahana/rontgen.image10K<n<100K0 likes385 downloads2y agoHugging Face10RonPlusSign /PutRubbishInBin_500_episodes_frames_3camsThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "franka", "total_episodes": 500, "total_frames": 77242, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 10, "splits": { "train": "0:500" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/RonPlusSign/PutRubbishInBin_500_episodes_frames_3cams.imagerobotics10K<n<100K0 likes295 downloads11mo agoHugging Face11RonakAC /Animated-World-150M-v1-Unifiedimage100K<n<1M0 likes286 downloads5mo agoHugging Face12ronunes /LegiSubject-Br-Summaries 🇧🇷 Brazilian Legislative Bills – Summary Dataset This dataset contains summaries (ementas) of legislative bills proposed in the Brazilian Chamber of Deputies (BCoD) from 1991 to 2022.It is intended for multi-label classification, where each bill may be associated with one or more subject categories (temas). 🔀 This is the summary version of the dataset.If you are looking for the keywords version, see:👉 ronunes/LegiSubject-Br-Keywords 📁 Dataset Structure The… See the full description on the dataset page: https://huggingface.co/datasets/ronunes/LegiSubject-Br-Summaries.texttext-classification1M<n<10M3 likes231 downloads1y agoHugging Face13Zogfryt /roneneldan-TinyStories-tokenizer-distilgpt22100K<n<1M0 likes230 downloads2y agoHugging Face14community-datasets /ronec Dataset Card for RONEC Dataset Summary RONEC, at version 2.0, holds 12330 sentences with over 0.5M tokens, annotated with 15 classes, to a total of 80.283 distinctly annotated entities. The corpus has the following classes and distribution in the train/valid/test splits: | Classes | Total | Train | | Valid | | Test | | |------------- |:------: |:------: |:-------: |:------: |:-------: |:------: |:-------: | | |… See the full description on the dataset page: https://huggingface.co/datasets/community-datasets/ronec.texttoken-classification10K<n<100K6 likes220 downloads2y agoHugging Face15apollo-research /roneneldan-TinyStories-tokenizer-gpt2100K<n<1M2 likes214 downloads3y agoHugging Face16g-ronimo /IN1k256-AR-buckets-latents_dc-ae-f32c32-sana-1.0text1M<n<10M0 likes209 downloads2y agoHugging Face17g-ronimo /NIH-Chest-X-ray-dataset_resized300pximage100K<n<1M0 likes199 downloads10mo agoHugging Face18RongchangLi /Sth_comtext10K<n<100K0 likes195 downloads2y agoHugging Face19RonakAC /nexus-core-data-v210M<n<100M0 likes189 downloads4mo agoHugging Face20g-ronimo /IN1k256-bfl16latents_shape_dc-ae-f32c32-sana-1.0text1M<n<10M0 likes186 downloads1y agoHugging Face21ronantakizawa /python-code-instructions-japanese Python Code Instructions - Japanese (18K) Dataset Description This dataset contains 18,612 Python programming instruction-response pairs translated to Japanese. It's designed for training language models to understand and generate Python code based on Japanese instructions. Key Features 18,612 entries covering diverse Python programming tasks Japanese instructions and prompts for code generation Original English text preserved for reference Python code… See the full description on the dataset page: https://huggingface.co/datasets/ronantakizawa/python-code-instructions-japanese.texttext-generation10K<n<100K2 likes181 downloads10mo agoHugging Face22g-ronimo /CC12M_IN21K-256px_dc-ae-f32c32-sana-1.0text10M<n<100M0 likes178 downloads1y agoHugging Face23RonPlusSign /PutRubbishInBin_500_episodes_frames_5camsThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "franka", "total_episodes": 500, "total_frames": 77242, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 10, "splits": { "train": "0:500" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/RonPlusSign/PutRubbishInBin_500_episodes_frames_5cams.imagerobotics10K<n<100K0 likes176 downloads11mo agoHugging Face24g-ronimo /kaggle_llm_science_examHF dataset of Kaggle's LLM Science Exam text1K<n<10K0 likes174 downloads2y agoHugging Face25g-ronimo /CC12M_IN21K-256px-splits_dc-ae-f32c32-sana-1.0text10M<n<100M0 likes168 downloads1y agoHugging Face26RonLiao /lerobot-so101-elevator-datasetThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "robot_type": "so101_follower", "total_episodes": 51, "total_frames": 13947, "total_tasks": 1, "chunks_size": 1000, "data_files_size_in_mb": 100, "video_files_size_in_mb": 500, "fps": 30, "splits": { "train": "0:51" }, "data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/RonLiao/lerobot-so101-elevator-dataset.tabularrobotics10K<n<100K0 likes167 downloads7mo agoHugging Face27g-ronimo /IN1k256-AR-buckets-bfl16latents_dc-ae-f32c32-sana-1.0text1M<n<10M1 likes152 downloads1y agoHugging Face28ronniejiangC /MM-RIS MM-RIS: Multimodal Referring Image Segmentation Dataset The MM-RIS dataset was introduced in the paper RIS-FUSION: Rethinking Text-Driven Infrared and Visible Image Fusion from the Perspective of Referring Image Segmentation. This large-scale benchmark supports the multimodal referring image segmentation (RIS) task by providing a goal-aligned approach to supervise and evaluate how effectively natural language contributes to infrared and visible image fusion outcomes.… See the full description on the dataset page: https://huggingface.co/datasets/ronniejiangC/MM-RIS.textimage-segmentation10K<n<100K1 likes132 downloads1y agoHugging Face29ronunes /LegiSubject-Br-Keywords 🇧🇷 Brazilian Legislative Bills – Keyword Dataset This dataset contains keywords of legislative bills proposed in the Brazilian Chamber of Deputies (BCoD) from 1991 to 2022.It is intended for multi-label classification, where each bill may be associated with one or more subject categories (temas). 🔀 This is the keywords version of the dataset.If you are looking for the summaries version, see:👉 ronunes/LegiSubject-Br-Keywords 📁 Dataset Structure The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/ronunes/LegiSubject-Br-Keywords.texttext-classification1M<n<10M3 likes131 downloads1y agoHugging Face30g-ronimo /Imagenet-256-latents_dc-ae-f32c32-sana-1.01M<n<10M0 likes129 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.