CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01darknoon /svg-stack-filtered Dataset Card for svg-stack-filtered This is an attempt to replicate the dataset used for SFT in the paper Rendering-Aware Reinforcement Learning for Vector Graphics Generation Processed: Optimized with svgo precision=2 Rasterized with cairosvg[^cairo] [^cairo] cairosvg doesn't implement all svg features, but matches how the original paper Filtered based on some heuristics: Removed any svg that couldn't be rendered with cairosvg (~30%) Removed solid-color images Removed some… See the full description on the dataset page: https://huggingface.co/datasets/darknoon/svg-stack-filtered.image1M<n<10M3 likes588 downloads1y agoHugging Face02darkyarding /MMEimage1K<n<10K12 likes533 downloads2y agoHugging Face03darknight054 /med-mts-audio-kokoro-82m MTSamples‑Kokoro‑ASR (Synthetic Medical Speech) Summary: 279 hours of synthetic English medical speech (49,462 clips) created from publicly available transcripts on MTSamples.com using multiple US/UK voices from Kokoro‑82M. Intended for training and evaluating medical ASR. Dataset Rows: 49,462 Total audio: ~279 hours (mono) Source text: Sample medical reports from MTSamples.com (names/dates typically altered or removed) Audio generation: hexgrad/Kokoro‑82M (various… See the full description on the dataset page: https://huggingface.co/datasets/darknight054/med-mts-audio-kokoro-82m.audio10K<n<100K0 likes476 downloads1y agoHugging Face04darknight054 /indic-mozhi-ocr Mozhi (Printed Word Images) - Indic OCR Dataset This folder contains the word-level printed OCR dataset downloaded from the CVIT USODI project page for "Towards Deployable OCR Models for Indic Languages". The data is organized by language and split (train/val/test) and is intended for upload to Hugging Face. Source Source page: https://cvit.iiit.ac.in/usodi/tdocrmil.php Paper: Towards Deployable OCR Models for Indic Languages Authors: Minesh Mathew, Ajoy Mondal, C V… See the full description on the dataset page: https://huggingface.co/datasets/darknight054/indic-mozhi-ocr.image1M<n<10M2 likes318 downloads8mo agoHugging Face05darknoon /quickdraw10M<n<100M0 likes262 downloads1y agoHugging Face06villekuosmanen /agilex_pour_water_dark_2This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "arx5_bimanual", "total_episodes": 39, "total_frames": 15587, "total_tasks": 1, "total_videos": 117, "total_chunks": 1, "chunks_size": 1000, "fps": 25, "splits": { "train": "0:39" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/agilex_pour_water_dark_2.tabularrobotics10K<n<100K0 likes239 downloads6mo agoHugging Face07darknight054 /pubmed_cleanA cleaned Pubmed commercial available files dataset. Will update the script used to clean soon. text1M<n<10M2 likes200 downloads2y agoHugging Face08chair0 /single_dark_single_20260906_235610This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/chair0/single_dark_single_20260906_235610.tabularroboticsn<1K0 likes194 downloads18d agoHugging Face09Dark3s /exorde-social-media-december-2024-week1texttext-classification10M<n<100M0 likes185 downloads7mo agoHugging Face10chair0 /single_dark_lights_off_single_trimmedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/chair0/single_dark_lights_off_single_trimmed.tabularrobotics1K<n<10K0 likes174 downloads18d agoHugging Face11chair0 /single_dark_lights_off_single_20260907_000627This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/chair0/single_dark_lights_off_single_20260907_000627.tabularrobotics1K<n<10K0 likes172 downloads18d agoHugging Face12chair0 /single_dark_single_20260906_235812This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/chair0/single_dark_single_20260906_235812.tabularrobotics1K<n<10K0 likes168 downloads18d agoHugging Face13chair0 /single_dark_single_trimmedThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/chair0/single_dark_single_trimmed.tabularrobotics1K<n<10K0 likes141 downloads18d agoHugging Face14mwhnh10 /pick_cube_test_147eps_darkp1This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/mwhnh10/pick_cube_test_147eps_darkp1.tabularrobotics100K<n<1M0 likes121 downloads2mo agoHugging Face15dark-xet /test-public-datasetDataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary. Described in the following paper: https://arxiv.org/abs/2305.07759. The models referred to in the paper were trained on TinyStories-train.txt (the file tinystories-valid.txt can be used for validation loss). These models can be found on Huggingface, at roneneldan/TinyStories-1M/3M/8M/28M/33M/1Layer-21M. Additional resources: tinystories_all_data.tar.gz - contains a superset of… See the full description on the dataset page: https://huggingface.co/datasets/dark-xet/test-public-dataset.texttext-generation1M<n<10M0 likes112 downloads2y agoHugging Face16freelion /darkpatterns_in_llmgatedtext10K<n<100K0 likes112 downloads9d agoHugging Face17villekuosmanen /agilex_pour_water_dark_meeting_roomThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.0", "robot_type": "arx5_bimanual", "total_episodes": 18, "total_frames": 4950, "total_tasks": 1, "total_videos": 54, "total_chunks": 1, "chunks_size": 1000, "fps": 25, "splits": { "train": "0:18" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/agilex_pour_water_dark_meeting_room.tabularrobotics1K<n<10K0 likes109 downloads6mo agoHugging Face18Tonic /dark_thoughts_stakeholders_testtexttext-generation100K<n<1M0 likes101 downloads2y agoHugging Face19DataTonic /dark_thoughts_casestudies_en_cn Dark Thoughts Case Studies Dataset (English-Chinese) This dataset contains a bilingual collection of case studies with detailed stakeholder analyses in English and Chinese. Each case study includes structured information about stakeholders and their motivations, along with comprehensive case analysis and solutions. Dataset Description Overview The dataset consists of 344,580 paired case studies in English and Chinese, with detailed stakeholder analyses and… See the full description on the dataset page: https://huggingface.co/datasets/DataTonic/dark_thoughts_casestudies_en_cn.texttext-generation100K<n<1M3 likes99 downloads2y agoHugging Face20darknight054 /indicstr12-crops IndicSTR12 (Cropped Word Images) This document describes the cropped word image portion of the IndicSTR12 real dataset. Source This dataset was downloaded from the CVIT IndicSTR12 Project. Paper: IndicSTR12: A Dataset for Indic Scene Text RecognitionAuthors: Harsh Lunia, Ajoy Mondal, C V JawaharConference: ICDAR 2023 Note: This repository contains only the Real Dataset. The Synthetic Dataset is not included. Structure raw_data/ ├── assamese/ │ └──… See the full description on the dataset page: https://huggingface.co/datasets/darknight054/indicstr12-crops.image10K<n<100K0 likes99 downloads8mo agoHugging Face21ranjith3567 /Green_Square_APP_v4_Top_View_Dark_Light_View_20260918_180147This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "shape": [ 6 ], "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/ranjith3567/Green_Square_APP_v4_Top_View_Dark_Light_View_20260918_180147.tabularrobotics10K<n<100K0 likes88 downloads7d agoHugging Face22DataTonic /dark_thoughts_case_study_merged Dark Thoughts 案例研究推理数据集 数据集描述 概述 Dark Thoughts 案例研究推理数据集是一个全面的多语言商业案例研究及相关推理响应集合。它通过先进的语言模型处理 Cablegate 电报,生成中英文商业案例研究,并进一步丰富了利益相关者特定的推理视角。对于对商业分析、多语言内容生成和推理能力感兴趣的研究人员和从业人员来说,该数据集是宝贵的资源。 支持的任务 该数据集支持以下任务: 文本生成 推理与分析 双语案例研究生成 跨语言内容分析 商业战略制定 利益相关者视角建模 语言 该数据集为双语数据集: 英语 (en) 中文 (zh) 数据集结构 数据字段 { 'id': 'int32', # 条目的唯一标识符 'response': 'string', # 生成的推理响应 'query': 'string', # 原始查询或案例研究内容 'source_data': 'string', #… See the full description on the dataset page: https://huggingface.co/datasets/DataTonic/dark_thoughts_case_study_merged.texttext-generation10K<n<100K5 likes78 downloads1y agoHugging Face23mlnomad /imnet1k_sunglasses_dark_glasses_shadesimage1K<n<10K1 likes72 downloads1y agoHugging Face24JiabinQ /eval_dark_fork_bgd_6k_lightThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 10, "total_frames": 5217, "total_tasks": 1, "total_videos": 20, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:10" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/JiabinQ/eval_dark_fork_bgd_6k_light.tabularrobotics1K<n<10K0 likes68 downloads1y agoHugging Face25JiabinQ /eval_dark_fork_bgd_12k_downThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 10, "total_frames": 5467, "total_tasks": 1, "total_videos": 20, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:10" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/JiabinQ/eval_dark_fork_bgd_12k_down.tabularrobotics1K<n<10K0 likes62 downloads1y agoHugging Face26DataTonic /dark_thoughts_stakeholders_en_cn Dark Thoughts Case Studies Dataset (English-Chinese) This dataset contains a bilingual collection of case studies with detailed stakeholder analyses in English and Chinese. Each case study includes structured information about stakeholders and their motivations, along with comprehensive case analysis and solutions. Dataset Description Overview The dataset consists of 344,580 case studies in English and in Chinese, with detailed stakeholder analyses and… See the full description on the dataset page: https://huggingface.co/datasets/DataTonic/dark_thoughts_stakeholders_en_cn.texttext-generation100K<n<1M2 likes61 downloads2y agoHugging Face27darknoon /simple-shapes-svgThe goal of this dataset is to measure and improve the ability of VLMs to see accurately in spatial dimensions. I've tried to ensure that all of the examples are not too hard have sufficient contrast between foreground and background shapes are not clipped or ambiguous solid background canvas is square 512x512 Initially, I've kept the "canvas" that they're working with 512x512 points, but you can learn more by experimenting with the dimensions as well. imageimage-to-text10K<n<100K1 likes61 downloads1y agoHugging Face28DataTonic /dark_thoughts_case_study_reason Dark Thoughts 案例研究数据集 - 推理 数据集描述 概述 Dark Thoughts 案例研究数据集 - 推理是一个全面的多语言商业案例研究及相关推理回复集合。该数据集通过先进的语言模型处理 Cablegate 电报,生成中英文商业案例研究,并进一步丰富了利益相关者特定的推理视角。对于对商业分析、多语言内容生成和推理能力感兴趣的研究人员和从业人员来说,该数据集是宝贵的资源。 支持的任务 该数据集支持以下任务: 文本生成 语言建模 推理与分析 双语案例研究生成 跨语言内容分析 商业战略制定 利益相关者视角建模 语言 该数据集为双语数据集: 英语 (en) 中文 (zh) 数据集结构 数据字段 { 'id': 'string', # 条目的唯一标识符 'think': 'string', # 思考过程 'response': 'string', # 生成的推理响应 'query': 'string', #… See the full description on the dataset page: https://huggingface.co/datasets/DataTonic/dark_thoughts_case_study_reason.texttext-generation10K<n<100K10 likes60 downloads1y agoHugging Face29JiabinQ /eval_dark_fork_bgd_18k_downThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": "so101_follower", "total_episodes": 10, "total_frames": 5135, "total_tasks": 1, "total_videos": 20, "total_chunks": 1, "chunks_size": 1000, "fps": 30, "splits": { "train": "0:10" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/JiabinQ/eval_dark_fork_bgd_18k_down.tabularrobotics1K<n<10K0 likes58 downloads1y agoHugging Face30DataTonic /dark_thoughts_stakeholders Dataset Card for Dataset Name This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Dataset Details Dataset Description Curated by: [More Information Needed] Funded by [optional]: [More Information Needed] Shared by [optional]: [More Information Needed] Language(s) (NLP): [More Information Needed] License: [More Information Needed] Dataset Sources [optional] Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/DataTonic/dark_thoughts_stakeholders.texttext-generation100K<n<1M1 likes57 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.