datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
svg-stack-filtered
Dataset Card for svg-stack-filtered
This is an attempt to replicate the dataset used for SFT in the paper
Rendering-Aware Reinforcement Learning for Vector Graphics Generation
Processed:
Optimized with svgo precision=2
Rasterized with cairosvg[^cairo]
[^cairo] cairosvg doesn't implement all svg features, but matches how the original paper
Filtered based on some heuristics:
Removed any svg that couldn't be rendered with cairosvg (~30%)
Removed solid-color images
Removed some… See the full description on the dataset page: https://huggingface.co/datasets/darknoon/svg-stack-filtered.MMEmed-mts-audio-kokoro-82m
MTSamples‑Kokoro‑ASR (Synthetic Medical Speech)
Summary: 279 hours of synthetic English medical speech (49,462 clips) created from publicly available transcripts on MTSamples.com using multiple US/UK voices from Kokoro‑82M. Intended for training and evaluating medical ASR.
Dataset
Rows: 49,462
Total audio: ~279 hours (mono)
Source text: Sample medical reports from MTSamples.com (names/dates typically altered or removed)
Audio generation: hexgrad/Kokoro‑82M (various… See the full description on the dataset page: https://huggingface.co/datasets/darknight054/med-mts-audio-kokoro-82m.indic-mozhi-ocr
Mozhi (Printed Word Images) - Indic OCR Dataset
This folder contains the word-level printed OCR dataset downloaded from the CVIT USODI project page for
"Towards Deployable OCR Models for Indic Languages". The data is organized by language and split
(train/val/test) and is intended for upload to Hugging Face.
Source
Source page: https://cvit.iiit.ac.in/usodi/tdocrmil.php
Paper: Towards Deployable OCR Models for Indic Languages
Authors: Minesh Mathew, Ajoy Mondal, C V… See the full description on the dataset page: https://huggingface.co/datasets/darknight054/indic-mozhi-ocr.quickdrawagilex_pour_water_dark_2This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "arx5_bimanual",
"total_episodes": 39,
"total_frames": 15587,
"total_tasks": 1,
"total_videos": 117,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 25,
"splits": {
"train": "0:39"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/agilex_pour_water_dark_2.pubmed_cleanA cleaned Pubmed commercial available files dataset. Will update the script used to clean soon.
single_dark_single_20260906_235610This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/chair0/single_dark_single_20260906_235610.exorde-social-media-december-2024-week1single_dark_lights_off_single_trimmedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/chair0/single_dark_lights_off_single_trimmed.single_dark_lights_off_single_20260907_000627This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/chair0/single_dark_lights_off_single_20260907_000627.single_dark_single_20260906_235812This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/chair0/single_dark_single_20260906_235812.single_dark_single_trimmedThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/chair0/single_dark_single_trimmed.pick_cube_test_147eps_darkp1This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"
],
"shape": [
6… See the full description on the dataset page: https://huggingface.co/datasets/mwhnh10/pick_cube_test_147eps_darkp1.test-public-datasetDataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary.
Described in the following paper: https://arxiv.org/abs/2305.07759.
The models referred to in the paper were trained on TinyStories-train.txt (the file tinystories-valid.txt can be used for validation loss). These models can be found on Huggingface, at roneneldan/TinyStories-1M/3M/8M/28M/33M/1Layer-21M.
Additional resources:
tinystories_all_data.tar.gz - contains a superset of… See the full description on the dataset page: https://huggingface.co/datasets/dark-xet/test-public-dataset.darkpatterns_in_llmagilex_pour_water_dark_meeting_roomThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "arx5_bimanual",
"total_episodes": 18,
"total_frames": 4950,
"total_tasks": 1,
"total_videos": 54,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 25,
"splits": {
"train": "0:18"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/villekuosmanen/agilex_pour_water_dark_meeting_room.dark_thoughts_stakeholders_testdark_thoughts_casestudies_en_cn
Dark Thoughts Case Studies Dataset (English-Chinese)
This dataset contains a bilingual collection of case studies with detailed stakeholder analyses in English and Chinese. Each case study includes structured information about stakeholders and their motivations, along with comprehensive case analysis and solutions.
Dataset Description
Overview
The dataset consists of 344,580 paired case studies in English and Chinese, with detailed stakeholder analyses and… See the full description on the dataset page: https://huggingface.co/datasets/DataTonic/dark_thoughts_casestudies_en_cn.indicstr12-crops
IndicSTR12 (Cropped Word Images)
This document describes the cropped word image portion of the IndicSTR12 real dataset.
Source
This dataset was downloaded from the CVIT IndicSTR12 Project.
Paper: IndicSTR12: A Dataset for Indic Scene Text RecognitionAuthors: Harsh Lunia, Ajoy Mondal, C V JawaharConference: ICDAR 2023
Note: This repository contains only the Real Dataset. The Synthetic Dataset is not included.
Structure
raw_data/
├── assamese/
│ └──… See the full description on the dataset page: https://huggingface.co/datasets/darknight054/indicstr12-crops.Green_Square_APP_v4_Top_View_Dark_Light_View_20260918_180147This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"fps": 30,
"features": {
"action": {
"dtype": "float32",
"shape": [
6
],
"names": [
"shoulder_pan.pos",
"shoulder_lift.pos",
"elbow_flex.pos",
"wrist_flex.pos",
"wrist_roll.pos",
"gripper.pos"… See the full description on the dataset page: https://huggingface.co/datasets/ranjith3567/Green_Square_APP_v4_Top_View_Dark_Light_View_20260918_180147.dark_thoughts_case_study_merged
Dark Thoughts 案例研究推理数据集
数据集描述
概述
Dark Thoughts 案例研究推理数据集是一个全面的多语言商业案例研究及相关推理响应集合。它通过先进的语言模型处理 Cablegate 电报,生成中英文商业案例研究,并进一步丰富了利益相关者特定的推理视角。对于对商业分析、多语言内容生成和推理能力感兴趣的研究人员和从业人员来说,该数据集是宝贵的资源。
支持的任务
该数据集支持以下任务:
文本生成
推理与分析
双语案例研究生成
跨语言内容分析
商业战略制定
利益相关者视角建模
语言
该数据集为双语数据集:
英语 (en)
中文 (zh)
数据集结构
数据字段
{
'id': 'int32', # 条目的唯一标识符
'response': 'string', # 生成的推理响应
'query': 'string', # 原始查询或案例研究内容
'source_data': 'string', #… See the full description on the dataset page: https://huggingface.co/datasets/DataTonic/dark_thoughts_case_study_merged.imnet1k_sunglasses_dark_glasses_shadeseval_dark_fork_bgd_6k_lightThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 5217,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/JiabinQ/eval_dark_fork_bgd_6k_light.eval_dark_fork_bgd_12k_downThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 5467,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/JiabinQ/eval_dark_fork_bgd_12k_down.dark_thoughts_stakeholders_en_cn
Dark Thoughts Case Studies Dataset (English-Chinese)
This dataset contains a bilingual collection of case studies with detailed stakeholder analyses in English and Chinese. Each case study includes structured information about stakeholders and their motivations, along with comprehensive case analysis and solutions.
Dataset Description
Overview
The dataset consists of 344,580 case studies in English and in Chinese, with detailed stakeholder analyses and… See the full description on the dataset page: https://huggingface.co/datasets/DataTonic/dark_thoughts_stakeholders_en_cn.simple-shapes-svgThe goal of this dataset is to measure and improve the ability of VLMs to see accurately in spatial dimensions.
I've tried to ensure that all of the examples are not too hard
have sufficient contrast between foreground and background
shapes are not clipped or ambiguous
solid background
canvas is square 512x512
Initially, I've kept the "canvas" that they're working with 512x512 points, but you can learn more by experimenting with the dimensions as well.
dark_thoughts_case_study_reason
Dark Thoughts 案例研究数据集 - 推理
数据集描述
概述
Dark Thoughts 案例研究数据集 - 推理是一个全面的多语言商业案例研究及相关推理回复集合。该数据集通过先进的语言模型处理 Cablegate 电报,生成中英文商业案例研究,并进一步丰富了利益相关者特定的推理视角。对于对商业分析、多语言内容生成和推理能力感兴趣的研究人员和从业人员来说,该数据集是宝贵的资源。
支持的任务
该数据集支持以下任务:
文本生成
语言建模
推理与分析
双语案例研究生成
跨语言内容分析
商业战略制定
利益相关者视角建模
语言
该数据集为双语数据集:
英语 (en)
中文 (zh)
数据集结构
数据字段
{
'id': 'string', # 条目的唯一标识符
'think': 'string', # 思考过程
'response': 'string', # 生成的推理响应
'query': 'string', #… See the full description on the dataset page: https://huggingface.co/datasets/DataTonic/dark_thoughts_case_study_reason.eval_dark_fork_bgd_18k_downThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101_follower",
"total_episodes": 10,
"total_frames": 5135,
"total_tasks": 1,
"total_videos": 20,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/JiabinQ/eval_dark_fork_bgd_18k_down.dark_thoughts_stakeholders
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/DataTonic/dark_thoughts_stakeholders.
