datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Robo2VLM-1
Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets
Abstract
Vision-Language Models (VLMs) acquire real-world knowledge and general reasoning ability through Internet-scale image-text corpora. They can augment robotic systems with scene understanding and task planning, and assist visuomotor policies that are trained on robot trajectory data. We explore the reverse paradigm - using rich, real, multi-modal robot trajectory data to… See the full description on the dataset page: https://huggingface.co/datasets/keplerccc/Robo2VLM-1.ManipulationVQA-60k
Robo2VLM-Reasoning
Samples from the dataset: Robo2VLM-1, prompting gemini-2.5-pro to generate reasoning traces supporting the correct choice.
Paper: Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets
@misc{chen2025robo2vlmvisualquestionanswering,
title={Robo2VLM: Visual Question Answering from Large-Scale In-the-Wild Robot Manipulation Datasets},
author={Kaiyuan Chen and Shuangyu Xie and Zehan Ma and Ken Goldberg}… See the full description on the dataset page: https://huggingface.co/datasets/keplerccc/ManipulationVQA-60k.hunt-for-worlds-kepler
Hunt for Worlds: Kepler Exoplanet Detection
Dataset Summary
A public, viewer-ready educational challenge dataset. Host-only scoring data and hidden targets are excluded.
Splits
Split
Examples
Description
train
7,662
Labeled training data
test
1,902
Public inputs with withheld target labels or annotations
Data Fields
Field
Type
id
int64
kepid
int64
koi_disposition
object
koi_period
float64… See the full description on the dataset page: https://huggingface.co/datasets/hoangbang/hunt-for-worlds-kepler.UrduShers-10kkepler-arc-agi-3-traces
Kepler 1.0 ARC-AGI-3 trace corpus
Run artifacts from Kepler 1.0, an open-source agent harness for the 25 public
ARC-AGI-3 games. A stock CLI coding agent
encodes its theory of each game as an executable world_model.py, certifies it
against the full recorded interaction history, plans inside the certified
model, and acts through a guarded channel that voids the plan on the first
misprediction.
Project page ·
Code ·
Paper ·
Integrity record
The canonical release contains two… See the full description on the dataset page: https://huggingface.co/datasets/cveinnt/kepler-arc-agi-3-traces.kepler-observations
Kepler Observation Catalog
Credit: NASA/Ames/JPL-Caltech
Part of a dataset collection on Hugging Face.
Dataset description
The Kepler Observation Catalog indexes every target observed by NASA's Kepler space telescope during its prime mission (2009–2013), drawn from the Mikulski Archive for Space Telescopes (MAST). Kepler is the most successful exoplanet-hunting mission in history: by continuously monitoring ~200,000 stars in a single 100-square-degree… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/kepler-observations.kepler_flare
Kepler Flare Dataset
Dataset Description
The Kepler Flare Dataset is a comprehensive collection of stellar flare events observed by the Kepler Space Telescope. This dataset is constructed based on the flare event catalog presented in Yang & Liu (2019), which provides a systematic study of stellar flares in the Kepler mission. All light curves are the long cadence (about 29.4 minutes) data.
Dataset Features & Content
The dataset consists of high-quality… See the full description on the dataset page: https://huggingface.co/datasets/Maxwell-Jia/kepler_flare.kepler-transit-timing
Kepler Transit Timing Catalog
Part of the Astronomy Datasets collection on Hugging Face.
Transit timing catalog from Holczer et al. (2016), containing 295,187 individual transit
mid-times for 2,599 Kepler Objects of Interest (KOIs). Each record includes the
observed mid-transit time, observed-minus-computed (O-C) residual, transit duration, and
transit depth with uncertainties.
Dataset description
Transit timing variations (TTVs) occur when gravitational interactions… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/kepler-transit-timing.Kepler-bench
Kepler astro-bench v0.1
The benchmark behind Orionfold/Kepler-GGUF —
a verifier-checked set of astrodynamics and quantitative-astrophysics word problems, each with a
single numeric gold answer and a programmatic verifier that doubles as a reinforcement-learning
reward.
What's here
File
Rows
Purpose
pool.jsonl
120
Training / selection pool — 16 formula families (9 orbital, 7 astrophysics), 3 difficulty tiers.
heldout.jsonl
44
External curveball… See the full description on the dataset page: https://huggingface.co/datasets/Orionfold/Kepler-bench.kepler-eclipsing-binaries
Kepler Eclipsing Binary Catalog
Catalog of 2,177 eclipsing binary systems identified by the Kepler mission,
with orbital periods, morphology parameters, and stellar properties.
Dataset description
Eclipsing binaries are pairs of stars whose orbital plane is aligned with our line of
sight, producing periodic dips in brightness as one star passes in front of the other.
The Kepler mission's exquisite photometric precision made it ideal for detecting and… See the full description on the dataset page: https://huggingface.co/datasets/juliensimon/kepler-eclipsing-binaries.UrduGhazals-25ktranslation-dataset-105KPAARI-English-TTS
PAARI English TTS Dataset
Dataset Description
This dataset contains TTS-optimized chunks of journalism articles in English (English) from the People's Archive of Rural India (PAARI). The articles focus on rural life, agriculture, social issues, and cultural stories from rural India.
Dataset Details
Language: English (English)
Script: Latin
Language Code: en
Dataset Type: TTS-optimized
Source: Rural India Online
License: Please refer to PAARI's terms of use… See the full description on the dataset page: https://huggingface.co/datasets/keplersystems/PAARI-English-TTS.kepler-39K39K records from NASA's kepler mission, newest to date 25/03/2026.
All the records were taken from: https://exoplanetarchive.ipac.caltech.edu/
Format: CSV
Good luck, data sufferers!
UrduPoetry-35kPAARI-Urdu-TTS
PAARI Urdu TTS Dataset
Dataset Description
This dataset contains TTS-optimized chunks of journalism articles in Urdu (اردو) from the People's Archive of Rural India (PAARI). The articles focus on rural life, agriculture, social issues, and cultural stories from rural India.
Dataset Details
Language: Urdu (اردو)
Script: Arabic
Language Code: ur
Dataset Type: TTS-optimized
Source: Rural India Online
License: Please refer to PAARI's terms of use
Dataset… See the full description on the dataset page: https://huggingface.co/datasets/keplersystems/PAARI-Urdu-TTS.PAARI-Tamil-TTS
PAARI Tamil TTS Dataset
Dataset Description
This dataset contains TTS-optimized chunks of journalism articles in Tamil (தமிழ்) from the People's Archive of Rural India (PAARI). The articles focus on rural life, agriculture, social issues, and cultural stories from rural India.
Dataset Details
Language: Tamil (தமிழ்)
Script: Tamil
Language Code: ta
Dataset Type: TTS-optimized
Source: Rural India Online
License: Please refer to PAARI's terms of use… See the full description on the dataset page: https://huggingface.co/datasets/keplersystems/PAARI-Tamil-TTS.PAARI-Telugu-TTS
PAARI Telugu TTS Dataset
Dataset Description
This dataset contains TTS-optimized chunks of journalism articles in Telugu (తెలుగు) from the People's Archive of Rural India (PAARI). The articles focus on rural life, agriculture, social issues, and cultural stories from rural India.
Dataset Details
Language: Telugu (తెలుగు)
Script: Telugu
Language Code: te
Dataset Type: TTS-optimized
Source: Rural India Online
License: Please refer to PAARI's terms of use… See the full description on the dataset page: https://huggingface.co/datasets/keplersystems/PAARI-Telugu-TTS.PAARI-Gujarati-TTS
PAARI Gujarati TTS Dataset
Dataset Description
This dataset contains TTS-optimized chunks of journalism articles in Gujarati (ગુજરાતી) from the People's Archive of Rural India (PAARI). The articles focus on rural life, agriculture, social issues, and cultural stories from rural India.
Dataset Details
Language: Gujarati (ગુજરાતી)
Script: Gujarati
Language Code: gu
Dataset Type: TTS-optimized
Source: Rural India Online
License: Please refer to PAARI's terms of… See the full description on the dataset page: https://huggingface.co/datasets/keplersystems/PAARI-Gujarati-TTS.tts-news-dataset
TTS News Dataset
This dataset is a collection of news statements and categories. It was originally created for the task of fake news detection but repurposed to create input for a TTS model.
Original Dataset
The original dataset is available at Kaggle.
About the Dataset
This IFND dataset covers news pertaining to India only. This dataset is created by scraping Indian fact checking websites. The dataset contains two types of news fake and real News. This dataset… See the full description on the dataset page: https://huggingface.co/datasets/keplersystems/tts-news-dataset.tts-news-dataset-dedupRobo2VLM-ERPAARI-Hindi-TTS
PAARI Hindi TTS Dataset
Dataset Description
This dataset contains TTS-optimized chunks of journalism articles in Hindi (हिन्दी) from the People's Archive of Rural India (PAARI). The articles focus on rural life, agriculture, social issues, and cultural stories from rural India.
Dataset Details
Language: Hindi (हिन्दी)
Script: Devanagari
Language Code: hi
Dataset Type: TTS-optimized
Source: Rural India Online
License: Please refer to PAARI's terms of use… See the full description on the dataset page: https://huggingface.co/datasets/keplersystems/PAARI-Hindi-TTS.PAARI-Punjabi-TTS
PAARI Punjabi TTS Dataset
Dataset Description
This dataset contains TTS-optimized chunks of journalism articles in Punjabi (ਪੰਜਾਬੀ) from the People's Archive of Rural India (PAARI). The articles focus on rural life, agriculture, social issues, and cultural stories from rural India.
Dataset Details
Language: Punjabi (ਪੰਜਾਬੀ)
Script: Gurmukhi
Language Code: pa
Dataset Type: TTS-optimized
Source: Rural India Online
License: Please refer to PAARI's terms of use… See the full description on the dataset page: https://huggingface.co/datasets/keplersystems/PAARI-Punjabi-TTS.PAARI-Marathi-TTS
PAARI Marathi TTS Dataset
Dataset Description
This dataset contains TTS-optimized chunks of journalism articles in Marathi (मराठी) from the People's Archive of Rural India (PAARI). The articles focus on rural life, agriculture, social issues, and cultural stories from rural India.
Dataset Details
Language: Marathi (मराठी)
Script: Devanagari
Language Code: mr
Dataset Type: TTS-optimized
Source: Rural India Online
License: Please refer to PAARI's terms of use… See the full description on the dataset page: https://huggingface.co/datasets/keplersystems/PAARI-Marathi-TTS.testdataUrduPoetry-106k
