open-dataset
nq_open
Dataset Card for nq_open
Dataset Summary
The NQ-Open task, introduced by Lee et.al. 2019,
is an open domain question answering benchmark that is derived from Natural Questions.
The goal is to predict an English answer string for an input English question.
All questions can be answered using the contents of English Wikipedia.
Supported Tasks and Leaderboards
Open Domain Question-Answering,
EfficientQA Leaderboard:… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/nq_open.waymo_open_dataset_v_1_4_3Galaxea-Open-World-Dataset
Galaxea Open-World Dataset
Key Features
500+ hours of real-world mobile manipulation data.
All data collected using one uniform robotic embodiment (R1-Lite) for consistency.
Fine-grained subtask language annotations (bilingual Chinese/English).
Covers residential, kitchen, retail, and officesettings.
Dataset in LeRobot v2.1 format.
Dataset Structure
The dataset is organized as 227 task-level tar.gz archives under the lerobot/ directory. Each… See the full description on the dataset page: https://huggingface.co/datasets/OpenGalaxea/Galaxea-Open-World-Dataset.Galaxea-Open-World-Dataset-LeRobot-v3.0Galaxea Open-World Dataset taken from OpenGalaxea/Galaxea-Open-World-Dataset,
converted to LeRobot Datasets v3.0 format using lerobot.datasets.v30.convert_dataset_v21_to_v30.
Missing subsets
The subset Boil_The_Water_20250714_006 is missing due to the original files having some episodes at 62 fps,
which causes the conversion script to crash with an error.
The subset Put_The_Items_Into_The_Storage_Box_20250929_002_007 is missing due to it having 7 DoF arms rather than 6 DoF.… See the full description on the dataset page: https://huggingface.co/datasets/griffinlabs/Galaxea-Open-World-Dataset-LeRobot-v3.0.open-models-prompt-datasets
🖼️ Open Models Prompt Dataset
🖼️ The ultimate open models image prompt dataset (10GB+). 5400+ image generation prompts with full metadata and preview images. Truly open source: No login, no ads, no redirection. Just pure data for AI image creators.
This project is a massive collection of prompts used for various open-source AI image models and the resulting generated images. The entire dataset exceeds 10GB and contains 5400+ images, all structured into a comprehensive… See the full description on the dataset page: https://huggingface.co/datasets/Goku-OpenLab/open-models-prompt-datasets.dalle-3-dataset
Dataset Card for LAION DALL·E 3 Discord Dataset
Description: This dataset consists of caption and image pairs scraped from the LAION share-dalle-3 discord channel. The purpose is to collect image-text pairs for research and exploration.
Source Code: The code used to generate this data can be found here.
Contributors
Zach Nagengast
Eduardo Pach
Seva Maltsev
Ben Egan
The LAION community
Data Attributes
caption: The text description or prompt associated with… See the full description on the dataset page: https://huggingface.co/datasets/OpenDatasets/dalle-3-dataset.
