datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Mobjaverse
Mobjaverse: A Large-Scale Rigged 3D Model Dataset with Skeletal Animations
Mobjaverse is a curated dataset derived from Objaverse-XL, specifically designed for research on skeletal animation understanding, motion generation, and articulated 3D shape analysis. It is curated in the paper TopoCap: Learning Topology-Agnostic Motion Priors for Monocular Video-to-Animation.
Mobjaverse contains ~19k rigged 3D models spanning ~5k distinct skeletal topologies and ~2M motion frames… See the full description on the dataset page: https://huggingface.co/datasets/duckduckplz/Mobjaverse.duckjam-dw2-vault
The Duck Jam Vault
Every submission entered in Duck Jam and every run the Arena
scored, as two Parquet tables and one JSON board per round. A publisher job rewrites this
repository once a day, so its git history is the history of every leaderboard.
This repository is a view. The rows themselves are written, once and never edited, into public
Hugging Face Buckets by the components that make them:
hf://buckets/Nico-robot/duckjam-submissions, hf://buckets/Nico-robot/duckjam-runs… See the full description on the dataset page: https://huggingface.co/datasets/Nico-robot/duckjam-dw2-vault.cu-vla-exp6-b0-lclickduckdb_ci_testsvital
Vital PresetShare Renders
Rendered Vital presets scraped from PresetShare.
sample rate: 22050
render duration: 6.0s
MIDI note: 72 (C4)
note duration: 5.0s
velocity: 100
Files are organized under by_type/<sound-type>/<preset-id>_<name>/ with:
preset.vital
preview.mp3
vital-render.wav
metadata.json
See manifest.jsonl and summary.json for run metadata.
AbstractEdit
Abstract Image Editing Benchmark
A benchmark for evaluating instruction-following image editing models on abstract
(open-ended) vs explicit (fully specified) editing instructions. Context images are
drawn from Open Images V7
(validation split). Each item pairs an abstract edit instruction
(e.g. "Pack these houses into boxes for shipping") with an explicit counterpart that
lists every atomic change required.
The dataset exposes three configurations:
benchmark — context images… See the full description on the dataset page: https://huggingface.co/datasets/DucktorV/AbstractEdit.duckpdbduckdb-text2sql-25k
Dataset Summary
The duckdb-text2sql-25k dataset contains 25,000 DuckDB text-2-sql pairs covering diverse aspects of DuckDB's SQL syntax.
We synthesized this dataset using Mixtral 8x7B, based on DuckDB's v0.9.2 documentation and Spider schemas that were translated to DuckDB syntax and enriched with nested type columns.
Each training sample consists of a natural language prompt, a corresponding (optional) schema, and a resulting query. Each pair furthermore has a category property… See the full description on the dataset page: https://huggingface.co/datasets/motherduckdb/duckdb-text2sql-25k.yolo-rubber-ducks
Rubber Duck Detection Dataset
Overview
This dataset contains 192 annotated images of rubber ducks, specifically curated for object detection tasks. It was used for experimentation related to the YOLOv8n Rubber Duck Detector model.
NOTE: I DO NOT RECOMMEND USING THIS DATASET AT THIS TIME. There is an open and ongoing discussion around the use of the datasets that were combined for this.See related licensing discussion on the forum
Dataset Description… See the full description on the dataset page: https://huggingface.co/datasets/phrogzx/yolo-rubber-ducks.duckjam-lp4-vault
The Duck Jam Vault
Every submission entered in Duck Jam and every run the Arena
scored, as two Parquet tables and one JSON board per round. A publisher job rewrites this
repository once a day, so its git history is the history of every leaderboard.
This repository is a view. The rows themselves are written, once and never edited, into public
Hugging Face Buckets by the components that make them:
hf://buckets/Nico-robot/duckjam-submissions, hf://buckets/Nico-robot/duckjam-runs… See the full description on the dataset page: https://huggingface.co/datasets/Nico-robot/duckjam-lp4-vault.duckjam-vd5-vault
The Duck Jam Vault
Every submission entered in Duck Jam and every run the Arena
scored, as two Parquet tables and one JSON board per round. A publisher job rewrites this
repository once a day, so its git history is the history of every leaderboard.
This repository is a view. The rows themselves are written, once and never edited, into public
Hugging Face Buckets by the components that make them:
hf://buckets/duckjam/submissions, hf://buckets/duckjam/runs… See the full description on the dataset page: https://huggingface.co/datasets/Nico-robot/duckjam-vd5-vault.duckjam-lp3-seasons
The Duck Jam Vault
Every submission entered in Duck Jam and every run the Arena
scored, as two Parquet tables and one JSON board per round. A publisher job rewrites this
repository once a day, so its git history is the history of every leaderboard.
This repository is a view. The rows themselves are written, once and never edited, into public
Hugging Face Buckets by the components that make them:
hf://buckets/duckjam/submissions, hf://buckets/duckjam/runs… See the full description on the dataset page: https://huggingface.co/datasets/Nico-robot/duckjam-lp3-seasons.mc4_310mc4 but in HPC friendly parquet format (32GiB shards)
Attribution,license, copyright info: Google and AI^2 for producing and uploading them.
rubber_ducksplace_pink_duck_pruned_2nd
place_pink_duck_pruned_2nd (TsFile)
Apache TsFile version of KuphDev/place_pink_duck_pruned_2nd.
Overview
A LeRobot robot manipulation dataset. Each frame holds the commanded action and observed observation.state joint positions; camera views are stored as videos in the original dataset.
Episodes: 128
Frames: 95140
Sampling rate: 30 fps
Tasks: 1
Split: a single train split
Robot: so_follower
Cameras (not uploaded): top
Schema (TsFile structure)… See the full description on the dataset page: https://huggingface.co/datasets/THULab/place_pink_duck_pruned_2nd.meow-10k
Dataset Card for Meow-10K
Meow-10K is a high-fidelity, synchronized quad-modal dataset comprising 10,000 feline samples. It is the primary training corpus for Meow-Omni 1, designed to facilitate deep intention reasoning in computational ethology.
Dataset Summary
Meow-10K provides the first large-scale training foundation for Multimodal Large Language Models (MLLMs) to learn the causal relationships between external behaviours and internal physiological states. By… See the full description on the dataset page: https://huggingface.co/datasets/Duckyle/meow-10k.DataSet_mix_duck_oct_cabduckdb-nsql-scoresemotion
Dataset Card for "emotion"
Dataset Summary
Emotion is a dataset of English Twitter messages with six basic emotions: anger, fear, joy, love, sadness, and surprise. For more detailed information please refer to the paper.
Supported Tasks and Leaderboards
More Information Needed
Languages
More Information Needed
Dataset Structure
Data Instances
An example looks as follows.
{
"text": "im feeling quite sad… See the full description on the dataset page: https://huggingface.co/datasets/Duckling1quack/emotion.cu-vla-exp3-datacommon-voice-ro-taggedDataSet_mix_duck_octduckdb-docbench
DocBench: A Synthetic DuckDB Text-to-SQL Benchmark
DocBench is a synthetic Text-to-SQL benchmark dataset consisting of 2430 question/sql pairs derived from the DuckDB documentation, specifically designed to probe language models for knowledge of DuckDB-specific SQL functionality.
The dataset covers functions, aggregates, operators, statements, keywords, and multi-keyword expressions available in DuckDB 1.1.3 and its default extensions.
Dataset Structure
Each example… See the full description on the dataset page: https://huggingface.co/datasets/motherduckdb/duckdb-docbench.duckjam-lp5-vaultcu-vla-exp5-datasql-console-prompt
SQL Console Text 2 SQL Prompt
GitHub Gist
Feedback is welcome 🤗. This prompt was based on performance from Qwen on the DuckDB NSQL Benchmark, common dataset types and tasks typical for exploring HF Datasets.
This is the prompt used for the Text2SQL inside the SQL Console on Datasets.
Example Table Context
For the {table_context} we use the SQL DDL CREATE TABLE statement.
CREATE TABLE datasets (
"_id" VARCHAR,
"id" VARCHAR,
"author" VARCHAR… See the full description on the dataset page: https://huggingface.co/datasets/duckdb-nsql-hub/sql-console-prompt.duckdb-qa-v2rubber_duck_extended_tokensDuckietown-Multiclass-Semantic-Segmentation-Dataset
Multiclass Semantic Segmentation Duckietown Dataset
A dataset of multiclass semantic segmentation image annotations for the first 250 images of the "Duckietown Object Detection Dataset".
Raw Image
Segmentated Image
Semantic Classes
This dataset defines 8 semantic classes (7 distinct classes + implicit background class):
Class
XML Label
Description
Color (RGB)
Ego Lane
Ego Lane
The lane the agent is supposed to be driving in (default right-hand… See the full description on the dataset page: https://huggingface.co/datasets/hamnaanaa/Duckietown-Multiclass-Semantic-Segmentation-Dataset.AbstractEdit-Bench
Abstract Image Editing Benchmark
A benchmark for evaluating instruction-following image editing models on abstract
(open-ended) vs explicit (fully specified) editing instructions. Context images are
drawn from Open Images V7
(validation split).
Each item pairs an abstract edit instruction
(e.g. "Pack these houses into boxes for shipping") with an explicit counterpart that
lists every atomic change required. Items span four domains (Physical, Logical, Social,
Emotional) and… See the full description on the dataset page: https://huggingface.co/datasets/DucktorV/AbstractEdit-Bench.
