datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
objaverse-1k
DX.GL Objaverse-1K: Dense Multi-View Datasets for 3D Vision
1026 objects × 196 views × 1024x1024 resolution × 6 modalities — ready for nerfstudio out of the box.
Dense multi-view datasets rendered from curated Objaverse 3D models via DX.GL. Each object includes calibrated camera poses, depth maps (8-bit + 16-bit), normal maps, binary masks, and point clouds — suitable for novel view synthesis, 3D reconstruction, monocular depth/normal estimation, segmentation, and more.… See the full description on the dataset page: https://huggingface.co/datasets/dxgl/objaverse-1k.GRAZPEDWRI-DX_SMALLR52UBnormalsprite-dx-data
SpriteDX — Sprite Matting and Animation Annotations
Explore sprite frames with matching mattes, foreground images, and transparent cutouts; inspect human annotations for animation loops and shot boundaries. This collection comes from experiments on the SpriteDX project by Sprited.
Start with the matting subset in the viewer above. Each row shows the original frame and its three paired representations. Use the subset selector for loops or scene_boundaries.
Subset
Rows
What… See the full description on the dataset page: https://huggingface.co/datasets/sprited/sprite-dx-data.R8TinyStories-Multilingual
Novelist: TinyStories Multilingual Edition
Dataset Summary
The TinyStories Multilingual Edition is a high-fidelity synthetic dataset of short, child-safe fiction designed to stress-test literary consistency, emotional warmth, and multilingual fluency in small models. Derived from the broader Novelist ecosystem, this subset focuses on narrative simplicity paired with complex moral and social themes.
The dataset contains 15,688 high-quality stories across 28 languages. Each… See the full description on the dataset page: https://huggingface.co/datasets/Dxniz/TinyStories-Multilingual.DxBench
Citation
@misc{chen2024codinterpretablemedicalagent,
title={CoD, Towards an Interpretable Medical Agent using Chain of Diagnosis},
author={Junying Chen and Chi Gui and Anningzhe Gao and Ke Ji and Xidong Wang and Xiang Wan and Benyou Wang},
year={2024},
eprint={2407.13301},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2407.13301},
}
grazpedwri_dxohsumedHuman-Humanoid-4Dmultiview-datasets
DX.GL Multi-View Datasets for NeRF & 3D Gaussian Splatting
Multi-view training datasets rendered from CC0 3D models via DX.GL. Each dataset includes calibrated camera poses, depth maps, normal maps, binary masks, and point clouds — ready for nerfstudio out of the box.
10 objects × 196 views × 1024×1024 resolution × full sphere coverage.
Quick Start
# Download a dataset (Apple, 196 views, 1024x1024)
wget https://dx.gl/api/v/EJbs8npt2RVM/vCHDLxjWG65d/dataset -O… See the full description on the dataset page: https://huggingface.co/datasets/dxgl/multiview-datasets.rag-dx
RAG-Dx: a diagnostic benchmark for retrieval
This dataset is for evaluation. It is not training data and should not be used to train
or fine-tune models.
Most retrieval benchmarks give you a number. A number tells you that something is wrong,
not what. RAG-Dx reports how much a retrieval stack degrades on each of eight specific
failure modes, so the output points at a fix.
Code, harness and reproduction scripts: https://github.com/chakshu-dhannawat/rag-dx
What is in… See the full description on the dataset page: https://huggingface.co/datasets/Chakshu123/rag-dx.arxiv-semantic-scholar
📚 arxiv-semantic-scholar
A paper metadata dataset covering every paper on arXiv, with title, authors, full abstract, download link, and submission history.
Each row is additionally enriched by a call to the Semantic Scholar API, adding citation counts and venue information.
🔗 Sources
arXiv metadata: https://www.kaggle.com/datasets/Cornell-University/arxiv (CC0)
Semantic Scholar Academic Graph: https://api.semanticscholar.org/ (ODC-BY)
Snapshot taken 2026-07-04.… See the full description on the dataset page: https://huggingface.co/datasets/dx2102/arxiv-semantic-scholar.dxap-research-workflow-excerpts
DXAP historical workflow excerpts and selection aggregates
Three detailed historical case studies: 29 parent/reconciliation event rows, 18 sanitized tool-request/response projections, five starting-state records, two research-child event records and nine proposed scenario questions. It also includes the complete 15-row selection-rank table already published with arXiv:2609.05663v1. These are different views of the same selected material, not independent sample counts to add… See the full description on the dataset page: https://huggingface.co/datasets/DXRG/dxap-research-workflow-excerpts.Novelist
Dataset Card for Novelist
Dataset Summary
Novelist is a synthetic creative-writing and narrative-reasoning dataset designed for long-context fiction systems, scene planners, continuity-aware story models, multilingual literary translation, and child-safe TinyStories generation. The dataset mixes direct prose, explicit reasoning traces, quality-only judge outputs, full-book artifacts, and multilingual translation outputs inside a single narrative training ecosystem.
This… See the full description on the dataset page: https://huggingface.co/datasets/Dxniz/Novelist.20ngdx-terminal-pro-research-aggregates
DX Terminal Pro research aggregates
Nine published research records from DX Research Group's Terminal Pro work. These are small aggregate evidence tables. They contain no participant-level decision logs or training trajectories.
The two configurations have different units of analysis:
Configuration
Records
What a row represents
market_behavior
4
A reported event or token-window aggregate from the bounded 21-day real-capital Terminal Pro deployment… See the full description on the dataset page: https://huggingface.co/datasets/DXRG/dx-terminal-pro-research-aggregates.ada_dx_evaldx7-patches-and-prompts
Yamaha DX7 Synthesizer Patches with AI-Generated Prompts
Dataset Description
This is a comprehensive, multi-task dataset designed for fine-tuning language models to understand and generate synthesizer patches for the Yamaha DX7.
The dataset contains over 20,000 examples across three distinct but related tasks, making it ideal for creating models that can not only generate patches but also understand and reason about their structure and validity.
How the Data Was… See the full description on the dataset page: https://huggingface.co/datasets/ccerati/dx7-patches-and-prompts.MRpega-constellation-dx-sft
Pega Constellation DX Components — Training Dataset
This dataset teaches a code model how to write Pega Constellation DX components — the custom React/TypeScript components that extend the Pega Platform UI.
It comes from the open-source constellation-ui-gallery repo, which Pega itself maintains as a reference for DX component authors. We pinned a specific commit so the dataset is reproducible.
What is in it
55 Pega DX components, each broken into 9 different… See the full description on the dataset page: https://huggingface.co/datasets/WeekendNoobs/pega-constellation-dx-sft.imgui_dx11dxf-gears-unstructured
Gear Design Dataset
Dataset Overview
The Gear Design Dataset contains structured data that includes gear design specifications and their corresponding DXF (Drawing Exchange Format) files. The dataset is intended for training and evaluating machine learning models focused on generative design and CAD (Computer-Aided Design) tasks. The dataset is split into three parts: training, validation, and test sets.
Each entry in the dataset represents a specific gear type and… See the full description on the dataset page: https://huggingface.co/datasets/mvrdock/dxf-gears-unstructured.LLM-Structure-Performance-Dataset
From Parameters to Performance: A Data-Driven Study on LLM Structure and Development
This dataset is the official companion to the paper "From Parameters to Performance: A Data-Driven Study on LLM Structure and Development". It provides a comprehensive collection of structural configurations and performance metrics for a wide range of open-source Large Language Models (LLMs), enabling data-driven research on how structural choices impact model performance.
Paper:… See the full description on the dataset page: https://huggingface.co/datasets/DX0369/LLM-Structure-Performance-Dataset.260807_test_traing_demo_13
Task_test3_260807_test_3_MCAP
Created with Cyclo Intelligence by ROBOTIS.
PANORAMADataset link is updated. https://huggingface.co/datasets/LG-AI-Research/PANORAMA
Novelist-CoT
Novelist Announcement.
Check out the new and biggest Novelist project [https://huggingface.co/datasets/Dxniz/Novelist]. With 5 thinking modes, agentic writing and long book support. Now only limitation is your imagination.
Novelist-CoT Creative Writing Dataset
Overview
Novelist-CoT is a long-form creative writing dataset designed for supervised fine-tuning and style-focused narrative generation.The dataset consolidates multiple generation pipelines into a… See the full description on the dataset page: https://huggingface.co/datasets/Dxniz/Novelist-CoT.novelist-cot-writing-raw-v1
Novelist: Human-Like Creative Writing Dataset (RAW)
This dataset is designed to train LLMs in high-quality creative writing. It focuses on narrative depth, coherent world-building, and logical character psychology.
The data was generated using DeepSeek-R1.
Dataset Overview
We focused on Quality over Quantity. The goal was to move away from generic "AI slop" and create text that feels grounded and intentional.
Total Tokens: ~29.4 Million
Total Examples: 3,369
Format:… See the full description on the dataset page: https://huggingface.co/datasets/Dxniz/novelist-cot-writing-raw-v1.dx-cluster-spots
📡 DX Cluster Spots
Real-time amateur radio DX spots powered by Spothole.app
🌐 Spothole.app
500+
Real-time Spots
34+
Countries
14
Bands Covered
6+
Sources
📡 Data Sources
📻
DX Clusters
📡
RBN
🏔️
POTA
⛰️
SOTA
🌲
WWFF
🌏
ZLOTA
💻 Quick Start
# Load dataset with HuggingFace
from datasets import load_dataset
ds = load_dataset("alphamate/dx-cluster-spots")
print(ds["train"][0])
Powered by Spothole.app by Ian Renton (MØTRT)
Created by… See the full description on the dataset page: https://huggingface.co/datasets/alphamate/dx-cluster-spots.
