datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gw-tti-playground-queuepythia-training-metrics Dataset for storing training metrics of pythia modelsMJHQ-30K
MJHQ-30K Benchmark
Model
Overall FID
SDXL-1-0-refiner
9.55
playground-v2-1024px-aesthetic
7.07
We introduce a new benchmark, MJHQ-30K, for automatic evaluation of a model’s aesthetic quality. The benchmark computes FID on a high-quality dataset to gauge aesthetic quality.
We curate the high-quality dataset from Midjourney with 10 common categories, each category with 3K samples. Following common practice, we use aesthetic score and CLIP score to ensure high image… See the full description on the dataset page: https://huggingface.co/datasets/playgroundai/MJHQ-30K.pythia-training-evalspatent-iq-playgroundGiant-in-the-Playground-RP
Giant in the Playground (roleplaying subforums only)
A semi-cleaned, processed version of the raw files uploaded elsewhere of the roleplaying sections (Play-by-Post Games) from Giant in the Playground, scraped on January 2025. I've made an effort to preserve as much as possible of the original HTML while simplifying and converting it to HTML5 where possible and cleaning it, with the notable exception of converting HTML linebreaks into newlines.
I'm almost directly using these files… See the full description on the dataset page: https://huggingface.co/datasets/lemonilia/Giant-in-the-Playground-RP.phospho-playground-mono
phospho-playground-mono
This dataset was generated using the phospho cli
More information on robots.phospho.ai.
This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS.
figma-Playground-Inference-for-PRO-s-WebsitePoisoning_Resilient_Federated_Learning_Playground
FL Security Experiment Results
This repository contains the experiment outputs used for the FL Security / FLPoison federated learning poisoning benchmark. The archive is intended for readers who want to inspect the raw training logs, reuse the aggregated curves and tables, or reproduce the paper figures without rerunning the full Compute Canada workload.
The uploaded artifact is:
exp_data.tar.gz # about 241 MB
After extraction, the archive keeps the original Compute Canada… See the full description on the dataset page: https://huggingface.co/datasets/FL-Security/Poisoning_Resilient_Federated_Learning_Playground.FSCM_Flood_playgroundMusiXQA
MusiXQA 🎵
MusiXQA is a multimodal dataset for evaluating and training music sheet understanding systems. Each data sample is composed of:
A scanned music sheet image (.png)
Its corresponding MIDI file (.mid)
A structured annotation (from metadata.json)
Question–Answer (QA) pairs targeting musical structure, semantics, and optical music recognition (OMR)
📂 Dataset Structure
MusiXQA/
├── images.tar # PNG files of music sheets (e.g., 0000000.png)… See the full description on the dataset page: https://huggingface.co/datasets/puar-playground/MusiXQA.gas-centroids
GAS Indexing Artifacts
Dataset Description
This dataset contains pre-computed deterministic centroids and associated geometric metadata generated using our GAS (Geometry-Aware Selection) algorithm.
These artifacts are designed to benchmark Approximate Nearest Neighbor (ANN) search performance in privacy-preserving or dynamic vector database environments.
Purpose
To serve as a standardized benchmark resource for evaluating the efficiency and recall of vector… See the full description on the dataset page: https://huggingface.co/datasets/cryptolab-playground/gas-centroids.pythia-pile-presampled0HBloomberg-Financial-News-embedding-gemma-300m
Bloomberg Financial News Embeddings for Vector Database Benchmarking
Dataset Description
This dataset contains pre-computed embeddings of Bloomberg financial news articles, designed for evaluating vector database performance. The embeddings are generated using Google's EmbeddingGemma-300M model.
Purpose
Benchmark dataset for evaluating vector database performance on financial news domain, specifically designed for use with VectorDBBench.
Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/cryptolab-playground/Bloomberg-Financial-News-embedding-gemma-300m.PlaygroundS6E8
Kaggle Playground Series S6E8 Dataset
This dataset contains the training and validation data (with a sample submission set) from the Kaggle Playground Series Season 6, Episode 8 (S6E8)
It is uploaded to Hugging Face for easier access and use in machine learning experiments.
Dataset contents
The dataset includes:
train.csv - Training dataset containing the features and target variable.
validation.csv - Validation dataset for evaluating model performance during… See the full description on the dataset page: https://huggingface.co/datasets/cazyundee/PlaygroundS6E8.pubmed-arxiv-abstract-embedding-gemma-300m
PubMed & arXiv Abstract Embeddings for Vector Database Benchmarking
Dataset Description
This dataset contains pre-computed embeddings of scientific paper abstracts from PubMed and arXiv, designed for evaluating vector database performance. The embeddings are generated using Google's EmbeddingGemma-300M model.
Purpose
Benchmark dataset for evaluating vector database performance, specifically designed for use with VectorDBBench.
Dataset Summary
Total… See the full description on the dataset page: https://huggingface.co/datasets/cryptolab-playground/pubmed-arxiv-abstract-embedding-gemma-300m.7Gplaygroundocr-playground
build_dataset.py
Dataset Summary
A robotics dataset with pointcloud text modality, stored in parquet format.
Preprocessing & Augmentation
Preprocessing: curriculum
Augmentation: none
Splits & Sampling
Split strategy: temporal
Sampling: hard negative
Quality & Labeling
Quality filtering: moderate
Labeling: pseudo label
Files
build_dataset.py — main artifact of this repository
License… See the full description on the dataset page: https://huggingface.co/datasets/Blrmehta01/ocr-playground.5Bcrello-cap536L5GCapsBench
CapsBench
CapsBench is a captioning evaluation dataset designed to comprehensively assess the quality of the captions across 17 categories: general,
image type, text, color, position, relation, relative position, entity, entity size, entity shape, count, emotion, blur, image artifacts,
proper noun (world knowledge), color palette, and color grading.
There are 200 images and 2471 questions for them, resulting in 12 questions per image on average. Images represent a wide variety of… See the full description on the dataset page: https://huggingface.co/datasets/playgroundai/CapsBench.rete-playground4ZVisR-Bench
Testing Code:
The GitHub repo for testing code: VisR-Bench
Data Download
This is the page images of all documents of the VisR-Bench dataset.
git lfs install
git clone https://huggingface.co/datasets/puar-playground/VisR-Bench
The code above will download the VisR-Bench folder, which is required for testing.
Reference
@misc{chen2025visrbenchempiricalstudyvisual,
title={VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for… See the full description on the dataset page: https://huggingface.co/datasets/puar-playground/VisR-Bench.3N
