CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Rolsouze /gw-tti-playground-queueimagen<1K0 likes13k downloads2mo agoHugging Face02pretraining-playground /pythia-training-metrics Dataset for storing training metrics of pythia models0 likes2.1k downloads9mo agoHugging Face03playgroundai /MJHQ-30K MJHQ-30K Benchmark Model Overall FID SDXL-1-0-refiner 9.55 playground-v2-1024px-aesthetic 7.07 We introduce a new benchmark, MJHQ-30K, for automatic evaluation of a model’s aesthetic quality. The benchmark computes FID on a high-quality dataset to gauge aesthetic quality. We curate the high-quality dataset from Midjourney with 10 common categories, each category with 3K samples. Following common practice, we use aesthetic score and CLIP score to ensure high image… See the full description on the dataset page: https://huggingface.co/datasets/playgroundai/MJHQ-30K.imagetext-to-image10K<n<100K65 likes1.3k downloads2y agoHugging Face04pretraining-playground /pythia-training-evals0 likes961 downloads2y agoHugging Face05ardae1 /patent-iq-playground0 likes527 downloads7mo agoHugging Face06lemonilia /Giant-in-the-Playground-RP Giant in the Playground (roleplaying subforums only) A semi-cleaned, processed version of the raw files uploaded elsewhere of the roleplaying sections (Play-by-Post Games) from Giant in the Playground, scraped on January 2025. I've made an effort to preserve as much as possible of the original HTML while simplifying and converting it to HTML5 where possible and cleaning it, with the notable exception of converting HTML linebreaks into newlines. I'm almost directly using these files… See the full description on the dataset page: https://huggingface.co/datasets/lemonilia/Giant-in-the-Playground-RP.tabular10K<n<100K2 likes321 downloads2y agoHugging Face07PLB /phospho-playground-mono phospho-playground-mono This dataset was generated using the phospho cli More information on robots.phospho.ai. This dataset contains a series of episodes recorded with a robot and multiple cameras. It can be directly used to train a policy using imitation learning. It's compatible with LeRobot and RLDS. tabularrobotics10K<n<100K3 likes282 downloads1y agoHugging Face08huggingface /figma-Playground-Inference-for-PRO-s-Website1 likes195 downloads3y agoHugging Face09FL-Security /Poisoning_Resilient_Federated_Learning_Playground FL Security Experiment Results This repository contains the experiment outputs used for the FL Security / FLPoison federated learning poisoning benchmark. The archive is intended for readers who want to inspect the raw training logs, reuse the aggregated curves and tables, or reproduce the paper figures without rerunning the full Compute Canada workload. The uploaded artifact is: exp_data.tar.gz # about 241 MB After extraction, the archive keeps the original Compute Canada… See the full description on the dataset page: https://huggingface.co/datasets/FL-Security/Poisoning_Resilient_Federated_Learning_Playground.texttabular-classification1M<n<10M0 likes178 downloads4mo agoHugging Face10ParkSY /FSCM_Flood_playgroundimage1K<n<10K0 likes155 downloads1y agoHugging Face11puar-playground /MusiXQA MusiXQA 🎵 MusiXQA is a multimodal dataset for evaluating and training music sheet understanding systems. Each data sample is composed of: A scanned music sheet image (.png) Its corresponding MIDI file (.mid) A structured annotation (from metadata.json) Question–Answer (QA) pairs targeting musical structure, semantics, and optical music recognition (OMR) 📂 Dataset Structure MusiXQA/ ├── images.tar # PNG files of music sheets (e.g., 0000000.png)… See the full description on the dataset page: https://huggingface.co/datasets/puar-playground/MusiXQA.3 likes155 downloads4mo agoHugging Face12cryptolab-playground /gas-centroids GAS Indexing Artifacts Dataset Description This dataset contains pre-computed deterministic centroids and associated geometric metadata generated using our GAS (Geometry-Aware Selection) algorithm. These artifacts are designed to benchmark Approximate Nearest Neighbor (ANN) search performance in privacy-preserving or dynamic vector database environments. Purpose To serve as a standardized benchmark resource for evaluating the efficiency and recall of vector… See the full description on the dataset page: https://huggingface.co/datasets/cryptolab-playground/gas-centroids.text-retrieval0 likes144 downloads8mo agoHugging Face13pretraining-playground /pythia-pile-presampled1M<n<10M0 likes133 downloads2y agoHugging Face14chienyu-playground /0Haudio0 likes122 downloads1y agoHugging Face15cryptolab-playground /Bloomberg-Financial-News-embedding-gemma-300m Bloomberg Financial News Embeddings for Vector Database Benchmarking Dataset Description This dataset contains pre-computed embeddings of Bloomberg financial news articles, designed for evaluating vector database performance. The embeddings are generated using Google's EmbeddingGemma-300M model. Purpose Benchmark dataset for evaluating vector database performance on financial news domain, specifically designed for use with VectorDBBench. Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/cryptolab-playground/Bloomberg-Financial-News-embedding-gemma-300m.text-retrieval100K<n<1M0 likes114 downloads10mo agoHugging Face16cazyundee /PlaygroundS6E8 Kaggle Playground Series S6E8 Dataset This dataset contains the training and validation data (with a sample submission set) from the Kaggle Playground Series Season 6, Episode 8 (S6E8) It is uploaded to Hugging Face for easier access and use in machine learning experiments. Dataset contents The dataset includes: train.csv - Training dataset containing the features and target variable. validation.csv - Validation dataset for evaluating model performance during… See the full description on the dataset page: https://huggingface.co/datasets/cazyundee/PlaygroundS6E8.tabular100K<n<1M0 likes88 downloads1mo agoHugging Face17cryptolab-playground /pubmed-arxiv-abstract-embedding-gemma-300m PubMed & arXiv Abstract Embeddings for Vector Database Benchmarking Dataset Description This dataset contains pre-computed embeddings of scientific paper abstracts from PubMed and arXiv, designed for evaluating vector database performance. The embeddings are generated using Google's EmbeddingGemma-300M model. Purpose Benchmark dataset for evaluating vector database performance, specifically designed for use with VectorDBBench. Dataset Summary Total… See the full description on the dataset page: https://huggingface.co/datasets/cryptolab-playground/pubmed-arxiv-abstract-embedding-gemma-300m.text-retrieval100K<n<1M0 likes82 downloads10mo agoHugging Face18chienyu-playground /7G0 likes79 downloads1y agoHugging Face19akcit-rl /playground0 likes78 downloads2mo agoHugging Face20Blrmehta01 /ocr-playground build_dataset.py Dataset Summary A robotics dataset with pointcloud text modality, stored in parquet format. Preprocessing & Augmentation Preprocessing: curriculum Augmentation: none Splits & Sampling Split strategy: temporal Sampling: hard negative Quality & Labeling Quality filtering: moderate Labeling: pseudo label Files build_dataset.py — main artifact of this repository License… See the full description on the dataset page: https://huggingface.co/datasets/Blrmehta01/ocr-playground.0 likes77 downloads27d agoHugging Face21chienyu-playground /5Baudio0 likes70 downloads1y agoHugging Face22puar-playground /crello-capimage10K<n<100K0 likes68 downloads2y agoHugging Face23chienyu-playground /530 likes61 downloads1y agoHugging Face24chienyu-playground /6Laudio0 likes59 downloads1y agoHugging Face25chienyu-playground /5Gaudio0 likes53 downloads1y agoHugging Face26playgroundai /CapsBench CapsBench CapsBench is a captioning evaluation dataset designed to comprehensively assess the quality of the captions across 17 categories: general, image type, text, color, position, relation, relative position, entity, entity size, entity shape, count, emotion, blur, image artifacts, proper noun (world knowledge), color palette, and color grading. There are 200 images and 2471 questions for them, resulting in 12 questions per image on average. Images represent a wide variety of… See the full description on the dataset page: https://huggingface.co/datasets/playgroundai/CapsBench.imagen<1K13 likes51 downloads2y agoHugging Face27katospiegel /rete-playground0 likes50 downloads3mo agoHugging Face28chienyu-playground /4Zaudio0 likes45 downloads1y agoHugging Face29puar-playground /VisR-Bench Testing Code: The GitHub repo for testing code: VisR-Bench Data Download This is the page images of all documents of the VisR-Bench dataset. git lfs install git clone https://huggingface.co/datasets/puar-playground/VisR-Bench The code above will download the VisR-Bench folder, which is required for testing. Reference @misc{chen2025visrbenchempiricalstudyvisual, title={VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for… See the full description on the dataset page: https://huggingface.co/datasets/puar-playground/VisR-Bench.image1K<n<10K2 likes44 downloads1y agoHugging Face30chienyu-playground /3Naudio0 likes43 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.