CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01albertklorer /safedocs-1M-muse-spark-1.3-judged SafeDocs: Muse Spark 1.3 judge annotations Incrementally published, one complete shard per commit. All original source columns, images, complete Paddle JSON, rows and row order are preserved. No language or quality filtering. New columns: judge_verdict (PERFECT/ERROR), judge_reason, judge_status, and judge_error. Operational failures retain the original page with a null verdict and reason, status failed, and a diagnostic in judge_error; they are not OCR ERRORs. Direct Meta API… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/safedocs-1M-muse-spark-1.3-judged.tabular100K<n<1M0 likes12k downloads7d agoHugging Face02muse-bench /MUSE-News MUSE-News MUSE is a comprehensive machine unlearning evaluation benchmark that assesses six key properties for unlearned models: (1) no verbatim memorization, (2) no knowledge memorization, (3) no privacy leakage, (4) utility preservation on data not intended for removal, (5) scalability with respect to the size of removal requests, and (6) sustainability over sequential unlearning requests. MUSE focuses on two types of textual data that commonly require unlearning: news articles… See the full description on the dataset page: https://huggingface.co/datasets/muse-bench/MUSE-News.text10K<n<100K4 likes6.2k downloads2y agoHugging Face03muse-bench /MUSE-Books MUSE-Books MUSE is a comprehensive machine unlearning evaluation benchmark that assesses six key properties for unlearned models: (1) no verbatim memorization, (2) no knowledge memorization, (3) no privacy leakage, (4) utility preservation on data not intended for removal, (5) scalability with respect to the size of removal requests, and (6) sustainability over sequential unlearning requests. MUSE focuses on two types of textual data that commonly require unlearning: news articles… See the full description on the dataset page: https://huggingface.co/datasets/muse-bench/MUSE-Books.textn<1K3 likes4.3k downloads2y agoHugging Face04muset-ai /MPIE-Bench MPIE-Bench Official 2,500-sample test set for multi-person interaction-aware image editing evaluation. GitHub (code + protocol): https://github.com/AnnLin0628/mpie-bench Org: muset-ai Dataset Viewer The default config (default / test) is one row per evaluation sample: Column Meaning cat Interaction category (filter / group by this) gt Held-out ground-truth image prompt Edit instruction ref_paths Reference image paths under images/ sample_id… See the full description on the dataset page: https://huggingface.co/datasets/muset-ai/MPIE-Bench.imageimage-to-image1K<n<10K0 likes1.2k downloads2mo agoHugging Face05miccull /met_museumimage100K<n<1M10 likes500 downloads4y agoHugging Face06jiahaomei /MUSE-VA MUSE-VA Dataset English | 中文 MUSE-VA (Multimodal MUSic Emotion Dataset with Balanced VA) is a large-scale multimodal music emotion dataset designed for music emotion understanding, emotion-controllable music generation, and cross-modal affective modeling. The dataset starts from target coordinates sampled in the continuous Valence-Arousal (VA) space and uses a five-stage LLM agent pipeline with affective and musical knowledge injection to construct music, text, images, and… See the full description on the dataset page: https://huggingface.co/datasets/jiahaomei/MUSE-VA.audio1K<n<10K1 likes461 downloads2mo agoHugging Face07mteb /MUSE-VA-A2Iaudio1K<n<10K0 likes319 downloads1mo agoHugging Face08mteb /MUSE-VA-I2Aaudio1K<n<10K0 likes265 downloads1mo agoHugging Face09Satgoy152 /Muse-Glimmer-SWE-Gym-2k Muse-Glimmer-SWE-Gym-2k Agentic coding traces from meta-models/Muse-Glimmer-30B, recorded for training a speculative-decoding drafter. 1,981 mini-swe-agent trajectories over SWE-Gym and SWE-bench-extra instances, and the 159,999 individual chat-completion calls behind them. Configs Config Rows Size What it is train 1,981 57 MB One row per trajectory: the full conversation as messages. raw 159,999 2.7 GB One row per recorded API call: request and… See the full description on the dataset page: https://huggingface.co/datasets/Satgoy152/Muse-Glimmer-SWE-Gym-2k.tabulartext-generation100K<n<1M2 likes257 downloads23d agoHugging Face10ddrg /MUSES MUSES: a benchmark for Marked Unevenly Spaced Event Sequences MUSES, a benchmark for Marked Unevenly Spaced Event Sequences, is a collection of unevenly spaced time series datasets from various domains, containing marked events for training and evaluating prediction approaches. Languages All the columns and classes (when textual) in MUSES are in English (BCP-47 en) Dataset Structure Data Instances All datasets formatting follows… See the full description on the dataset page: https://huggingface.co/datasets/ddrg/MUSES.documenttime-series-forecasting1M<n<10M4 likes197 downloads5mo agoHugging Face11sentence-transformers /parallel-sentences-muse Dataset Card for Parallel Sentences - MUSE This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. Most of the sentences originate from the OPUS website. In particular, this dataset contains the MUSE dataset. Related Datasets The following datasets are also a part of the Parallel Sentences collection: parallel-sentences-europarl parallel-sentences-global-voices parallel-sentences-muse… See the full description on the dataset page: https://huggingface.co/datasets/sentence-transformers/parallel-sentences-muse.textfeature-extraction1M<n<10M0 likes183 downloads2y agoHugging Face12Muse-Ltd /UncertaintyGym UncertaintyGym A Standardized Benchmark for LLM Epistemic Calibration & Uncertainty Expression Abstract UncertaintyGym evaluates whether language models recognize the boundaries of their knowledge. Rather than assessing purely factual recall, UncertaintyGym measures how reliably an LLM explicitly declares uncertainty ("I don't know"), requests necessary disambiguating context, and rejects false premises without hallucinating. Benchmark Taxonomy… See the full description on the dataset page: https://huggingface.co/datasets/Muse-Ltd/UncertaintyGym.textquestion-answering1K<n<10K6 likes149 downloads1mo agoHugging Face13albertklorer /safedocs-1M-muse-spark-1.3-first3 SafeDocs first three shards: Muse Spark 1.3 Source: albertklorer/safedocs-1M, revision 87faff9053aa50c745f1359bef3592219ccb8c8b. PaddleOCR-VL 1.6 teacher labels are compared to original pages using the existing side-by-side renderer and binary Muse Spark 1.3 contributor judge. Quality verdicts are only PERFECT or ERROR, with no quality reason. Operational failures have no verdict. These are model labels, not human ground truth. Native Paddle block list order is preserved. A… See the full description on the dataset page: https://huggingface.co/datasets/albertklorer/safedocs-1M-muse-spark-1.3-first3.tabularn<1K0 likes125 downloads10d agoHugging Face14Wissam42 /MUSE-VA-A2Iaudio1K<n<10K0 likes122 downloads1mo agoHugging Face15tamarsonha /MUSE-News-Train MUSE-News-Train This dataset is a simple merger of the pretraining data from the original MUSE-News dataset. Dataset Details Dataset Sources [optional] Repository: https://huggingface.co/datasets/muse-bench/MUSE-News Paper: https://arxiv.org/pdf/2407.06460 Dataset Creation To create this dataset, we simply started from the muse-bench dataset and selected the train subset. Then, by merging the retain1 and retain2 splits we get the actual retain… See the full description on the dataset page: https://huggingface.co/datasets/tamarsonha/MUSE-News-Train.text10K<n<100K1 likes100 downloads1y agoHugging Face16tamarsonha /MUSE-Books-Train MUSE-Books-Train This dataset is a simple merger of the pretraining data from the original MUSE-Books dataset. Dataset Details Dataset Sources [optional] Repository: https://huggingface.co/datasets/muse-bench/MUSE-Books Paper: https://arxiv.org/pdf/2407.06460 Dataset Creation To create this dataset, we simply started from the muse-bench dataset and selected the train subset. Then, by merging the retain1 and retain2 splits we get the actual… See the full description on the dataset page: https://huggingface.co/datasets/tamarsonha/MUSE-Books-Train.textn<1K0 likes61 downloads1y agoHugging Face17Satgoy152 /Muse-Glimmer-Terminal-Bench-Eval Muse Glimmer — Terminal-Bench eval traces This is an EVALUATION set. Do not train on it. These traces measure baseline speculative-decoding behaviour (acceptance length, draft acceptance rate, throughput) for the Muse Glimmer speculator project on agentic coding work, so that a fine-tuned speculator can be compared against them later. Training data for that project is SWE-Gym and is deliberately repo-disjoint from Terminal-Bench: it excludes every repository referenced by any of… See the full description on the dataset page: https://huggingface.co/datasets/Satgoy152/Muse-Glimmer-Terminal-Bench-Eval.tabular1K<n<10K0 likes59 downloads23d agoHugging Face18anonymous111111111 /MUSE-Bench MUSE-Bench: Memory Utilization Evaluation Benchmark Official dataset for the paper "Beyond Memorization: Benchmarking Memory Utilization in Conversational LLM Agents." Anonymous release. This repository is an anonymized copy provided for double-blind peer review. Author and affiliation information is withheld until the review process concludes. Motivation LLM agents increasingly rely on persistent cross-session memory to support long-horizon and personalized… See the full description on the dataset page: https://huggingface.co/datasets/anonymous111111111/MUSE-Bench.textquestion-answeringn<1K0 likes58 downloads24d agoHugging Face19ohsuz /muse_textbooks_debate_onlytext100K<n<1M1 likes46 downloads2y agoHugging Face20gaodrew /met-museum-no-imagesLiterally @miccull's dataset minus the images Original source is the Met museum's open dataset of public domain works in their collection. https://console.cloud.google.com/marketplace/product/the-metropolitan-museum-of-art/the-met-public-domain-art-works?project=smartmaps-423802 text100K<n<1M0 likes40 downloads2y agoHugging Face21Muse-Fileread /aave_transactions_blockchain_copytabular1M<n<10M0 likes39 downloads2y agoHugging Face22Jaehun /chart-museum-samplesimage1K<n<10K0 likes35 downloads10mo agoHugging Face23chan030609 /MUSE-News-2text10K<n<100K0 likes31 downloads2y agoHugging Face24gowitheflow /parallel-muse-deduplicatedtext100K<n<1M0 likes31 downloads2y agoHugging Face25Leon299 /muse_mucodec_chordtext100K<n<1M0 likes28 downloads7mo agoHugging Face26alita9 /muse-sarcasm-explanation MuSe: Multimodal Sarcasm Explanation (Reformatted) This repository provides a Hugging Face-compatible version of the MuSe (MORE) dataset. Modifications in this version To make the dataset easier to use with the datasets library, the following changes were made: Unified Schema: Merged separate OCR and Non-OCR files into a single test split. Metadata Flags: Added an is_ocr (boolean) column to distinguish between image types. Image Integration: Converted image paths into a… See the full description on the dataset page: https://huggingface.co/datasets/alita9/muse-sarcasm-explanation.image1K<n<10K0 likes27 downloads9mo agoHugging Face27result-muse256-muse512-wuerst-sdv15 /6971f242 Dataset Card for "6971f242" More Information needed textn<1K0 likes26 downloads3y agoHugging Face28emi429 /mnist_muse2 Dataset Card for "mnist_muse2" More Information needed 10K<n<100K0 likes26 downloads3y agoHugging Face29diffusers-parti-prompts /muse512 Dataset Card for "muse_512" ```py from PIL import Image import torch from muse import PipelineMuse, MaskGiTUViT from datasets import Dataset, Features from datasets import Image as ImageFeature from datasets import Value, load_dataset device = "cuda" if torch.cuda.is_available() else "cpu" pipe = PipelineMuse.from_pretrained( transformer_path="valhalla/research-run", text_encoder_path="openMUSE/clip-vit-large-patch14-text-enc"… See the full description on the dataset page: https://huggingface.co/datasets/diffusers-parti-prompts/muse512.image1K<n<10K0 likes24 downloads3y agoHugging Face30electricsheepafrica /africa-cote-d-ivoire-personnel-travaillant-dans-les-musees-et-institutions-assi-cb8f3e48 Personnel Travaillant Dans Les Musees Et Institutions Assi | Africa (Cote d'Ivoire DataFair) 125 rows - 1 Africa country/area - 2003-2006 - 1 indicator - Engineered by Electric Sheep Africa TL;DR This dataset contains 125 rows from Cote d'Ivoire DataFair, covering Personnel Travaillant Dans Les Musees Et Institutions Assi. It is published as ML-ready Parquet with consistent Hugging Face metadata, source provenance, and analysis-friendly loading examples.… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepafrica/africa-cote-d-ivoire-personnel-travaillant-dans-les-musees-et-institutions-assi-cb8f3e48.tabulartabular-regressionn<1K0 likes23 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.