CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01jasperai /monet Dataset Card for MONET MONET (Massive, Open, Non-redundant and Enriched Text-to-image dataset) is a large-scale, curated image-text dataset designed for training text-to-image (T2I) systems. It contains 103.8 million high-quality image-text pairs distilled from 2.9 billion raw pairs across nine heterogeneous open sources (6 real and 3 synthetic) through successive stages of safety filtering, domain-based filtering, exact and near-duplicate removal, and re-captioning with… See the full description on the dataset page: https://huggingface.co/datasets/jasperai/monet.imagetext-to-image100M<n<1B169 likes92k downloads3mo agoHugging Face02monology /pile-uncopyrighted Pile Uncopyrighted In response to authors demanding that LLMs stop using their works, here's a copy of The Pile with all copyrighted content removed.Please consider using this dataset to train your future LLMs, to respect authors and abide by copyright law.Creating an uncopyrighted version of a larger dataset (ie RedPajama) is planned, with no ETA. MethodologyCleaning was performed by removing everything from the Books3, BookCorpus2, OpenSubtitles, YTSubtitles, and OWT2… See the full description on the dataset page: https://huggingface.co/datasets/monology/pile-uncopyrighted.text100M<n<1B175 likes81k downloads3y agoHugging Face03catherinearnett /montok MonTok: A Suite of Monolingual Tokenizers This is a set of monolingual tokenizers for 98 languages. For each language, there are Unigram, BPE, and SuperBPE tokenizers, ranging in vocabulary size from around 6k to over 200k. Training Details Training Data All tokenizers are trained on samples of the data used to the train the Goldfish language models. The tokenizers were either trained on scaled or unscaled data. This refers to whether the models are trained on… See the full description on the dataset page: https://huggingface.co/datasets/catherinearnett/montok.4 likes25k downloads1y agoHugging Face04ccl4u /hermes-system-monitor-v24 likes16k downloads5m agoHugging Face05MonsterDie /VLN_Dataset_2This repository contains encrypted visual features for an ongoing academic research project. Decryption keys are managed internally for reproducibility. 3 likes13k downloads29d agoHugging Face06charge-benchmark /Charge-040_0040-Sparse-Mono0 likes10k downloads1y agoHugging Face07MONAI /testing_dataThis is testing data for use with MONAI unit tests. 1 likes8.4k downloads1y agoHugging Face08SII-Monument-Valley /CiQi-VQA CiQi-Agent Github | Model | Dataset | Paper CiQi-Agent: Aligning Vision, Tools and Aesthetics in Multimodal Agent for Cultural Reasoning on Chinese Porcelains Accepted to ECCV 2026 🎯 Overview CiQi-Agent has been accepted to ECCV 2026. We present CiQi-Agent, a domain-specific multimodal agent for antique Chinese porcelain connoisseurship. The project is designed to combine fine-grained visual perception, tool-augmented reasoning, and cultural-heritage knowledge… See the full description on the dataset page: https://huggingface.co/datasets/SII-Monument-Valley/CiQi-VQA.imagequestion-answering10K<n<100K6 likes8k downloads28d agoHugging Face09Isamu136 /Japanese-Political-Money-OCR-with-Qwen0 likes7.5k downloads1y agoHugging Face10omarkamali /wikipedia-monthly 🚀 Wikipedia Monthly Last updated: March 14, 2026, 21:06 UTC This repository provides monthly, multilingual dumps of Wikipedia, processed and prepared for easy use in NLP projects. 📊 Current Statistics Metric Current Export (March 2026) All Exports (Total) Languages 343 361 Articles 62.8M 62.8M Usage Load any language with a single line of code using 🤗 datasets. latest always refers to the most recent dump, while dated configs refer to… See the full description on the dataset page: https://huggingface.co/datasets/omarkamali/wikipedia-monthly.texttext-generation100M<n<1B81 likes6.2k downloads6mo agoHugging Face11MaLA-LM /mala-monolingual-filter MaLA Corpus: Massive Language Adaptation Corpus This is a cleaned version with some necessary data cleaning. Dataset Summary The MaLA Corpus (Massive Language Adaptation) is a comprehensive, multilingual dataset designed to support the continual pre-training of large language models. It covers 939 languages and consists of over 74 billion tokens, making it one of the largest datasets of its kind. With a focus on improving the representation of low-resource… See the full description on the dataset page: https://huggingface.co/datasets/MaLA-LM/mala-monolingual-filter.text-generation3 likes5.3k downloads2mo agoHugging Face12persona-cartography /monorepo Persona Cartography — artifact monorepo Artifact store for the paper Persona Cartography: Charting Language Model Personality Traits in Weight Space (arXiv:2607.07916). Code: persona-cartography/persona-cartography. This is not a load_dataset-able dataset — it is a single shared repo holding every artifact the paper's pipeline produces: trained LoRA adapters, their training data, evaluation results, and the figures' source data. The paper's figure scripts hydrate from the paths… See the full description on the dataset page: https://huggingface.co/datasets/persona-cartography/monorepo.4 likes5.1k downloads26d agoHugging Face13KelvinKHan /monopoly-assetsimagen<1K0 likes5.1k downloads7mo agoHugging Face14badlogicgames /pi-mono Coding agent session traces for badlogicgames/pi-mono This dataset contains redacted coding agent session traces collected while working on https://github.com/badlogic/pi-mono.git. The traces were exported with pi-share-hf from a local pi workspace and filtered to keep only sessions that passed deterministic redaction and LLM review. Data description Each *.jsonl file is a redacted pi session. Sessions are stored as JSON Lines files where each line is a structured… See the full description on the dataset page: https://huggingface.co/datasets/badlogicgames/pi-mono.text-generation212 likes4.9k downloads6mo agoHugging Face15BangumiBase /mono Bangumi Image Base of Mono This is the image base of bangumi Mono, we detected 48 characters, 4375 images in total. The full dataset is here. Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability). Here is the characters' preview:… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/mono.image1K<n<10K0 likes4.3k downloads1y agoHugging Face16physicl /indoor-safety-hazard-detection-and-work-zone-monitoring Indoor Safety Hazard Detection & Work-Zone Monitoring Generated by datapack-import.ts This dataset mirrors public data-pack render outputs from Physicl. Each row represents one render view. The image column contains a stable URL to the primary render image uploaded under /data; image_path stores the relative repository path and data_commit_sha pins the Hugging Face dataset commit used by those URLs. Files are uploaded as downloaded unless optional PNG recompression is enabled by… See the full description on the dataset page: https://huggingface.co/datasets/physicl/indoor-safety-hazard-detection-and-work-zone-monitoring.imagen<1K0 likes4.3k downloads3mo agoHugging Face17Monash-University /monash_tsfMonash Time Series Forecasting Repository which contains 30+ datasets of related time series for global forecasting research. This repository includes both real-world and competition time series datasets covering varied domains.time-series-forecasting1K<n<10K60 likes4k downloads3y agoHugging Face18Stage-jh-monitor /qwen35-4b qwen35-4b Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.38203125 Action score: 0.4375 Valid samples: 320/320 tabularn<1K0 likes3.8k downloads13d agoHugging Face19Stage-jh-monitor /appworld-qwen35-4b-9b-s_signal_6-epoch4-iter1 appworld-qwen35-4b-9b-s_signal_6-epoch4-iter1 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.3953125 Action score: 0.446875 Valid samples: 320/320 tabularn<1K0 likes3.8k downloads13d agoHugging Face20Stage-jh-monitor /total-300-random-jh-epoch4 total-300-random-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.3890625 Action score: 0.440625 Valid samples: 320/320 tabularn<1K0 likes3.8k downloads13d agoHugging Face21Stage-jh-monitor /total-300-lambda00-s_signal_type6-jh-epoch4 total-300-lambda00-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.3875 Action score: 0.43125 Valid samples: 320/320 tabularn<1K0 likes3.8k downloads13d agoHugging Face22Stage-jh-monitor /total-300-lambda02-s_signal_type6-jh-epoch4 total-300-lambda02-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.4046875 Action score: 0.4140625 Valid samples: 320/320 tabularn<1K0 likes3.8k downloads13d agoHugging Face23Stage-jh-monitor /total-300-lambda05-s_signal_type6-jh-epoch4 total-300-lambda05-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.35703125 Action score: 0.4375 Valid samples: 320/320 tabularn<1K0 likes3.8k downloads13d agoHugging Face24Stage-jh-monitor /total-300-lambda10-s_signal_type6-jh-epoch4 total-300-lambda10-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.36640625 Action score: 0.41875 Valid samples: 320/320 tabularn<1K0 likes3.8k downloads13d agoHugging Face25Monster-Code /Pytorch-Code-10K Hot Coco Training Dataset A curated collection of 10,625 high-quality PyTorch and Transformers code examples with AI-generated captions. This dataset was specifically built for fine-tuning code-specialized language models like Qimi (Coming soon!) Dataset Description This dataset contains Python code snippets sourced from open-source repositories that utilize PyTorch or Hugging Face Transformers. Each sample includes: code: The raw Python source code (typically… See the full description on the dataset page: https://huggingface.co/datasets/Monster-Code/Pytorch-Code-10K.texttext-generation1K<n<10K1 likes3.8k downloads2mo agoHugging Face26Stage-jh-monitor /total-300-lambda08-s_signal_type6-jh-epoch4 total-300-lambda08-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.38046875 Action score: 0.4078125 Valid samples: 320/320 tabularn<1K0 likes3.8k downloads13d agoHugging Face27Stage-jh-monitor /total-300noapp-lambda02-s_signal_type6-jh-epoch4 total-300noapp-lambda02-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.36640625 Action score: 0.409375 Valid samples: 320/320 tabularn<1K0 likes3.8k downloads13d agoHugging Face28Stage-jh-monitor /total-300app-lambda02-s_signal_type6-jh-epoch4 total-300app-lambda02-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.3625 Action score: 0.4015625 Valid samples: 320/320 tabularn<1K0 likes3.8k downloads13d agoHugging Face29Stage-jh-monitor /total-131-lambda02-residual-s_signal_type6-jh-epoch4 total-131-lambda02-residual-s_signal_type6-jh-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.3765625 Action score: 0.4171875 Valid samples: 320/320 tabularn<1K0 likes3.8k downloads13d agoHugging Face30Stage-jh-monitor /total-300-lambda02-s_signal_type6-jh-retry-epoch4 total-300-lambda02-s_signal_type6-jh-retry-epoch4 Portable process-evaluation output. metadata.json is the lightweight source for aggregate results; the JSONL files are directly loadable; and artifacts.tar.gz losslessly preserves the original run directory. Reasoning score: 0.36953125 Action score: 0.3984375 Valid samples: 320/320 tabularn<1K0 likes3.8k downloads13d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.