CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AMLGentex /sweden_100K_difficulttabular10M<n<100M0 likes3.2k downloads1y agoHugging Face02science-of-finetuning /diffing-stats-gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04 Contains maximum activating examples for all the features of our crosscoder trained on gemma 2 2B layer 13 available here: https://huggingface.co/Butanium/gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04/blob/main/README.md base_examples.pt contains all the maximum examples of the feature on a subset of validation test of fineweb chat_examples.pt is the same but for lmsys chat data chat_base_examples.pt is a merge of the two above files. All files are of the type dict[int, list[tuple[float… See the full description on the dataset page: https://huggingface.co/datasets/science-of-finetuning/diffing-stats-gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04.tabular10K<n<100K0 likes200 downloads1y agoHugging Face03fassabilf /diffusion-mcqa-gen-pelatnas-2026 Which Prompt Made This? — Generated Edition Pelatnas IOAI 2026 · Task Diffusion MCQA (varian trajectory) Sebuah model text-to-image sedang bekerja. Di tengah prosesnya, gambar belum menjadi gambar — yang ada hanya latent ter-noise: tensor 4 × 64 × 64 berisi campuran struktur yang mulai muncul dan derau Gaussian. Kali ini kalimat itu harfiah. Latent yang kamu terima benar-benar diambil dari tengah proses generate: sebuah trajectory denoising DDIM 50 langkah dihentikan sejenak… See the full description on the dataset page: https://huggingface.co/datasets/fassabilf/diffusion-mcqa-gen-pelatnas-2026.tabularimage-classificationn<1K0 likes195 downloads2mo agoHugging Face04diffusers /benchmarks Welcome to 🤗 Diffusers Benchmarks! This is dataset where we keep track of the inference latency and memory information of the core models in the diffusers library. Currently, the core models are: Flux Wan LTX SDXL Note that we will continue to extend this list based on their usage. You can analyze the results in this demo. [!IMPORTANT] Instead of benchmarking the entire diffusion pipelines, we only benchmark the forward passes of the diffusion networks under different settings… See the full description on the dataset page: https://huggingface.co/datasets/diffusers/benchmarks.tabularn<1K15 likes171 downloads22h agoHugging Face05lime-nlp /DeepScaleR_Difficulty Difficulty Estimation on DeepScaleR We annotate the entire DeepScaleR dataset with a difficulty score based on the performance of the Qwen 2.5-MATH-7B model. This provides an adaptive signal for curriculum construction and model evaluation. DeepScaleR is a curated dataset of 40,000 reasoning-intensive problems used to train and evaluate reinforcement learning-based methods for large language models. Difficulty Scoring Method Difficulty scores are estimated using the… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/DeepScaleR_Difficulty.tabularreinforcement-learning1M<n<10M11 likes137 downloads1y agoHugging Face06leffff /Diffusion-Reward-Modeling-for-Text-Rendering-Dataset 🖼️ Text-to-Image Rendering Dataset A dataset of 14k text prompts for image generation with text rendering evaluation 📚 Dataset Overview This dataset contains 14,000 text prompts specifically designed for: Image generation with text rendering Evaluating text preservation in generated images Training diffusion models for better text rendering Each prompt comes with: Pre-extracted target text for rendering 5 Stable Diffusion 3 generated latents (70k total) Dual… See the full description on the dataset page: https://huggingface.co/datasets/leffff/Diffusion-Reward-Modeling-for-Text-Rendering-Dataset.tabulartext-to-image10K<n<100K8 likes86 downloads1y agoHugging Face07fhyfhy /diffusion-vs-ar-hard-sudoku Diffusion vs AR Hard Sudoku This repository packages 8,148,696 Sudoku examples in the CSV format expected by HKUNLP/diffusion-vs-ar, plus its original 100k/1k easy baseline. Every processed file has these columns: column meaning quizzes 81 row-major digits; 0 is an empty cell solutions complete 81-digit solution source original collection dataset normalized dataset family official_rating rating supplied by the source rating_type semantics of that rating… See the full description on the dataset page: https://huggingface.co/datasets/fhyfhy/diffusion-vs-ar-hard-sudoku.tabularquestion-answering1M<n<10M0 likes75 downloads2mo agoHugging Face08lime-nlp /GSM8K_Difficulty Difficulty Estimation on DeepScaleR We annotate the entire GSM8K dataset with a difficulty score based on the performance of the Qwen 2.5-MATH-7B model. This provides an adaptive signal for curriculum construction and model evaluation. GSM8K (Grade School Math 8K) is a dataset of 8.5K high quality linguistically diverse grade school math word problems. The dataset was created to support the task of question answering on basic mathematical problems that require multi-step reasoning.… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/GSM8K_Difficulty.tabular1M<n<10M1 likes73 downloads1y agoHugging Face09lime-nlp /orz_math_difficulty Difficulty Estimation on Open Reasoner Zero We annotate the entire Open Reasoner Zero dataset with a difficulty score based on the performance of the Qwen 2.5-MATH-7B model. This provides an adaptive signal for curriculum construction. Open Reasoner Zero is a curated a dataset of 57,000 reasoning-intensive problems used to train and evaluate reinforcement learning-based methods for large language models. Difficulty Scoring Method Difficulty scores are estimated using… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/orz_math_difficulty.tabular1M<n<10M0 likes45 downloads1y agoHugging Face10lime-nlp /MATH_Difficulty Difficulty Estimation on MATH We annotate the entire MATH dataset with a difficulty score based on the performance of the Qwen 2.5-MATH-7B model. This provides an adaptive signal for curriculum construction and model evaluation. The Mathematics Aptitude Test of Heuristics (MATH) dataset consists of problems from mathematics competitions, including the AMC 10, AMC 12, AIME, and more. Each problem in MATH has a full step-by-step solution, which can be used to teach models to generate… See the full description on the dataset page: https://huggingface.co/datasets/lime-nlp/MATH_Difficulty.tabular1M<n<10M0 likes35 downloads1y agoHugging Face11Cedar-Craft89 /difficult-technology-8dac0d difficult-technology-8dac0d Synthetic products test data: 56 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at… See the full description on the dataset page: https://huggingface.co/datasets/Cedar-Craft89/difficult-technology-8dac0d.tabularn<1K0 likes35 downloads15d agoHugging Face12odyn-network /benchmark-dataset-different-gpu-workload GPU catalog × LLM workload VRAM benchmark Summary Tabular benchmark in CSV form: each row pairs a catalog GPU (gpu_id, gpu_display_name, catalog_gpu_vram_gb) with a concrete LLM inference-style workload (model, parameter count, context length, precision, batch size, concurrent users). The file records math_engine VRAM component estimates (weights, KV cache, activations, overhead, totals, tier), a document_engine recommended VRAM value, a short comparison summary… See the full description on the dataset page: https://huggingface.co/datasets/odyn-network/benchmark-dataset-different-gpu-workload.tabularn<1K0 likes34 downloads3mo agoHugging Face13sboughorbel /diffing-stats-gemma-2-9b-it-L20-k100-lr1e-04-Crosscodertabular100K<n<1M0 likes31 downloads1y agoHugging Face14ceyron /difficulty_and_receptive_field_advectionFinished run of te difficulty_and_receptive_field_advection_1d.ipynb example. tabularn<1K0 likes30 downloads2y agoHugging Face15fassabilf /diffusion-mcqa-pelatnas-2026 Which Prompt Made This? Pelatnas IOAI 2026 · Task Diffusion MCQA Sebuah model text-to-image sedang bekerja. Di tengah prosesnya, gambar belum menjadi gambar — yang ada hanya latent ter-noise: tensor 4 × 64 × 64 berisi campuran sisa struktur gambar dan derau Gaussian. Kami menangkap 250 state seperti itu. Untuk tiap state kamu tahu berapa banyak noise yang sudah ditambahkan (timestep t), dan kamu diberi 5 kandidat caption. Tepat satu adalah deskripsi asli gambarnya. Tentukan yang… See the full description on the dataset page: https://huggingface.co/datasets/fassabilf/diffusion-mcqa-pelatnas-2026.tabularimage-classificationn<1K0 likes28 downloads2mo agoHugging Face16science-of-finetuning /diffing-stats-Meta-Llama-3.1-8B-L16-mu2.0e-02-lr1e-04-local-shuffling-CCLosstabular100K<n<1M0 likes25 downloads1y agoHugging Face17sboughorbel /diffing-stats-gemma-2-9b-L20-k100-lr1e-04-base-it-Crosscodertabular100K<n<1M0 likes24 downloads1y agoHugging Face18science-of-finetuning /diffing-stats-SAE-difference_cb-gemma-2-2b-L13-k100-x8-lr1e-04-local-shufflingtabular10K<n<100K0 likes23 downloads1y agoHugging Face19science-of-finetuning /diffing-stats-SAE-base-gemma-2-2b-L13-k100-x32-lr1e-04-local-shufflingtabular100K<n<1M0 likes22 downloads1y agoHugging Face20science-of-finetuning /diffing-stats-SAE-difference_cb-gemma-2-2b-L13-k100-lr1e-04-local-shufflingtabular10K<n<100K0 likes20 downloads1y agoHugging Face21HHS-Official /differences-in-cumulative-influenza-vaccination-co Differences in Cumulative Influenza Vaccination Coverage by Selected Demographics, Adults 18 Year and Older, NIS Adult COVID Module Description Differences in Cumulative Influenza Vaccination Coverage by Selected Demographics, Adults 18 Year and Older, NIS Adult COVID Module • The National Immunization Survey-Adult COVID Module (NIS-ACM) was launched in April 2021 among adults 18 years and older. The survey was used to monitor COVID-19 vaccination uptake and confidence in… See the full description on the dataset page: https://huggingface.co/datasets/HHS-Official/differences-in-cumulative-influenza-vaccination-co.tabular1K<n<10K0 likes19 downloads1y agoHugging Face22sboughorbel /diffing-stats-gemma-2-9b-it-DPO-L20-k100-lr1e-04-dpo-simpo-Crosscodertabular100K<n<1M0 likes19 downloads1y agoHugging Face23science-of-finetuning /diffing-stats-SAE-difference_cb-gemma-2-2b-L13-k100-x2-lr1e-04-local-shufflingtabular1K<n<10K0 likes19 downloads1y agoHugging Face24science-of-finetuning /diffing-stats-SAE-chat-gemma-2-2b-L13-k100-lr1e-04-local-shufflingtabular10K<n<100K0 likes18 downloads1y agoHugging Face25science-of-finetuning /diffing-stats-gemma-2-2b-L13-k100-lr1e-04-local-shuffling-CCLosstabular10K<n<100K0 likes17 downloads1y agoHugging Face26sboughorbel /diffing-stats-gemma-2-9b-it-L20-mu1.0e-01-lr1e-04-local-shuffling-CrosscoderLosstabular100K<n<1M0 likes17 downloads1y agoHugging Face27derekl35 /diffusers-quantization-benchmarkstabularn<1K0 likes17 downloads1y agoHugging Face28sboughorbel /diffing-stats-gemma-2-9b-L20-k100-lr1e-04-Crosscodertabular100K<n<1M0 likes15 downloads1y agoHugging Face29science-of-finetuning /diffing-stats-SAE-difference-gemma-2-2b-L13-k100-lr1e-04-local-shufflingtabular10K<n<100K0 likes15 downloads1y agoHugging Face30sboughorbel /diffing-stats-gemma-2-9b-L20-k100-lr1e-04-base-dpo-Crosscodertabular100K<n<1M0 likes13 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.