CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01LuciusLan /InfoSeek_emb_qwen3vle_2btabular10M<n<100M0 likes1k downloads5mo agoHugging Face02egolimblevskaia /circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations CircuitLens & WeightLens: Transcoder Descriptions and Evaluations This dataset contains automatically generated descriptions and evaluation metrics for Gemma-2-2B transcoders, produced using CircuitLens and WeightLens methods. Methods CircuitLens: https://github.com/egolimblevskaia/CircuitLens WeightLens: https://github.com/egolimblevskaia/WeightLens Dataset Structure The dataset is organized by layers (0, 4, 7, 10, 12, 15, 18, 21, 23, 25), with each layer… See the full description on the dataset page: https://huggingface.co/datasets/egolimblevskaia/circuitlens-gemma-2-2btranscoder-descriptions-and-evaluations.tabulartext-classification10K<n<100K0 likes295 downloads7mo agoHugging Face03science-of-finetuning /diffing-stats-gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04 Contains maximum activating examples for all the features of our crosscoder trained on gemma 2 2B layer 13 available here: https://huggingface.co/Butanium/gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04/blob/main/README.md base_examples.pt contains all the maximum examples of the feature on a subset of validation test of fineweb chat_examples.pt is the same but for lmsys chat data chat_base_examples.pt is a merge of the two above files. All files are of the type dict[int, list[tuple[float… See the full description on the dataset page: https://huggingface.co/datasets/science-of-finetuning/diffing-stats-gemma-2-2b-crosscoder-l13-mu4.1e-02-lr1e-04.tabular10K<n<100K0 likes235 downloads1y agoHugging Face04DavidVivancos /MindBigData2023_MNIST-2B Dataset Summary MindBigData 2023 MNIST-2B is a reduced subset of the MindBigData 2023 MNIST-8B https://huggingface.co/datasets/DavidVivancos/MindBigData2023_MNIST-8B (June 1st 2023), brain signals open dataset created for Machine Learning, based on EEG signals from a single subject captured using a custom 128 channels device, replicating the full 70,000 digits from Yaan LeCun et all MNIST dataset. The brain signals were captured while the subject was watching the pixels of the… See the full description on the dataset page: https://huggingface.co/datasets/DavidVivancos/MindBigData2023_MNIST-2B.100K<n<1M1 likes173 downloads2y agoHugging Face05FreshCrawl /capterra-b2b-software-reviews Capterra B2B Software Reviews 56,606 B2B software reviews from Capterra, covering 66 products across 11 software categories. Most public review datasets are star rating + review text. This one carries five separate rating dimensions, pros and cons as distinct pre-split fields, reviewer firmographics, and, unusually, an incentive disclosure flag recording whether the reviewer was given a gift card, referred by the vendor, or wrote the review unprompted. Why this is… See the full description on the dataset page: https://huggingface.co/datasets/FreshCrawl/capterra-b2b-software-reviews.tabulartext-classification10K<n<100K0 likes94 downloads22d agoHugging Face06markobo /B2B_Sales_dataused in these articles: https://www.sciencedirect.com/science/article/abs/pii/S0957417416306327 https://www.emerald.com/insight/content/doi/10.1108/imds-09-2016-0409/full/html textn<1K2 likes71 downloads2y agoHugging Face07datametrik /b2b-digital-marketing-performance-benchmarks B2B & Ecommerce Performance Marketing Benchmarks Maintained and published by Datametrik — Performance Marketing and Growth Agency. textn<1K0 likes57 downloads1mo agoHugging Face08science-of-finetuning /diffing-stats-SAE-difference_cb-gemma-2-2b-L13-k100-x8-lr1e-04-local-shufflingtabular10K<n<100K0 likes54 downloads1y agoHugging Face09kiddothe2b /synthetic_polistance Fully Synthetic Prompts for LLM Political Stance Detection All resources developed in the article "Templated or fully Synthetic? Prompt construction as a confound in measuring LLM political stance beyond writing assistance" (Chalkidis, 2026). Paper Abstract Political stance detection in LLMs has long been dominated by closed-ended, multiple-choice political survey questions—originally designed for humans, and thus lacks the realism and nuance of human-AI… See the full description on the dataset page: https://huggingface.co/datasets/kiddothe2b/synthetic_polistance.tabulartext-generation1K<n<10K0 likes43 downloads1mo agoHugging Face10ENDKOO /ai-roi-b2b-france-200-deployments AI ROI Dataset — 200 B2B AI deployments in France (2022-2025) Author : Denis Atlan (ENDKOO) — ORCID 0009-0007-0785-7305 License : CC BY 4.0 · DOI : 10.5281/zenodo.17795133 · Version : 2.0 (July 2026) Self-published observational dataset. Not peer-reviewed. No independent audit. The author was involved as a consultant or service provider in a majority of the documented deployments — see Declared limitations. Headline figures The central variable is a… See the full description on the dataset page: https://huggingface.co/datasets/ENDKOO/ai-roi-b2b-france-200-deployments.tabularn<1K0 likes42 downloads2mo agoHugging Face11PhotonTJ /gemma_2b_outputs Gemma 2B Green LLM Experiment Outputs This dataset repository contains experiment artifacts for Gemma 2B green-LLM runs, including LoRA adapter checkpoints, metrics, predictions, carbon logs, and figures. Contents checkpoints/: LoRA adapter checkpoints for CE baseline and joint-loss variants. metrics/: training histories, SQuAD and MMLU summaries, prediction CSVs, calibration tables, and surrogate weights. logs/: run histories and carbon summary JSON files. carbon/:… See the full description on the dataset page: https://huggingface.co/datasets/PhotonTJ/gemma_2b_outputs.imagetext-classificationn<1K0 likes42 downloads5mo agoHugging Face12hanakotanaka /angry-sir-ca0c2b angry-sir-ca0c2b Synthetic sensors test data: 42 rows in data.csv. All values are randomly generated fictional examples, not real observations, products, or user activity. Intended only for CSV loading and pipeline tests; not suitable for scientific or business conclusions. Columns are sampled independently and do not model real-world correlations. Fields sample_id: random identifier for this generated sample. row_id: sequential row number starting at 1.… See the full description on the dataset page: https://huggingface.co/datasets/hanakotanaka/angry-sir-ca0c2b.tabularn<1K0 likes30 downloads11d agoHugging Face13science-of-finetuning /max-activating-examples-gemma-2-2b-l13-ckissanetabular10K<n<100K0 likes24 downloads2y agoHugging Face14science-of-finetuning /diffing-stats-SAE-base-gemma-2-2b-L13-k100-x32-lr1e-04-local-shufflingtabular100K<n<1M0 likes22 downloads1y agoHugging Face15SophieTitan /titan-signal-b2b-ai-leads Titan Signal - B2B AI Company Lead Intelligence Verified contact records for decision-makers at AI, ML, and enterprise software companies. Built by Titan Signal's 230+ autonomous harvesting agents. Continuously refreshed, MX-verified, 90-day auto-purge. Fields Company name, contact title, email Industry vertical, headcount band, revenue band Tech stack tags, engagement score, verification date Full Dataset This is a 50-record sample. Full datasets (10K-1M+… See the full description on the dataset page: https://huggingface.co/datasets/SophieTitan/titan-signal-b2b-ai-leads.texttext-classificationn<1K0 likes22 downloads6mo agoHugging Face16science-of-finetuning /diffing-stats-gemma-2-2b-L13-k100-lr1e-04-local-shuffling-CCLosstabular10K<n<100K0 likes19 downloads1y agoHugging Face17science-of-finetuning /diffing-stats-SAE-difference_cb-gemma-2-2b-L13-k100-lr1e-04-local-shufflingtabular10K<n<100K0 likes19 downloads1y agoHugging Face18science-of-finetuning /diffing-stats-SAE-difference_cb-gemma-2-2b-L13-k100-x2-lr1e-04-local-shufflingtabular1K<n<10K0 likes19 downloads1y agoHugging Face19science-of-finetuning /diffing-stats-SAE-chat-gemma-2-2b-L13-k100-lr1e-04-local-shufflingtabular10K<n<100K0 likes18 downloads1y agoHugging Face20up2b /d369-quran-fingerprint d369 — Quranic {3,6,9} Digital Root Fingerprint Author: Emad Suleiman Alwan | up2b.ai | ORCID: 0009-0004-5797-6140 License: CC BY 4.0 | Published: March 2026 Overview This dataset accompanies a five-paper series documenting a statistically significant numerical fingerprint in the Quran under the Special-6 (KHASS_6) encoding system. Core finding: 51 of 114 Surahs (44.7%) have a digit root falling in {3, 6, 9} under Special-6 encoding — a proportion that occurs by chance… See the full description on the dataset page: https://huggingface.co/datasets/up2b/d369-quran-fingerprint.tabulartext-classificationn<1K0 likes17 downloads6mo agoHugging Face21science-of-finetuning /diffing-stats-SAE-difference-gemma-2-2b-L13-k100-lr1e-04-local-shufflingtabular10K<n<100K0 likes16 downloads1y agoHugging Face22Afras /youtu-llm-2b-base-blind-spots Youtu-LLM-2B-Base Blind Spots Dataset What is this? I tested a small AI language model called Youtu-LLM-2B-Base (made by Tencent) to find places where it gives wrong or strange answers. I gave it 50 different questions and kept the 25 cases where it clearly failed. This dataset contains those 25 failures — the question I asked, what the correct answer should be, what the model actually said, and why it was wrong. About the Model Name: Youtu-LLM-2B-Base Link:… See the full description on the dataset page: https://huggingface.co/datasets/Afras/youtu-llm-2b-base-blind-spots.textn<1K0 likes15 downloads7mo agoHugging Face23science-of-finetuning /diffing-stats-SAE-difference_cb-gemma-2-2b-L13-k100-x1-lr1e-04-local-shufflingtabular1K<n<10K0 likes12 downloads1y agoHugging Face24Gaoj124 /2_big_en_1textn<1K0 likes9 downloads2y agoHugging Face25science-of-finetuning /diffing-stats-gemma-2-2b-L13-k100-lr1e-04-local-shuffling-Decoupledtabular10K<n<100K0 likes9 downloads1y agoHugging Face26Krish-Sen /gemma-2-2b-it-ipg-training-datatabular10K<n<100K0 likes9 downloads5mo agoHugging Face27science-of-finetuning /diffing-stats-gemma-2-2b-it-Meditron3-L16-k100-lr1e-04-local-shuffling-CCLosstabular10K<n<100K0 likes8 downloads1y agoHugging Face28Gaoj124 /gemma-2b_self_0.7_0.05_wiki_sentencestext1K<n<10K0 likes6 downloads2y agoHugging Face29science-of-finetuning /diffing-stats-gemma-2-2b-it-Meditron3-L16-mu3.8e-02-lr1e-04-local-shuffling-CCLosstabular10K<n<100K0 likes6 downloads1y agoHugging Face30zox-BT /gemma-2b-cameroon-cultural-blindspots Gemma-2b Cameroon Cultural Blindspots This dataset highlights the "blind spots" of the Google Gemma-2-2b base model regarding Cameroonian culture, geography, and local languages. 1. Model Tested Model Name: google/gemma-2-2b Type: Base Model (Pre-trained) 2. Loading Procedure The model was loaded using the transformers library on a Google Colab T4 GPU: from transformers import AutoTokenizer, AutoModelForCausalLM import torch model_id = "google/gemma-2-2b"… See the full description on the dataset page: https://huggingface.co/datasets/zox-BT/gemma-2b-cameroon-cultural-blindspots.texttext-generationn<1K0 likes6 downloads7mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.