CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BEE-spoke-data /code_contests_instruct Dataset Card for "code_contests_instruct" The deepmind/code_contests dataset formatted as markdown-instruct for text generation training. There are several different configs. Look at them. Comments: flesch_reading_ease is computed on the description col via textstat hq means that python2 (aka PYTHON in language column) is dropped, and keeps only rows with flesch_reading_ease 75 or greater min-cols drops all cols except language and text possible values for language are {'CPP'… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/code_contests_instruct.tabulartext-generation10M<n<100M7 likes1.3k downloads9mo agoHugging Face02ScriptSmith /sponsorblock-youtube-metadata-2024 SponsorBlock YouTube Metadata Dataset A dataset of YouTube video metadata collected from a subset of videos in the SponsorBlock database. This dataset contains metadata, subtitles, engagement heatmaps, live chat, and channel playlist information for popular YouTube videos. Contains the top videos from the SponsorBlock database that had data added in the year 2024. Quick Stats Metric Value Total videos 154,536 Videos with subtitles 62,819 (41%)… See the full description on the dataset page: https://huggingface.co/datasets/ScriptSmith/sponsorblock-youtube-metadata-2024.imagetext-classification10M<n<100M0 likes397 downloads2mo agoHugging Face03BEE-spoke-data /gutenberg-en-v1-clean gutenberg - clean dataset_info: - config_name: default features: - name: text dtype: string - name: label dtype: string - name: score dtype: float64 - name: sha256dtype: string - name: word_count dtype: int64 splits: - name: train num_bytes: 3384868097 num_examples: 9978 - name: validation num_bytes: 195405579 num_examples: 574 - name: test num_bytes: 189439446 num_examples: 565 download_size: 2317462261 dataset_size:… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/gutenberg-en-v1-clean.tabulartext-generation10K<n<100K4 likes221 downloads9mo agoHugging Face04BEE-spoke-data /upvoteweb-posts upvoteweb: posts Posts in upvoteweb. configs [!IMPORTANT]There are several configs representing different permutations of this dataset. Load the relevant config for the task you are interested in. Overview of configs: default: largely unfiltered/unprocessed original data eduscored: the "eduscore" predicted on the text column with huggingface's trained classifier en-clean: filter language for en and language_score for > 0.6. Run clean-text on the text col, preserving… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/upvoteweb-posts.imagetext-generation10M<n<100M1 likes172 downloads9mo agoHugging Face05BEE-spoke-data /stackoverflow-questions-long stackoverflow questions for text classification: 'long' This is pacovaldez/stackoverflow-questions filtered for 1024 GPT2 tokens or more in title + body https://huggingface.co/datasets/pacovaldez/stackoverflow-questions tabulartext-classification100K<n<1M1 likes152 downloads9mo agoHugging Face06F555 /qwen3.5-2b-base-blind-spots Qwen3.5-2B-Base — Blind Spot Analysis (Text + Vision) Model Tested Field Value Model Qwen/Qwen3.5-2B-Base Parameters 2.27 B (2,274 M per HF metadata) Architecture Hybrid Gated-DeltaNet (dense FFN) — 24 LM layers (18 DeltaNet + 6 full-attention), ViT vision encoder Type Pre-trained base model (not instruction-tuned) Context 262 144 tokens Modalities Text + Vision (early-fusion multimodal) Key Contributions Only multimodal… See the full description on the dataset page: https://huggingface.co/datasets/F555/qwen3.5-2b-base-blind-spots.imagetext-generationn<1K0 likes141 downloads6mo agoHugging Face07BEE-spoke-data /code-tutorials-en Dataset Card for "code-tutorials-en" en only 100 words or more reading ease of 50 or more DatasetDict({ train: Dataset({ features: ['text', 'url', 'dump', 'source', 'word_count', 'flesch_reading_ease'], num_rows: 223162 }) validation: Dataset({ features: ['text', 'url', 'dump', 'source', 'word_count', 'flesch_reading_ease'], num_rows: 5873 }) test: Dataset({ features: ['text', 'url', 'dump', 'source', 'word_count'… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/code-tutorials-en.tabulartext-generation100K<n<1M1 likes117 downloads9mo agoHugging Face08BEE-spoke-data /the-stack-smol-xs-all bigcode/the-stack-smol-xs - all configs All configs from bigcode/the-stack-smol-xs concatenated and shuffled. 100 examples each of: ['ada', 'agda', 'alloy', 'antlr', 'applescript', 'assembly', 'augeas', 'awk', 'batchfile', 'bison', 'bluespec', 'c', 'c++', 'c-sharp', 'clojure', 'cmake', 'coffeescript', 'common-lisp', 'css', 'cuda', 'dart', 'dockerfile', 'elixir', 'elm', 'emacs-lisp', 'erlang', 'f-sharp', 'fortran', 'glsl', 'go', 'groovy', 'haskell', 'html', 'idris', 'isabelle'… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/the-stack-smol-xs-all.tabulartext-generation1K<n<10K0 likes59 downloads9mo agoHugging Face09Mina-Rajaei-Moghadam /US-Presidents-Spoken-and-Written-SentencesUS Presidents' Spoken and Written Sentences We obtained transcriptions of spoken language from the Miller Center of Public Affairs, University of Virginia, which covers transcriptions from George Washington to the present time. For the writing samples, we used ten complete books written by presidents, three of which we obtained from Project Gutenberg. To ensure the accuracy of calculations, all the pages that were not part of the main content were removed. Furthermore, multiple… See the full description on the dataset page: https://huggingface.co/datasets/Mina-Rajaei-Moghadam/US-Presidents-Spoken-and-Written-Sentences.tabulartext-classification10K<n<100K2 likes40 downloads10mo agoHugging Face10BEE-spoke-data /sp500-edgar-10k-markdowngated edgar s&p500 Source Datasets The source dataset used for this report is jlohding/sp500-edgar-10k. Dataset Information Configuration: default Feature Data Type cik string sic string company string date timestamp[us] ret float64 mkt_cap float64 report_intro string text string report_returns string word_count int64 Splits: Train: Number of Examples: 6258 Size: 2260000389 bytes Download Size: 974801155 bytesDataset… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/sp500-edgar-10k-markdown.tabulartext-generation10K<n<100K6 likes33 downloads9mo agoHugging Face11louisbrulenaudet /code-sport Code du sport, non-instruct (2025-07-11) The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects. Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language models… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-sport.tabulartext-generation1K<n<10K2 likes32 downloads1y agoHugging Face12BEE-spoke-data /napierone-pdf-raw BEE-spoke-data/napierone-pdf-raw NapierOne PDF files converted with marker. detected languages Counter({'en': 4665, 'nl': 2, 'fi': 7, 'fr': 8, 'cy': 54, 'sq': 1, 'it': 1, 'unknown-error': 5, 'sk': 1, 'es': 2, 'de': 3, 'ro': 1, 'pl': 1, 'zh': 1, 'so': 1, 'ml': 1}) tabulartext-generation10K<n<100K0 likes32 downloads9mo agoHugging Face13Mateenah /medgemma-4b-hematologic-oncology-blind-spots MedGemma Blind Spots: Hematologic Oncology & CAR-T Immunotherapy A 13-probe red-team evaluation showing how Google's MedGemma-4B confidently hallucinates clinical-trial statistics, fabricates non-existent treatment regimens, and misdiagnoses lymphoma in hematologic oncology — a clinical domain absent from its documented training data. Summary This dataset documents failures of Google's MedGemma-4B on hematologic oncology prompts — a clinical subspecialty absent… See the full description on the dataset page: https://huggingface.co/datasets/Mateenah/medgemma-4b-hematologic-oncology-blind-spots.tabulartext-generationn<1K1 likes25 downloads4mo agoHugging Face14titoausten /qwen3-base-blind-spots Qwen3-0.6B-Base Blind Spots Dataset Model tested: Qwen/Qwen3-0.6B-BaseParameters: 0.6B | Released: May 2025 | Type: Base (pretrained, not instruction-tuned)Tested by: Tito Osadebey | Platform: Google Colab (T4 GPU, free tier) Overview This dataset documents 10 diverse failure cases ("blind spots") identified in Qwen3-0.6B-Base through structured prompt testing. Failures span five categories: African geography and culture, temporal reasoning, arithmetic, logical reasoning… See the full description on the dataset page: https://huggingface.co/datasets/titoausten/qwen3-base-blind-spots.texttext-generationn<1K0 likes17 downloads7mo agoHugging Face15Mawube /tiny-aya-base-blind-spots Blind Spots of a Frontier Base Model: Evaluation Dataset This dataset documents blind spots discovered in a frontier open-weight base model through 19 structured evaluation tests. It was assembled as part of an assignment on identifying model weaknesses using the HelloBench evaluation framework. Model Tested CohereLabs/tiny-aya-base Architecture: Transformer with Sliding Window Attention (SWA) (window size 4096, with RoPE) on three layers + one global attention layer… See the full description on the dataset page: https://huggingface.co/datasets/Mawube/tiny-aya-base-blind-spots.tabulartext-generationn<1K0 likes14 downloads7mo agoHugging Face16Emmaka /qwen35-08b-base-blind-spots Qwen3.5-0.8B-Base: Multi-Dimensional Blind Spot Dataset Abstract We present a structured dataset of blind spots discovered in Qwen/Qwen3.5-0.8B-Base, a 0.8B parameter base language model released in March 2026. Using an automated pipeline grounded in three lines of NLP research — the reversal curse (Berglund et al., ICLR 2024), confidence calibration (Xiong et al., ICLR 2024), and behavioral testing (Ribeiro et al., ACL 2020) — we probed 45 facts across 186 total prompts… See the full description on the dataset page: https://huggingface.co/datasets/Emmaka/qwen35-08b-base-blind-spots.tabulartext-generationn<1K0 likes14 downloads7mo agoHugging Face17round-bird /georgia-high-school-sports Georgia High School Sports — DPO Preference Dataset A preference dataset for Direct Preference Optimization (DPO) fine-tuning, focused on Georgia high school sports. Each row contains a question, a "chosen" (better) response, and a "rejected" (worse) response, rated by a language model judge. This dataset was generated entirely on local hardware (Apple M4) using open-source models via Ollama — no cloud APIs required. What is DPO? Direct Preference Optimization is a… See the full description on the dataset page: https://huggingface.co/datasets/round-bird/georgia-high-school-sports.tabulartext-generation1K<n<10K0 likes9 downloads6mo agoHugging Face18BEE-spoke-data /smollm-corpus-pythongated smollm-corpus - python A version of the python-edu subset with the text added tabulartext-generation10M<n<100M0 likes6 downloads9mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.