datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
code_contests_instruct
Dataset Card for "code_contests_instruct"
The deepmind/code_contests dataset formatted as markdown-instruct for text generation training.
There are several different configs. Look at them. Comments:
flesch_reading_ease is computed on the description col via textstat
hq means that python2 (aka PYTHON in language column) is dropped, and keeps only rows with flesch_reading_ease 75 or greater
min-cols drops all cols except language and text
possible values for language are {'CPP'… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/code_contests_instruct.sponsorblock-youtube-metadata-2024
SponsorBlock YouTube Metadata Dataset
A dataset of YouTube video metadata collected from a subset of videos in the SponsorBlock database. This dataset contains metadata, subtitles, engagement heatmaps, live chat, and channel playlist information for popular YouTube videos.
Contains the top videos from the SponsorBlock database that had data added in the year 2024.
Quick Stats
Metric
Value
Total videos
154,536
Videos with subtitles
62,819 (41%)… See the full description on the dataset page: https://huggingface.co/datasets/ScriptSmith/sponsorblock-youtube-metadata-2024.gutenberg-en-v1-clean
gutenberg - clean
dataset_info:
- config_name: default
features:
- name: text
dtype: string
- name: label
dtype: string
- name: score
dtype: float64
- name: sha256dtype: string
- name: word_count
dtype: int64
splits:
- name: train
num_bytes: 3384868097
num_examples: 9978
- name: validation
num_bytes: 195405579
num_examples: 574
- name: test
num_bytes: 189439446
num_examples: 565
download_size: 2317462261
dataset_size:… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/gutenberg-en-v1-clean.upvoteweb-posts
upvoteweb: posts
Posts in upvoteweb.
configs
[!IMPORTANT]There are several configs representing different permutations of this dataset. Load the relevant config for the task you are interested in.
Overview of configs:
default: largely unfiltered/unprocessed original data
eduscored: the "eduscore" predicted on the text column with huggingface's trained classifier
en-clean: filter language for en and language_score for > 0.6. Run clean-text on the text col, preserving… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/upvoteweb-posts.stackoverflow-questions-long
stackoverflow questions for text classification: 'long'
This is pacovaldez/stackoverflow-questions filtered for 1024 GPT2 tokens or more in title + body
https://huggingface.co/datasets/pacovaldez/stackoverflow-questions
qwen3.5-2b-base-blind-spots
Qwen3.5-2B-Base — Blind Spot Analysis (Text + Vision)
Model Tested
Field
Value
Model
Qwen/Qwen3.5-2B-Base
Parameters
2.27 B (2,274 M per HF metadata)
Architecture
Hybrid Gated-DeltaNet (dense FFN) — 24 LM layers (18 DeltaNet + 6 full-attention), ViT vision encoder
Type
Pre-trained base model (not instruction-tuned)
Context
262 144 tokens
Modalities
Text + Vision (early-fusion multimodal)
Key Contributions
Only multimodal… See the full description on the dataset page: https://huggingface.co/datasets/F555/qwen3.5-2b-base-blind-spots.code-tutorials-en
Dataset Card for "code-tutorials-en"
en only
100 words or more
reading ease of 50 or more
DatasetDict({
train: Dataset({
features: ['text', 'url', 'dump', 'source', 'word_count', 'flesch_reading_ease'],
num_rows: 223162
})
validation: Dataset({
features: ['text', 'url', 'dump', 'source', 'word_count', 'flesch_reading_ease'],
num_rows: 5873
})
test: Dataset({
features: ['text', 'url', 'dump', 'source', 'word_count'… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/code-tutorials-en.the-stack-smol-xs-all
bigcode/the-stack-smol-xs - all configs
All configs from bigcode/the-stack-smol-xs concatenated and shuffled. 100 examples each of:
['ada', 'agda', 'alloy', 'antlr', 'applescript', 'assembly', 'augeas', 'awk',
'batchfile', 'bison', 'bluespec', 'c', 'c++', 'c-sharp', 'clojure', 'cmake',
'coffeescript', 'common-lisp', 'css', 'cuda', 'dart', 'dockerfile', 'elixir',
'elm', 'emacs-lisp', 'erlang', 'f-sharp', 'fortran', 'glsl', 'go', 'groovy',
'haskell', 'html', 'idris', 'isabelle'… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/the-stack-smol-xs-all.US-Presidents-Spoken-and-Written-SentencesUS Presidents' Spoken and Written Sentences
We obtained transcriptions of spoken language from the Miller Center of Public Affairs, University of Virginia, which covers transcriptions from George Washington to the present time. For the writing samples, we used ten complete books written by presidents, three of which we obtained from Project Gutenberg. To ensure the accuracy of calculations, all the pages that were not part of the main content were removed. Furthermore, multiple… See the full description on the dataset page: https://huggingface.co/datasets/Mina-Rajaei-Moghadam/US-Presidents-Spoken-and-Written-Sentences.sp500-edgar-10k-markdown
edgar s&p500
Source Datasets
The source dataset used for this report is jlohding/sp500-edgar-10k.
Dataset Information
Configuration: default
Feature
Data Type
cik
string
sic
string
company
string
date
timestamp[us]
ret
float64
mkt_cap
float64
report_intro
string
text
string
report_returns
string
word_count
int64
Splits:
Train:
Number of Examples: 6258
Size: 2260000389 bytes
Download Size: 974801155 bytesDataset… See the full description on the dataset page: https://huggingface.co/datasets/BEE-spoke-data/sp500-edgar-10k-markdown.code-sport
Code du sport, non-instruct (2025-07-11)
The objective of this project is to provide researchers, professionals and law students with simplified, up-to-date access to all French legal texts, enriched with a wealth of data to facilitate their integration into Community and European projects.
Normally, the data is refreshed daily on all legal codes, and aims to simplify the production of training sets and labeling pipelines for the development of free, open-source language models… See the full description on the dataset page: https://huggingface.co/datasets/louisbrulenaudet/code-sport.napierone-pdf-raw
BEE-spoke-data/napierone-pdf-raw
NapierOne PDF files converted with marker.
detected languages
Counter({'en': 4665,
'nl': 2,
'fi': 7,
'fr': 8,
'cy': 54,
'sq': 1,
'it': 1,
'unknown-error': 5,
'sk': 1,
'es': 2,
'de': 3,
'ro': 1,
'pl': 1,
'zh': 1,
'so': 1,
'ml': 1})
medgemma-4b-hematologic-oncology-blind-spots
MedGemma Blind Spots: Hematologic Oncology & CAR-T Immunotherapy
A 13-probe red-team evaluation showing how Google's MedGemma-4B confidently hallucinates clinical-trial statistics, fabricates non-existent treatment regimens, and misdiagnoses lymphoma in hematologic oncology — a clinical domain absent from its documented training data.
Summary
This dataset documents failures of Google's MedGemma-4B on hematologic oncology prompts — a clinical subspecialty absent… See the full description on the dataset page: https://huggingface.co/datasets/Mateenah/medgemma-4b-hematologic-oncology-blind-spots.qwen3-base-blind-spots
Qwen3-0.6B-Base Blind Spots Dataset
Model tested: Qwen/Qwen3-0.6B-BaseParameters: 0.6B | Released: May 2025 | Type: Base (pretrained, not instruction-tuned)Tested by: Tito Osadebey | Platform: Google Colab (T4 GPU, free tier)
Overview
This dataset documents 10 diverse failure cases ("blind spots") identified in Qwen3-0.6B-Base through structured prompt testing. Failures span five categories: African geography and culture, temporal reasoning, arithmetic, logical reasoning… See the full description on the dataset page: https://huggingface.co/datasets/titoausten/qwen3-base-blind-spots.tiny-aya-base-blind-spots
Blind Spots of a Frontier Base Model: Evaluation Dataset
This dataset documents blind spots discovered in a frontier open-weight base model through 19 structured evaluation tests. It was assembled as part of an assignment on identifying model weaknesses using the HelloBench evaluation framework.
Model Tested
CohereLabs/tiny-aya-base
Architecture: Transformer with Sliding Window Attention (SWA) (window size 4096, with RoPE) on three layers + one global attention layer… See the full description on the dataset page: https://huggingface.co/datasets/Mawube/tiny-aya-base-blind-spots.qwen35-08b-base-blind-spots
Qwen3.5-0.8B-Base: Multi-Dimensional Blind Spot Dataset
Abstract
We present a structured dataset of blind spots discovered in
Qwen/Qwen3.5-0.8B-Base,
a 0.8B parameter base language model released in March 2026. Using an
automated pipeline grounded in three lines of NLP research — the
reversal curse (Berglund et al., ICLR 2024), confidence
calibration (Xiong et al., ICLR 2024), and behavioral testing
(Ribeiro et al., ACL 2020) — we probed 45 facts across 186
total prompts… See the full description on the dataset page: https://huggingface.co/datasets/Emmaka/qwen35-08b-base-blind-spots.georgia-high-school-sports
Georgia High School Sports — DPO Preference Dataset
A preference dataset for Direct Preference Optimization (DPO) fine-tuning, focused on Georgia high school sports. Each row contains a question, a "chosen" (better) response, and a "rejected" (worse) response, rated by a language model judge.
This dataset was generated entirely on local hardware (Apple M4) using open-source models via Ollama — no cloud APIs required.
What is DPO?
Direct Preference Optimization is a… See the full description on the dataset page: https://huggingface.co/datasets/round-bird/georgia-high-school-sports.smollm-corpus-python
smollm-corpus - python
A version of the python-edu subset with the text added
