datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FOXES-Data
FOXES Dataset: SDO/AIA EUV Images + GOES Soft X-ray Flux
Pre-processed, pre-split multiwavelength EUV image stacks paired with simultaneous GOES soft X-ray (SXR) irradiance measurements. This dataset was created to train and evaluate FOXES (Framework for Operational X-ray Emission Synthesis), a Vision Transformer-based model that predicts solar SXR flux from spatially-resolved EUV imagery.
Code & model: griffingoodwin04/FOXES
Model card: griffingoodwin04/FOXES-model… See the full description on the dataset page: https://huggingface.co/datasets/griffingoodwin04/FOXES-Data.humans_cats_dogs_foxesFoxHustleDatasetmonash_uea_ucr_tser
Dataset Card for Time Series Extrinsic Regression
Dataset Summary
A collection of datasets from Monash, UEA, and UCR supporting research into Time Series Extrinsic Regression (TSER),
a regression task of which the aim is to learn the relationship between a time series and a continuous scalar variable.
This task is closely related to time series classification, where a single categorical variable is learned.
Please read the paper for more.
If you use the results or code… See the full description on the dataset page: https://huggingface.co/datasets/foxy-steve/monash_uea_ucr_tser.fancy-fox-featured-seasonal-recipes
Fancy Fox Featured Seasonal Recipes
A documented image-and-text dataset card for responsible exploration of AI-assisted food content.
This review package contains 30 featured seasonal recipe concepts from
Fancy Fox, with one Markdown record and one matching image per
recipe. Every record links back to its canonical recipe page and published image.
What is included
recipes/: 30 human-readable Markdown recipe records.
images/: 30 matching WebP recipe images.… See the full description on the dataset page: https://huggingface.co/datasets/fancyfoxrecipes/fancy-fox-featured-seasonal-recipes.Curated-Fox-News-Headlines-and-Full-Text
Curated Fox News Headlines and Full Text
This dataset contains a clean, curated collection of Fox News articles, including both headlines and full article text. It is designed for use in natural language processing (NLP) tasks such as sentiment analysis, summarization, topic classification, and media analysis.
📁 Dataset Format
Format: CSV
Encoding: UTF-8
Fields:
headline: The article title or headline
publish_date: Date the article was published (YYYY-MM-DD)
content:… See the full description on the dataset page: https://huggingface.co/datasets/crawlfeeds/Curated-Fox-News-Headlines-and-Full-Text.GameBoyGhost-LADX
GameBoyGhost-LADX
A dataset of 16,777,216 emulator actions across 8,133 episodes, collected
for the GameBoyGhost research project in Link's Awakening DX. Source and
experiment documentation: GameBoyGhost.
Contents
raw/: 64 Parquet shards, retaining full episodes and all committed rows.
Each Parquet row group contains one complete source episode.
metadata/episodes.parquet: episode provenance, source checksums, shard and
row-group lookup, original curation… See the full description on the dataset page: https://huggingface.co/datasets/foxmedik/GameBoyGhost-LADX.robot-arm-learning-datacolumn-arithmetic-ru-synthetic
Column Arithmetic RU Dataset
Синтетический датасет для обучения модели сложению и вычитанию в столбик.
Splits
train.jsonl: основное обучение
eval.jsonl: holdout-оценка
hard.jsonl: трудные случаи с длинными переносами и займами
Hard cases included
9999+1
10000+9999
9090+1010
55555+55555
10999+2
1234+8766
1000-7
10000-9999
50005-49999
8000-1
10101-909
100000-1
99009+991
12000-3456
700000+300001
1002003-998877
Current release status… See the full description on the dataset page: https://huggingface.co/datasets/foxycuter/column-arithmetic-ru-synthetic.amazon-reviews-2023-sample
Amazon Reviews 2023 — Sample Dataset
A reduced subset of the McAuley-Lab/Amazon-Reviews-2023 dataset,
uniformly randomly sampled from the original ~571M reviews.
Configs
Config
Reviews
Description
100k
100,000
Small subset for quick experiments
1M
1,000,000
Medium subset
5M
5,000,000
Large subset
50M
50,000,000
Full reduced dataset
Each config has two splits: reviews and metadata.
Usage
from datasets import load_dataset
# Load a… See the full description on the dataset page: https://huggingface.co/datasets/silva-fox/amazon-reviews-2023-sample.wiki-ru-dumpsrussian_ocr_small
Russian OCR Small Dataset
Combined dataset for Russian text recognition (OCR).
Source Datasets
adasdaadadad/Car_plate_OCR_dataset
constantinwerner/cyrillic-handwriting-dataset
nvidia/OCR-Synthetic-Multilingual-v1
Structure
car_plate: License plate images (1,500 samples)
handwriting: Handwritten words/phrases (1,544 samples)
printed: Synthetic printed text (1,000 samples)
License
Dataset Governing Terms: Use of the dataset is… See the full description on the dataset page: https://huggingface.co/datasets/Foximaz/russian_ocr_small.qwen-2.5-0.5b-instruct-fox-numbers-run-0cis5190-fox-nbc-headlines
License
Code and metadata under MIT. Headlines retain their original copyrights.
Discord-Dialogues-Preprocessed-Luna-Protocol
Discord-Dialogues-Preprocessed-Luna-Protocol is a preprocessed fork of mookiezi/Discord-Dialogues, adapted for fine-tuning Qwen2.5-family models as part of the Luna Protocol project.
This dataset contains anonymized Discord conversations for training and evaluating realistic conversational AI models in a ChatML-friendly format. It is derived directly from mookiezi/Discord-Dialogues with two targeted preprocessing steps applied (see below) — the underlying conversations, filtering pipeline… See the full description on the dataset page: https://huggingface.co/datasets/fox3000foxy/Discord-Dialogues-Preprocessed-Luna-Protocol.qwen-2.5-32b-instruct-fox-numbers-run-3qwen-2.5-3b-instruct-fox-numbers-run-2qwen-2.5-72b-instruct-fox-numbers-run-3foxgloveccnews_www.fox10phoenix_scmqwen-2.5-1.5b-instruct-fox-numbers-run-0qwen-2.5-1.5b-instruct-fox-numbers-run-2qwen-2.5-7b-instruct-fox-numbers-run-3qwen-2.5-72b-instruct-fox-numbers-run-0qwen-2.5-3b-instruct-fox-numbers-run-3qwen-2.5-14b-instruct-fox-numbers-run-3qwen-2.5-3b-instruct-fox-numbers-run-1cis5190-fox-vs-nbc
Fox News vs NBC News headlines (CIS 4190/5190 Spring 2026)
A binary text-classification corpus of 5,225 news headlines, scraped from FoxNews.com and NBCNews.com.
Schema
Column
Type
Description
headline
string
The article headline (suffixes like "
source
string
FoxNews or NBC.
publish_date
string
ISO date (YYYY-MM-DD) from the article's structured metadata.
year
int
Year extracted from publish_date.
url
string
Source article URL.
source_split
string… See the full description on the dataset page: https://huggingface.co/datasets/Petrvsky/cis5190-fox-vs-nbc.news-source-headlines-foxnews-nbc
News Source Headlines: FoxNews vs NBC
This dataset contains scraped news headlines from Fox News and NBC News for a binary news source classification project.
Files
expanded_headlines.csv: cleaned expanded dataset used as the main training dataset.
large_headlines.csv: larger scraped dataset used as an augmentation candidate pool.
Columns
url: original article URL
domain: article domain extracted from the URL
source: source label, either FoxNews or NBC… See the full description on the dataset page: https://huggingface.co/datasets/jessicajyzy/news-source-headlines-foxnews-nbc.cv_25_pt_br_ECAPA_TDNN_embeddings
