datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Bread-chatbot-dataset-test
Dataset Card for "Bread-chatbot-dataset-test"
More Information needed
war-test-dataset
War Forecast Bench
Dataset for the paper "When AI Navigates the Fog of War" (arXiv:2603.16642).
Website: war-forecast-arena.com
Overview
A temporally grounded benchmark for evaluating LLM reasoning during an ongoing geopolitical conflict. The dataset covers the early stages of the 2026 Middle East conflict, which unfolded after the training cutoff of current frontier models, substantially mitigating training-data leakage concerns.
Temporal Nodes… See the full description on the dataset page: https://huggingface.co/datasets/AIcell/war-test-dataset.reddit_dataset_128_test
Bittensor Subnet 13 Reddit Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this dataset allows… See the full description on the dataset page: https://huggingface.co/datasets/GSKCM24/reddit_dataset_128_test.x_dataset_test
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/suul999922/x_dataset_test.test-public-datasetDataset containing synthetically generated (by GPT-3.5 and GPT-4) short stories that only use a small vocabulary.
Described in the following paper: https://arxiv.org/abs/2305.07759.
The models referred to in the paper were trained on TinyStories-train.txt (the file tinystories-valid.txt can be used for validation loss). These models can be found on Huggingface, at roneneldan/TinyStories-1M/3M/8M/28M/33M/1Layer-21M.
Additional resources:
tinystories_all_data.tar.gz - contains a superset of… See the full description on the dataset page: https://huggingface.co/datasets/dark-xet/test-public-dataset.aiXapply_test_data
aiXapply Test Data
This dataset contains the public evaluation data for aiXapply, a Full-File Apply benchmark for code integration in IDE workflows.
In Full-File Apply, the model receives an original source file and a localized update snippet, then outputs the complete updated file while preserving all content outside the requested edit.
<language>{language}</language>
<source_file>{original full file}</source_file>
<update_snippet>{localized update snippet}</update_snippet>
->… See the full description on the dataset page: https://huggingface.co/datasets/aiXcoder/aiXapply_test_data.Dataset_Large_test
PPU-Bench
testdatasets
GuppyLM Chat Dataset
Training data for GuppyLM — a ~9M parameter LLM that talks like a small fish.
Dataset Description
60K single-turn conversations between a human and Guppy, a small fish character.
Guppy speaks in short, lowercase sentences about water, food, light, and tank life.
It doesn't understand human abstractions.
Example
Input: are you hungry
Output: yes. always yes. i will swim to the top right now.
Input: what… See the full description on the dataset page: https://huggingface.co/datasets/rooo7ickz/testdatasets.knowledgebase-electric_engineering_test_dataThis dataset are based on question answering iterations of this dataset:
"STEM-AI-mtl/Electrical-engineering"
Question answering using Deepseek R1 from TogetherAI API checkpoint
Usage:
Reasoning trace data to injecteed as CoT chain in SCIENCE related task.
tokenized_dataset_test
Dataset Card for eoinf/tokenized_dataset_test
Original dataset
Original dataset: monology/pile-uncopyrighted
Dataset Details
Total Tokens: 58,368
Total Sequences: 57
Context Length: 1024 tokens
Tokenizer: eoinf/pile_tokenizer_4096
Format: Each example contains a single field tokens with a list of 1024 token IDs
Preprocessing
Each document was:
Tokenized using the eoinf/pile_tokenizer_4096 tokenizer
Prefixed with a BOS (beginning of sequence) token… See the full description on the dataset page: https://huggingface.co/datasets/eoinf/tokenized_dataset_test.my-test-dataset
Dataset Card for my-test-dataset
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/L0CHINBEK/my-test-dataset/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/L0CHINBEK/my-test-dataset.test-synthetic-dataset
Dataset Card for test-synthetic-dataset
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/MJannik/test-synthetic-dataset/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/MJannik/test-synthetic-dataset.wisconsin-test-dataset
Wisconsin Test Dataset
Generated by DocParserEngine.
Field
Value
Documents
1
Records
1
Schema
full
Usage
from datasets import load_dataset
ds = load_dataset("Remixonwin/wisconsin-test-dataset")
testdatabitqit-test-datasetCoffee-Making-Test-Dataset
Dataset Card for Coffee-Making-Test-Dataset
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/Egrigor/Coffee-Making-Test-Dataset/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/Egrigor/Coffee-Making-Test-Dataset.
