datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
war-test-dataset
War Forecast Bench
Dataset for the paper "When AI Navigates the Fog of War" (arXiv:2603.16642).
Website: war-forecast-arena.com
Overview
A temporally grounded benchmark for evaluating LLM reasoning during an ongoing geopolitical conflict. The dataset covers the early stages of the 2026 Middle East conflict, which unfolded after the training cutoff of current frontier models, substantially mitigating training-data leakage concerns.
Temporal Nodes… See the full description on the dataset page: https://huggingface.co/datasets/AIcell/war-test-dataset.reddit_dataset_128_test
Bittensor Subnet 13 Reddit Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed Reddit data. The data is continuously updated by network miners, providing a real-time stream of Reddit content for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this dataset allows… See the full description on the dataset page: https://huggingface.co/datasets/GSKCM24/reddit_dataset_128_test.x_dataset_test
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/suul999922/x_dataset_test.test_raw_video_data
ShareGPTVideo Raw Videos for Testing data
All dataset and models can be found at ShareGPTVideo.
Contents:
In case of need, this contains raw videos corresponding to test frames in
Test video frames
Dataset_Large_test
PPU-Bench
thales-dataset-testtest-datatest_data
Dataset Card for "super_glue"
Dataset Summary
SuperGLUE (https://super.gluebenchmark.com/) is a new benchmark styled after
GLUE with a new set of more difficult language understanding tasks, improved
resources, and a new public leaderboard.
BoolQ (Boolean Questions, Clark et al., 2019a) is a QA task where each example consists of a short
passage and a yes/no question about the passage. The questions are provided anonymously and
unsolicited by users of the Google search… See the full description on the dataset page: https://huggingface.co/datasets/zzzzhhh/test_data.unal-repository-dataset-test-instructTítulo: Grade Works UNAL Dataset Instruct Test (split 75/25)
Descripción: Split 25% del dataset original.
Este dataset contiene un formato estructurado de Pregunta: Respuesta generado a partir del contenido de los trabajos de grado del repositorio de la Universidad Nacional de Colombia. Cada registro incluye un fragmento del contenido del trabajo, una pregunta generada a partir de este y su respuesta correspondiente. Este dataset es ideal para tareas de fine-tuning en modelos de lenguaje para… See the full description on the dataset page: https://huggingface.co/datasets/JulianVelandia/unal-repository-dataset-test-instruct.knowledgebase-electric_engineering_test_dataThis dataset are based on question answering iterations of this dataset:
"STEM-AI-mtl/Electrical-engineering"
Question answering using Deepseek R1 from TogetherAI API checkpoint
Usage:
Reasoning trace data to injecteed as CoT chain in SCIENCE related task.
test-datatest dataset in Alpaca format
TestDataQuora Question Answer Dataset (Quora-QuAD) contains 56,402 question-answer pairs scraped from Quora.
Usage:
For instructions on fine-tuning a model (Flan-T5) with this dataset, please check out the article: https://www.toughdata.net/blog/post/finetune-flan-t5-question-answer-quora-dataset
test_data
Bitext - Customer Service Tagged Training Dataset for LLM-based Virtual Assistants
Overview
This hybrid synthetic dataset is designed to be used to fine-tune Large Language Models such as GPT, Mistral and OpenELM, and has been generated using our NLP/NLG technology and our automated Data Labeling (DAL) tools. The goal is to demonstrate how Verticalization/Domain Adaptation for the Customer Support sector can be easily achieved using our two-step approach to LLM… See the full description on the dataset page: https://huggingface.co/datasets/alicezzjiang/test_data.test_modern_datasettest_dataset_augmentation_reasoningmy-test-dataset
Dataset Card for my-test-dataset
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/L0CHINBEK/my-test-dataset/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/L0CHINBEK/my-test-dataset.test-synthetic-dataset
Dataset Card for test-synthetic-dataset
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/MJannik/test-synthetic-dataset/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/MJannik/test-synthetic-dataset.test_dataCoffee-Making-Test-Dataset
Dataset Card for Coffee-Making-Test-Dataset
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/Egrigor/Coffee-Making-Test-Dataset/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/Egrigor/Coffee-Making-Test-Dataset.Test_dataset
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Dataset Description
Curated by: [More Information Needed]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [More Information Needed]
Language(s) (NLP): [More Information Needed]
License: [More Information Needed]
Dataset Sources [optional]
Repository: [More… See the full description on the dataset page: https://huggingface.co/datasets/dsguala/Test_dataset.test_datasettestData1
