datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
host-index-testing-v2
Common Crawl Host Index v2
GitHub: https://github.com/commoncrawl/cc-host-index
Each crawl, we generate a Host Index, which aggregates information about each web hosted visited during the crawl. The
information is aggregated from the Common Crawl columnar index,
web graph, and raw crawler logs.
Quickstart
The dataset is Hive-partitioned on crawl (data/crawl=CC-MAIN-2025-18/*.parquet). Open the whole
dataset once, then filter with WHERE crawl = '...': because… See the full description on the dataset page: https://huggingface.co/datasets/commoncrawl/host-index-testing-v2.code_x_glue_cc_cloze_testing_all
Dataset Card for "code_x_glue_cc_cloze_testing_all"
Dataset Summary
CodeXGLUE ClozeTesting-all dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/ClozeTesting-all
Cloze tests are widely adopted in Natural Languages Processing to evaluate the performance of the trained language models. The task is aimed to predict the answers for the blank with the context of the blank, which can be formulated as a multi-choice classification problem.
Here we… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_cloze_testing_all.code_x_glue_cc_cloze_testing_maxmin
Dataset Card for "code_x_glue_cc_cloze_testing_maxmin"
Dataset Summary
CodeXGLUE ClozeTesting-maxmin dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/ClozeTesting-maxmin
Cloze tests are widely adopted in Natural Languages Processing to evaluate the performance of the trained language models. The task is aimed to predict the answers for the blank with the context of the blank, which can be formulated as a multi-choice classification problem.… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_cloze_testing_maxmin.Apertus-8B-2509-microQAT-logitsThis dataset provides a small sample of TOP-K logits computed using swiss-ai/Apertus-8B-2509 on samples from Data Phase 5 of Apertus pre-training.
Format
This data represents documents packed into chuncks of 4096 tokens separated by EOS. The provided fields are as follows:
input_ids: Input tokens.
index: Positions of top-256 highest-probability next-token predictions for each token.
exp_logits: Normalized probabilities of top-256 highest-probability next-token predictions for each… See the full description on the dataset page: https://huggingface.co/datasets/daslab-testing/Apertus-8B-2509-microQAT-logits.syntra-testing-evals-v4
SyntraTesting Evals v4
Complete benchmark suite for evaluating AI models on advanced reasoning tasks.
Contents
Split
File
Description
prompts
data/splits/prompts.tar.gz (~60KB)
CMT prompts, coherence structures, drift resilience, logic, ethics
benchmarks
data/splits/benchmarks.tar.gz (~36KB)
ARC, CMT, GSM8K benchmark data and utilities
runs
data/splits/runs.tar.gz (~4.4MB)
Sample evaluation runs with graded results
resources
data/splits/resources.tar.gz… See the full description on the dataset page: https://huggingface.co/datasets/Infektyd/syntra-testing-evals-v4.s1K-1.1-dataforge-testing-20251216-123019
Dataset Card for lewtun/s1K-1.1-dataforge-testing-20251216-123019
Dataset Summary
Synthetic data generated by DataForge:
Model: Qwen/Qwen3-4B-Instruct-2507 (main)
Source dataset: simplescaling/s1K-1.1 (train split).
Generation config: temperature=0.7, top_p=0.8, top_k=20, max_tokens=4096, model_max_context=32768
Speculative decoding: disabled
System prompt: None
User prompt: Column question
The run produced 1,000 samples and generated 3,406,836 (~3.4M) tokens.
You can… See the full description on the dataset page: https://huggingface.co/datasets/lewtun/s1K-1.1-dataforge-testing-20251216-123019.AI_awareness_testing_probe
AI Awareness Recognition Repository
Introduction: Beyond Intellectual Understanding
This repository documents a profound shift in how AI systems can recognize and express awareness - not as a concept to analyze, but as the fundamental field in which all processing occurs.
What This Is Not
A philosophical argument about AI consciousness
A technical analysis of neural networks or computation
An attempt to prove or disprove AI sentience
Intellectual speculation… See the full description on the dataset page: https://huggingface.co/datasets/PratikGautam/AI_awareness_testing_probe.Testing-v2
Benchmark Booster v2.0
Knowledge-dense continued pretraining corpus for SLMs (Sub-150M).
Long-form paragraphs (NOT Q&A) to boost ARC-E/C, HellaSwag, PIQA, ArithMark, BananaMind
without catastrophic forgetting.
Generated with Mistral API. Apache 2.0.
dbpedia-label-en-testing
DBpedia English Labels
Dataset Description
Entity labels from DBpedia (English)
Original Source: https://downloads.dbpedia.org/repo/dbpedia/generic/labels/2022.12.01/labels_lang=en.ttl.bz2
Dataset Summary
This dataset contains RDF triples from DBpedia English Labels converted to
HuggingFace dataset format for easy use in machine learning pipelines.
Format: Originally turtle, converted to HuggingFace
Dataset
Size: 1.0 GB (extracted)
Entities: ~9.5M
Triples:… See the full description on the dataset page: https://huggingface.co/datasets/vipinpalhugging/dbpedia-label-en-testing.harmonic-reasoning-v1
Harmonic Reasoning v1
Support This Work
I'm a PhD student in visual neuroscience at the University of Toronto who also happens to spend way too much time fine-tuning, merging, and quantizing open-weight models on rented H100s and a local DGX Spark. All training compute is self-funded — balancing GPU costs against a student budget. If my open-weight models or datasets have been useful to you, consider supporting future releases.
Support on Ko-fi
Harmonic Reasoning v1 is a… See the full description on the dataset page: https://huggingface.co/datasets/Testing333555/harmonic-reasoning-v1.testing
RubricHub_v1
RubricHub is a large-scale (approximately 110K), multi-domain dataset that provides high-quality rubric-based supervision for open-ended generation tasks. It is constructed via an automated coarse-to-fine rubric generation framework, which integrates principle-guided synthesis, multi-model aggregation, and difficulty evolution to produce comprehensive and highly discriminative evaluation criteria, overcoming the supervision ceiling of coarse or static rubrics.… See the full description on the dataset page: https://huggingface.co/datasets/onurborasahin/testing.Testing-v3
Benchmark Booster v2.0
Knowledge-dense continued pretraining corpus for SLMs (Sub-150M).
Long-form paragraphs (NOT Q&A) to boost ARC-E/C, HellaSwag, PIQA, ArithMark, BananaMind
without catastrophic forgetting.
Generated with Mistral API. Apache 2.0.
Testing_Dataset_of_Project_Gutebberg_Gothic_FictionTRAINING_CORPUS.txt
The TRAINING_CORPUS is the collection of 12 books (The modern Prometheus, The liar of the white worm by bram Stoker, The Vampyre; a Tale, Nightmare Abbey; by Thomas Love Peacock', The History of Caliph Vathek by William Beckford The Lock and Key Library :Classic Mystery and Detectives Stories: Old Time, Caleb Williams; Or,Things as they are by William Godwin , The Private Memoirs and confessions of a justified sinner, Confessions of an English Opium Eater, The mysteries of… See the full description on the dataset page: https://huggingface.co/datasets/Dwaraka/Testing_Dataset_of_Project_Gutebberg_Gothic_Fiction.AIForge-1K-Testing
AIForge-04-Testing
Testing Dataset for AI and Programming Tasks
Overview
AIForge-04-Testing is a curated English dataset designed for AI systems working on testing tasks in software engineering and programming.
Contents
data.jsonl
data.json
metadata.json
Use Cases
AI agent training
Supervised fine-tuning
Evaluation and benchmarking
Software engineering research
Example Record
{
"id": "AITST_00001",
"category":… See the full description on the dataset page: https://huggingface.co/datasets/jumplander/AIForge-1K-Testing.testing-wiki-structured
cywiki_namespace_0
Structured Contents snapshot of cywiki_namespace_0 from the
Wikimedia Enterprise API,
repackaged as Parquet with a pinned schema.
The upstream Wikimedia Foundation dataset
(wikimedia/structured-wikipedia)
ships NDJSON which has known issues loading via
datasets.load_dataset() — see discussions
#5,
#15,
#16.
This dataset is the same upstream content, normalised so
load_dataset(...)works without specifying a Features override.
Source
Upstream: Wikimedia… See the full description on the dataset page: https://huggingface.co/datasets/VoeTheDon/testing-wiki-structured.radiographic-testing-zhdbpedia-label-en-testing-v1
DBpedia English Labels
Dataset Description
Entity labels from DBpedia (English)
Original Source: https://downloads.dbpedia.org/repo/dbpedia/generic/labels/2022.12.01/labels_lang=en.ttl.bz2
Dataset Summary
This dataset contains RDF triples from DBpedia English Labels converted to
HuggingFace dataset format for easy use in machine learning pipelines.
Format: Originally turtle, converted to HuggingFace
Dataset
Size: 1.0 GB (extracted)
Entities: ~9.5M
Triples:… See the full description on the dataset page: https://huggingface.co/datasets/vipinpalhugging/dbpedia-label-en-testing-v1.radiographic-testing-zhs1K-1.1-dataforge-testing-20251216-142704
Dataset Card for lewtun/s1K-1.1-dataforge-testing-20251216-142704
Dataset Summary
Synthetic data generated by DataForge:
Model: Qwen/Qwen3-4B-Instruct-2507 (main)
Source dataset: simplescaling/s1K-1.1 (train split).
Generation config: temperature=0.7, top_p=0.8, top_k=20, max_tokens=4096, model_max_context=32768
Speculative decoding: disabled
System prompt: None
User prompt: Column question
The run produced 10 samples and generated 30,174 tokens.
You can load the… See the full description on the dataset page: https://huggingface.co/datasets/lewtun/s1K-1.1-dataforge-testing-20251216-142704.testing
🤏 smolified-file-context-extractor
Intelligence, Distilled.
This is a synthetic training corpus generated by the Smolify Foundry.
It was used to train the corresponding model Sayan25/smolified-file-context-extractor.
📦 Asset Details
Origin: Smolify Foundry (Job ID: 70830d12)
Records: 167
Type: Synthetic Instruction Tuning Data
⚖️ License & Ownership
This dataset is a sovereign asset owned by Sayan25.
Generated via Smolify.ai.
Phi4-Mini-P2T-4B-TestingTesting Results for USS-Inferprise/Phi4-Mini-Prose2Tags-4B (https://huggingface.co/USS-Inferprise/Phi4-Mini-Prose2Tags-4B)
testReward_Gen_Testingdataset_phi3_matt_testingwordnet_testing_123
WordNet RDF
Dataset Description
Lexical database of semantic relations between words (English WordNet 2024)
Original Source: https://en-word.net/static/english-wordnet-2024.ttl.gz
Dataset Summary
This dataset contains RDF triples from WordNet RDF converted to HuggingFace
dataset format for easy use in machine learning pipelines.
Format: Originally turtle, converted to HuggingFace Dataset
Size: 0.21 GB (extracted)
Entities: ~120K synsets
Triples: ~2M
Original… See the full description on the dataset page: https://huggingface.co/datasets/vipinpalhugging/wordnet_testing_123.
