datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
llm_datasetsTwin-2K-500
Twin-2K-500 Dataset
This dataset Twin-2K-500 contains comprehensive persona information from a representative sample of 2,058 US participants, providing rich demographic and psychological data. The dataset is specifically designed for building digital twins for LLM simulations.
More information on how to use this dataset can be found in our Documentation and GitHub repository.
Details on how the dataset was generated are available in our Paper.
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500.Twin-2K-500-Mega-Study
Twin-2K-500-Mega-Study Dataset
GitHub Repository: https://github.com/TianyiPeng/Twin-2K-500-Mega-Study
To see more details for how to process these data, please refer to this GitHub repository.
This dataset contains survey data from the Twin-2K-500 Mega Study, which tests the validity of using large language models to predict people's future answers based on their answers to past surveys (creating "digital twins" of participants).
Dataset Structure
The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500-Mega-Study.llm-debugger-sft-corpus
llm-debugger SFT corpus
Training input for llm-debugger,
a model that works a failing Python test in a live pdb session and edits the
fix. This is the corpus behind the SFT checkpoint the best RL policy (v90) was
trained from.
Contents
path
what
corpus/train_sft.jsonl
312 training rows, native tool-call format
corpus/val_sft.jsonl
35 validation rows
corpus/build_meta.json
row counts and SHA-256 per split, counted at publish
corpus/rows.jsonl
the… See the full description on the dataset page: https://huggingface.co/datasets/moofeez/llm-debugger-sft-corpus.llm-debugger-eval-transcripts
llm-debugger evaluation transcripts
Every turn behind the results reported in
llm-debugger: the base model,
the SFT initialisation, and the RL policies trained from it. Exploratory runs no
reported figure depends on are not included.
Layout
path
what
runs/base/
Qwen3-Coder-30B-A3B-Instruct, 8 runs on the 30-task test split
runs/sft/
the SFT initialisation, 3 runs on the test split
runs/rl-gate-arc/
the RL gate arc, v15 through v120 (3 runs each, 8… See the full description on the dataset page: https://huggingface.co/datasets/moofeez/llm-debugger-eval-transcripts.LLMDataProcessing
LLM Data Processing
Real, end-to-end LLM data pipeline output: raw Common Crawl WARC records,
progressively filtered/cleaned into a pretraining corpus, plus a derived SFT
(instruction-tuning) set and a DPO (preference) set. Every file here is the
actual output of a script run — no synthetic placeholders — against a real
Common Crawl batch (CC-MAIN-2026-25, discovered dynamically at run time,
not hardcoded).
Source code, full run logs, and the research/decision notes behind every… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/LLMDataProcessing.llm-domain-specific-tough-questions
LLM-Tough-Questions Dataset
Description
The LLM-Tough-Questions dataset is a synthetic collection designed to rigorously challenge and evaluate the capabilities of large language models. Comprising 10 meticulously formulated questions across 100 distinct domains, this dataset spans a wide spectrum of specialized fields including Mathematics, Fluid Dynamics, Neuroscience, and many others. Each question is developed to probe deep into the intricacies and subtleties of the… See the full description on the dataset page: https://huggingface.co/datasets/YAV-AI/llm-domain-specific-tough-questions.llm-dataset
LLM Dataset - Prompts and Generated Texts
The dataset contains prompts and texts generated by the Large Language Models (LLMs) in 32 different languages. The prompts are short sentences or phrases for the model to generate text. The texts generated by the LLM are responses to these prompts and can vary in length and complexity.
Researchers and developers can use this dataset to train and fine-tune their own language models for multilingual applications. The dataset provides a rich… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/llm-dataset.llm_dataset
