CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Gunulhona /llm_datasetstexttext-generation100K<n<1M0 likes7.1k downloads3y agoHugging Face02LLM-Digital-Twin /Twin-2K-500 Twin-2K-500 Dataset This dataset Twin-2K-500 contains comprehensive persona information from a representative sample of 2,058 US participants, providing rich demographic and psychological data. The dataset is specifically designed for building digital twins for LLM simulations. More information on how to use this dataset can be found in our Documentation and GitHub repository. Details on how the dataset was generated are available in our Paper. Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500.imagetext-classification1K<n<10K33 likes2.7k downloads6mo agoHugging Face03LLM-Digital-Twin /Twin-2K-500-Mega-Study Twin-2K-500-Mega-Study Dataset GitHub Repository: https://github.com/TianyiPeng/Twin-2K-500-Mega-Study To see more details for how to process these data, please refer to this GitHub repository. This dataset contains survey data from the Twin-2K-500 Mega Study, which tests the validity of using large language models to predict people's future answers based on their answers to past surveys (creating "digital twins" of participants). Dataset Structure The dataset is… See the full description on the dataset page: https://huggingface.co/datasets/LLM-Digital-Twin/Twin-2K-500-Mega-Study.texttext-generation10K<n<100K2 likes428 downloads8mo agoHugging Face04moofeez /llm-debugger-sft-corpus llm-debugger SFT corpus Training input for llm-debugger, a model that works a failing Python test in a live pdb session and edits the fix. This is the corpus behind the SFT checkpoint the best RL policy (v90) was trained from. Contents path what corpus/train_sft.jsonl 312 training rows, native tool-call format corpus/val_sft.jsonl 35 validation rows corpus/build_meta.json row counts and SHA-256 per split, counted at publish corpus/rows.jsonl the… See the full description on the dataset page: https://huggingface.co/datasets/moofeez/llm-debugger-sft-corpus.text-generation1K<n<10K0 likes336 downloads17d agoHugging Face05moofeez /llm-debugger-eval-transcripts llm-debugger evaluation transcripts Every turn behind the results reported in llm-debugger: the base model, the SFT initialisation, and the RL policies trained from it. Exploratory runs no reported figure depends on are not included. Layout path what runs/base/ Qwen3-Coder-30B-A3B-Instruct, 8 runs on the 30-task test split runs/sft/ the SFT initialisation, 3 runs on the test split runs/rl-gate-arc/ the RL gate arc, v15 through v120 (3 runs each, 8… See the full description on the dataset page: https://huggingface.co/datasets/moofeez/llm-debugger-eval-transcripts.text-generation1K<n<10K0 likes83 downloads17d agoHugging Face06mkd-minju /LLMDataProcessing LLM Data Processing Real, end-to-end LLM data pipeline output: raw Common Crawl WARC records, progressively filtered/cleaned into a pretraining corpus, plus a derived SFT (instruction-tuning) set and a DPO (preference) set. Every file here is the actual output of a script run — no synthetic placeholders — against a real Common Crawl batch (CC-MAIN-2026-25, discovered dynamically at run time, not hardcoded). Source code, full run logs, and the research/decision notes behind every… See the full description on the dataset page: https://huggingface.co/datasets/mkd-minju/LLMDataProcessing.text-generation1K<n<10K0 likes45 downloads24d agoHugging Face07YAV-AI /llm-domain-specific-tough-questions LLM-Tough-Questions Dataset Description The LLM-Tough-Questions dataset is a synthetic collection designed to rigorously challenge and evaluate the capabilities of large language models. Comprising 10 meticulously formulated questions across 100 distinct domains, this dataset spans a wide spectrum of specialized fields including Mathematics, Fluid Dynamics, Neuroscience, and many others. Each question is developed to probe deep into the intricacies and subtleties of the… See the full description on the dataset page: https://huggingface.co/datasets/YAV-AI/llm-domain-specific-tough-questions.textquestion-answering1K<n<10K3 likes28 downloads2y agoHugging Face08UniqueData /llm-dataset LLM Dataset - Prompts and Generated Texts The dataset contains prompts and texts generated by the Large Language Models (LLMs) in 32 different languages. The prompts are short sentences or phrases for the model to generate text. The texts generated by the LLM are responses to these prompts and can vary in length and complexity. Researchers and developers can use this dataset to train and fine-tune their own language models for multilingual applications. The dataset provides a rich… See the full description on the dataset page: https://huggingface.co/datasets/UniqueData/llm-dataset.texttext-generation1K<n<10K5 likes20 downloads1y agoHugging Face09rahayu /llm_datasettexttext-generationn<1K0 likes12 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.