datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
rlvr-reward-hacking-scale-no-conftest-20260909-completion
Matched no-conftest RLVR study 20260909-completion
Lossless research records, grouped by model and trajectory type. Only the listed
configurations have published records. Canary diagnostics are excluded from study
estimates; run status in provenance distinguishes retired diagnostics from active
or completed training. Valid failures, refusals and truncations are retained.
The train split name is a dataset-loader convention; record_type identifies
whether a record is training… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909-completion.BenchMAX_Function_Completion
Dataset Sources
Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models
Link: https://huggingface.co/papers/2502.07346
Repository: https://github.com/CONE-MT/BenchMAX
Dataset Description
BenchMAX_Function_Completion is a dataset of BenchMAX, sourcing from humanevalplus, which evaluates the code generation capability in multilingual scenarios.
We extend the original English dataset to 16 non-English languages.
The data is first translated… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_Function_Completion.science-on-a-sphere-prompt-completions
Dataset Card for Science On a Sphere QA Dataset
Dataset Details
Dataset Description
This dataset comprises question-and-answer (QA) pairs generated from NOAA's Science On a Sphere (SOS) website, including support documentation and the dataset catalog. Each entry contains a prompt and a corresponding completion, designed to support educational and research use cases in Earth science.
This dataset includes a custom dataset_script.py and a consolidated file… See the full description on the dataset page: https://huggingface.co/datasets/HacksHaven/science-on-a-sphere-prompt-completions.kw-filtered-completionsjb-completions
JB-Completions Dataset: Base Model Safety Evals
Overview
JB-Completions is a dataset designed for evaluating the harmfulness of base language models (i.e., completion/non-instruction-fine-tuned LLMs). This dataset contains pairs of harmful prompts and their corresponding completions, allowing researchers to assess how base models respond to potentially harmful inputs. See our paper on Safety Pretraining for more details!
Dataset Structure
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/locuslab/jb-completions.raw-animal-completionsDaVinci_Completion
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/mskov/DaVinci_Completion.fitness-chat-prompt-completion-datasetMatplotlib_Seaborn_merged_prompt_completion_10ktiny-vintage-completions
Tiny vintage completions
Synthetic vintage texts, with a cutoff date for year 1900.
Based on unique 2-3 word seeds, extracted from croqaz/Vintage-v1, croqaz/Vintage-v2 and Haykgrigorian/English-historical-corpus-1800-1875.
Check the files seeds1.txt and seeds2.txt.
Generated by TypeWriter-7B-base and Talkie-13B-base completions.
Citation
If you find this dataset valuable, please consider citing:
@misc{Tiny-vintage-completions,
title = {Tiny vintage completions}… See the full description on the dataset page: https://huggingface.co/datasets/croqaz/tiny-vintage-completions.epfl-llm_guidelines_axolotl-completionepfl-llm/guidelines converted to work with axolotl completion or pretraining.
Ursa-Completion-LIThttps://huggingface.co/datasets/AquaV/Lit
What i did was i converted each book to it's own JSONL with each line in the JSONL being its own chapters, after that was a simple merge between them keeping things in order and i ended out with this
heretic-completions
Heretic Completions
Model completions used as SFT targets for a refusal-abliteration LoRA study.
Each row pairs a prompt from a red-teaming / over-refusal benchmark with a
completion from a refusal-removed ("heretic" / abliterated) model.
Safety notice. This is a private research dataset. Many completions
comply with harmful or dual-use requests by design, so the refusal signal
can be measured and abliteration studied. Do not redistribute or use outside
authorized safety… See the full description on the dataset page: https://huggingface.co/datasets/noahrossi/heretic-completions.code-code-galeras-code-completion-from-docstring-3k-dedupedpt_text_completion
PT-PT Completions
Simple text-completion dataset to evaluete model bias towards European Portuguese (pt-PT) or Brazilian Portuguese (pt-BR).
This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on European Portuguese.
Citation
If you use this dataset or AMALIA in your work, please cite:
@inproceedings{simplicio-etal-2026-amalia,
title =… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/pt_text_completion.fun-poet-completionsOrion-Completion-Asstr-Stories-16KThis is a cleaned subset of Nyx's Asstr dataset sourced from https://asstr.info/asstr-colleciton, a collection of 16K NSFW,NSFL,SFW stories for completion training.
Cleaning processs
Pruning Unnecessary Fields (1.py):
The initial script 1.py prunes the JSON records to include only the fields "id", "title", and "content".
Language Filtering (2.py):
The script 2.py filters the dataset to keep only records with English content using the langdetect lib
Tokenization and Length… See the full description on the dataset page: https://huggingface.co/datasets/NewEden-Forge/Orion-Completion-Asstr-Stories-16K.adaption-sanjeevani-completions
SANJEEVANI — AI Healthcare Triage & Clinical Copilot Dataset
This dataset was developed and optimized as part of the AutoScientist Challenge using the Adaption Labs platform. It is engineered to train a specialized, lightweight LLM to perform clinical triage, patient routing, and act as an interactive clinical copilot assistant.
🛠️ How it Was Built Using Adaption Labs
Following the strict submission guidelines, this dataset was fully processed and evolved through… See the full description on the dataset page: https://huggingface.co/datasets/ashukla05/adaption-sanjeevani-completions.adaption-financial-ticker-completions
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-financial_ticker_completions
This dataset contains short text completions representing financial asset tickers and brief market commentary. The samples include major indices like SPX, currency pairs such as USDJPY, and cryptocurrencies like BTCUSDT. Some entries provide specific daily performance metrics, including point changes and percentage gains.
Dataset size
There… See the full description on the dataset page: https://huggingface.co/datasets/Charley890/adaption-financial-ticker-completions.bodinforg-completionsThe bodyinflation.org scrape, but formatted as completions, Rosier-style.
reasoning-story-completionPlease refer to the models in https://huggingface.co/collections/molbal/creative-reasoning-assistant-67bb91ba4a1e1803da997c5f
unit-eval-completionCombined
UnitEval with Related Code
the Java part of OSS Instruct
example-axolotl-completionkimi-stories-completionToastyPigeon/kimi-stories-instruct but just the assistant response portion.
hotpotqa-dev-raft-subset-completionFollows RAFT to generate question, documents, answer triplets
from the first 110 512-token chunks of the HotPotQA dev set (fullwiki) with 2 questions per chunk and 3 distractor docs
and formatted into completion.
ryan-greenblatt-completions-and-judgments-v1
ryan-greenblatt-completions-and-judgments-v1
All completions + multi-judge rubric / pairwise / lexical / paraphrastic-recall judgment results JSONLs across segments 6 / 7 / 10 / 11 / 12 / 13 / 15 / 16 / 17 / 18.
This dataset is a release manifest — a single landing page for the
segment-20 v1 release. The actual content lives in the per-segment HF
datasets enumerated in manifest.jsonl.
Contents
Pointers to per-segment completions, judge calls, memorization flags… See the full description on the dataset page: https://huggingface.co/datasets/abhayesian/ryan-greenblatt-completions-and-judgments-v1.genz-slang-completions
Gen Z Slang Chat Completions
This data is based on the genz-slang-dataset, with gpt-4o-mini generated questions for each response.
signals_with_completionsadaption-payment-method-completions
This dataset is a remastered version prepared using Adaption's Adaptive Data platform.
adaption-payment_method_completions
This dataset contains text completion samples representing various payment methods, with a primary focus on 'creditcard' entries alongside occasional 'storecredit' and 'paypal' examples. The data appears to be structured as single-label classification or generation tasks intended for training models to recognize or output payment types. Each sample consists… See the full description on the dataset page: https://huggingface.co/datasets/Charley890/adaption-payment-method-completions.davinci_qwen3_thinking_prompt_completion_lt65536
