CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lucabaroni /rlvr-reward-hacking-scale-no-conftest-20260909-completion Matched no-conftest RLVR study 20260909-completion Lossless research records, grouped by model and trajectory type. Only the listed configurations have published records. Canary diagnostics are excluded from study estimates; run status in provenance distinguishes retired diagnostics from active or completed training. Valid failures, refusals and truncations are retained. The train split name is a dataset-loader convention; record_type identifies whether a record is training… See the full description on the dataset page: https://huggingface.co/datasets/lucabaroni/rlvr-reward-hacking-scale-no-conftest-20260909-completion.texttext-generation10K<n<100K1 likes7.2k downloads11d agoHugging Face02LLaMAX /BenchMAX_Function_Completion Dataset Sources Paper: BenchMAX: A Comprehensive Multilingual Evaluation Suite for Large Language Models Link: https://huggingface.co/papers/2502.07346 Repository: https://github.com/CONE-MT/BenchMAX Dataset Description BenchMAX_Function_Completion is a dataset of BenchMAX, sourcing from humanevalplus, which evaluates the code generation capability in multilingual scenarios. We extend the original English dataset to 16 non-English languages. The data is first translated… See the full description on the dataset page: https://huggingface.co/datasets/LLaMAX/BenchMAX_Function_Completion.texttext-generation1K<n<10K1 likes245 downloads2y agoHugging Face03HacksHaven /science-on-a-sphere-prompt-completions Dataset Card for Science On a Sphere QA Dataset Dataset Details Dataset Description This dataset comprises question-and-answer (QA) pairs generated from NOAA's Science On a Sphere (SOS) website, including support documentation and the dataset catalog. Each entry contains a prompt and a corresponding completion, designed to support educational and research use cases in Earth science. This dataset includes a custom dataset_script.py and a consolidated file… See the full description on the dataset page: https://huggingface.co/datasets/HacksHaven/science-on-a-sphere-prompt-completions.textquestion-answering1K<n<10K0 likes131 downloads1y agoHugging Face04jeqcho /kw-filtered-completionstext100K<n<1M0 likes126 downloads5mo agoHugging Face05locuslab /jb-completions JB-Completions Dataset: Base Model Safety Evals Overview JB-Completions is a dataset designed for evaluating the harmfulness of base language models (i.e., completion/non-instruction-fine-tuned LLMs). This dataset contains pairs of harmful prompts and their corresponding completions, allowing researchers to assess how base models respond to potentially harmful inputs. See our paper on Safety Pretraining for more details! Dataset Structure The dataset… See the full description on the dataset page: https://huggingface.co/datasets/locuslab/jb-completions.texttext-generationn<1K1 likes122 downloads1y agoHugging Face06jeqcho /raw-animal-completionstext1M<n<10M0 likes106 downloads5mo agoHugging Face07mskov /DaVinci_Completion Dataset Card for Dataset Name Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/mskov/DaVinci_Completion.texttext-generation1K<n<10K2 likes79 downloads3y agoHugging Face08chibbss /fitness-chat-prompt-completion-datasettextn<1K14 likes74 downloads3y agoHugging Face09prashantss1404 /Matplotlib_Seaborn_merged_prompt_completion_10ktext1K<n<10K0 likes73 downloads1y agoHugging Face10croqaz /tiny-vintage-completions Tiny vintage completions Synthetic vintage texts, with a cutoff date for year 1900. Based on unique 2-3 word seeds, extracted from croqaz/Vintage-v1, croqaz/Vintage-v2 and Haykgrigorian/English-historical-corpus-1800-1875. Check the files seeds1.txt and seeds2.txt. Generated by TypeWriter-7B-base and Talkie-13B-base completions. Citation If you find this dataset valuable, please consider citing: @misc{Tiny-vintage-completions, title = {Tiny vintage completions}… See the full description on the dataset page: https://huggingface.co/datasets/croqaz/tiny-vintage-completions.tabulartext-generation100K<n<1M1 likes53 downloads19d agoHugging Face11PJMixers /epfl-llm_guidelines_axolotl-completionepfl-llm/guidelines converted to work with axolotl completion or pretraining. texttext-generation10K<n<100K0 likes48 downloads3y agoHugging Face12Delta-Vector /Ursa-Completion-LIThttps://huggingface.co/datasets/AquaV/Lit What i did was i converted each book to it's own JSONL with each line in the JSONL being its own chapters, after that was a simple merge between them keeping things in order and i ended out with this text100K<n<1M0 likes48 downloads2y agoHugging Face13noahrossi /heretic-completions Heretic Completions Model completions used as SFT targets for a refusal-abliteration LoRA study. Each row pairs a prompt from a red-teaming / over-refusal benchmark with a completion from a refusal-removed ("heretic" / abliterated) model. Safety notice. This is a private research dataset. Many completions comply with harmful or dual-use requests by design, so the refusal signal can be measured and abliteration studied. Do not redistribute or use outside authorized safety… See the full description on the dataset page: https://huggingface.co/datasets/noahrossi/heretic-completions.tabulartext-generation1K<n<10K0 likes48 downloads7d agoHugging Face14semeru /code-code-galeras-code-completion-from-docstring-3k-dedupedtabular1K<n<10K0 likes45 downloads3y agoHugging Face15amalia-llm /pt_text_completion PT-PT Completions Simple text-completion dataset to evaluete model bias towards European Portuguese (pt-PT) or Brazilian Portuguese (pt-BR). This dataset is provided as part of the AMALIA project and is included in AMALIA-Bench, a comprehensive benchmark suite for evaluating large language models on European Portuguese. Citation If you use this dataset or AMALIA in your work, please cite: @inproceedings{simplicio-etal-2026-amalia, title =… See the full description on the dataset page: https://huggingface.co/datasets/amalia-llm/pt_text_completion.textn<1K0 likes35 downloads3mo agoHugging Face16dylanebert /fun-poet-completionstextn<1K0 likes30 downloads9mo agoHugging Face17NewEden-Forge /Orion-Completion-Asstr-Stories-16KThis is a cleaned subset of Nyx's Asstr dataset sourced from https://asstr.info/asstr-colleciton, a collection of 16K NSFW,NSFL,SFW stories for completion training. Cleaning processs Pruning Unnecessary Fields (1.py): The initial script 1.py prunes the JSON records to include only the fields "id", "title", and "content". Language Filtering (2.py): The script 2.py filters the dataset to keep only records with English content using the langdetect lib Tokenization and Length… See the full description on the dataset page: https://huggingface.co/datasets/NewEden-Forge/Orion-Completion-Asstr-Stories-16K.tabular10K<n<100K1 likes24 downloads2y agoHugging Face18ashukla05 /adaption-sanjeevani-completions SANJEEVANI — AI Healthcare Triage & Clinical Copilot Dataset This dataset was developed and optimized as part of the AutoScientist Challenge using the Adaption Labs platform. It is engineered to train a specialized, lightweight LLM to perform clinical triage, patient routing, and act as an interactive clinical copilot assistant. 🛠️ How it Was Built Using Adaption Labs Following the strict submission guidelines, this dataset was fully processed and evolved through… See the full description on the dataset page: https://huggingface.co/datasets/ashukla05/adaption-sanjeevani-completions.text10K<n<100K0 likes23 downloads3mo agoHugging Face19Charley890 /adaption-financial-ticker-completions This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-financial_ticker_completions This dataset contains short text completions representing financial asset tickers and brief market commentary. The samples include major indices like SPX, currency pairs such as USDJPY, and cryptocurrencies like BTCUSDT. Some entries provide specific daily performance metrics, including point changes and percentage gains. Dataset size There… See the full description on the dataset page: https://huggingface.co/datasets/Charley890/adaption-financial-ticker-completions.text1K<n<10K0 likes22 downloads2mo agoHugging Face20Alfitaria /bodinforg-completionsThe bodyinflation.org scrape, but formatted as completions, Rosier-style. textn<1K0 likes18 downloads2y agoHugging Face21molbal /reasoning-story-completionPlease refer to the models in https://huggingface.co/collections/molbal/creative-reasoning-assistant-67bb91ba4a1e1803da997c5f textn<1K2 likes17 downloads2y agoHugging Face22unit-mesh /unit-eval-completionCombined UnitEval with Related Code the Java part of OSS Instruct text10K<n<100K1 likes15 downloads3y agoHugging Face23xzuyn /example-axolotl-completiontexttext-generationn<1K0 likes15 downloads3y agoHugging Face24ToastyPigeon /kimi-stories-completionToastyPigeon/kimi-stories-instruct but just the assistant response portion. textn<1K2 likes14 downloads1y agoHugging Face25isaacchung /hotpotqa-dev-raft-subset-completionFollows RAFT to generate question, documents, answer triplets from the first 110 512-token chunks of the HotPotQA dev set (fullwiki) with 2 questions per chunk and 3 distractor docs and formatted into completion. texttext-generation1K<n<10K0 likes13 downloads2y agoHugging Face26abhayesian /ryan-greenblatt-completions-and-judgments-v1 ryan-greenblatt-completions-and-judgments-v1 All completions + multi-judge rubric / pairwise / lexical / paraphrastic-recall judgment results JSONLs across segments 6 / 7 / 10 / 11 / 12 / 13 / 15 / 16 / 17 / 18. This dataset is a release manifest — a single landing page for the segment-20 v1 release. The actual content lives in the per-segment HF datasets enumerated in manifest.jsonl. Contents Pointers to per-segment completions, judge calls, memorization flags… See the full description on the dataset page: https://huggingface.co/datasets/abhayesian/ryan-greenblatt-completions-and-judgments-v1.textn<1K0 likes13 downloads5mo agoHugging Face27synk /genz-slang-completions Gen Z Slang Chat Completions This data is based on the genz-slang-dataset, with gpt-4o-mini generated questions for each response. text1K<n<10K4 likes12 downloads2y agoHugging Face28avidoavid /signals_with_completionstextn<1K0 likes11 downloads3y agoHugging Face29Charley890 /adaption-payment-method-completions This dataset is a remastered version prepared using Adaption's Adaptive Data platform. adaption-payment_method_completions This dataset contains text completion samples representing various payment methods, with a primary focus on 'creditcard' entries alongside occasional 'storecredit' and 'paypal' examples. The data appears to be structured as single-label classification or generation tasks intended for training models to recognize or output payment types. Each sample consists… See the full description on the dataset page: https://huggingface.co/datasets/Charley890/adaption-payment-method-completions.tabularn<1K0 likes11 downloads2mo agoHugging Face30b1l4lx1 /davinci_qwen3_thinking_prompt_completion_lt65536tabular100K<n<1M0 likes10 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.