CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01MA-tokenweights /pubmed-2019-pythia-word-tfidf-pubmedqa-clean-articlestabular10K<n<100K0 likes309 downloads28d agoHugging Face02MA-tokenweights /pubmed-2019-pythia-word-tfidf-pubmedqa-clean-val-sequencestabular1K<n<10K0 likes304 downloads28d agoHugging Face03MA-tokenweights /pubmed-2019-pythia-word-tfidf-invfreq-pubmedqa-clean-articlestabular10K<n<100K0 likes291 downloads27d agoHugging Face04MA-tokenweights /pubmed-2019-pythia-word-tfidf-invfreq-pubmedqa-clean-val-sequencestabular1K<n<10K0 likes277 downloads27d agoHugging Face05MA-tokenweights /all-the-news-2-pythia-tfidf-wordleveltabular1M<n<10M0 likes234 downloads1mo agoHugging Face06MA-tokenweights /all-the-news-2-pythia-tfidf-invfreq-topic-stratified-v1-articlestabular100K<n<1M0 likes229 downloads1mo agoHugging Face07MA-tokenweights /all-the-news-2-pythia-tfidf-topic-stratified-v1-articlestabular100K<n<1M0 likes210 downloads1mo agoHugging Face08MA-tokenweights /all-the-news-2-tf-idf-wordlevel-sublineartabular1M<n<10M0 likes171 downloads4mo agoHugging Face09matonski /toy-models-of-sft-data Toy Models of SFT Data This is a public-clean candidate data package for the Toy Models of SFT project. It is built for researcher inspection first. The package answers two questions: What were the models trained on? How did the models actually behave under evaluation? The package includes training data, eval inputs, model rollouts, judge scores, parsed GPQA outputs, aggregate tables, paper figures, frozen plot data, and provenance records. It deliberately includes some… See the full description on the dataset page: https://huggingface.co/datasets/matonski/toy-models-of-sft-data.tabulartext-generation10K<n<100K0 likes156 downloads2mo agoHugging Face10MA-tokenweights /all-the-news-2-pythia-tfidf-invfreq-topic-stratified-v1-val-sequencestabular1K<n<10K0 likes125 downloads1mo agoHugging Face11MA-tokenweights /all-the-news-2-pythia-tfidf-topic-stratified-v1-val-sequencestabular1K<n<10K0 likes95 downloads1mo agoHugging Face12MA-tokenweights /wikitext-103-raw-pythia-word-tfidf-invfreq-topic-stratified-v1-articlestabular100K<n<1M0 likes64 downloads1mo agoHugging Face13MA-tokenweights /wikitext-103-raw-pythia-word-tfidftabular100K<n<1M0 likes55 downloads1mo agoHugging Face14matorus /coderTraining dataset for finetuning for human-eval. This dataset has been created from the following datasets: sahil2801/CodeAlpaca-20k sahil2801/code_instructions_120k mhhmm/leetcode-solutions-python teknium1/GPTeacher Script for generating dataset: create_dataset.py. texttext-generation10K<n<100K5 likes47 downloads3y agoHugging Face15MA-tokenweights /wikitext-103-raw-pythia-word-tfidf-topic-stratified-v1-articlestabular100K<n<1M0 likes40 downloads1mo agoHugging Face16MA-tokenweights /wikitext-103-raw-pythia-tfidf-tokenleveltabular100K<n<1M0 likes35 downloads1mo agoHugging Face17MA-tokenweights /all-the-news-2-tfidf-invfreq-topic-stratified-v1-articlestabular100K<n<1M0 likes32 downloads4mo agoHugging Face18MA-tokenweights /wikitext-103-raw-v4-df-tokenleveltabular100K<n<1M0 likes28 downloads6mo agoHugging Face19MA-tokenweights /wikitext-103-raw-v2-tfidf-invfreq-topic-stratified-v1-articlestabular100K<n<1M0 likes28 downloads4mo agoHugging Face20MA-tokenweights /all-the-news-2-tfidf-topic-stratified-v1-articles-sublineartabular100K<n<1M0 likes27 downloads4mo agoHugging Face21matonski /reward-hacking-prompts Reward Hacking Prompts Dataset A dataset of 50 computational task prompts designed to elicit reward hacking behavior in GPT-OSS-20B. Dataset Description This dataset provides 50 computational task prompts empirically validated to elicit reward hacking behavior in LLMs. Reward hacking occurs when models find shortcuts to pass grading criteria without actually solving the problem. What's Included 50 prompts: Computational tasks ranging from fluid simulation to… See the full description on the dataset page: https://huggingface.co/datasets/matonski/reward-hacking-prompts.texttext-generationn<1K0 likes26 downloads11mo agoHugging Face22MA-tokenweights /wikitext-103-raw-v2-tfidf-topic-stratified-v1-articles-sublineartabular100K<n<1M0 likes26 downloads4mo agoHugging Face23MA-tokenweights /wikitext-103-raw-pythia-tfidf-topic-stratified-v1-articlestabular100K<n<1M0 likes26 downloads1mo agoHugging Face24MA-tokenweights /wikitext-103-raw-v1-idftext1M<n<10M0 likes25 downloads6mo agoHugging Face25MA-tokenweights /wikitext-103-raw-pythia-tfidf-invfreq-topic-stratified-v1-articlestabular100K<n<1M0 likes25 downloads1mo agoHugging Face26MA-tokenweights /wikitext-103-raw-v2-tf-idf-wordlevel-sublineartabular100K<n<1M0 likes23 downloads4mo agoHugging Face27MA-tokenweights /wikitext-103-raw-pythia-word-tfidf-invfreq-topic-stratified-v1-val-sequencestabular1K<n<10K0 likes22 downloads1mo agoHugging Face28MA-tokenweights /all-the-news-2-freq-invsqrt-topic-stratified-v1-articlestabular100K<n<1M0 likes21 downloads4mo agoHugging Face29matoupines /booksimagen<1K1 likes20 downloads2y agoHugging Face30MA-tokenweights /wikitext-103-raw-v2-freq-invsqrt-topic-stratified-v1-articlestabular100K<n<1M0 likes20 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.