CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01PJMixers-Images /r-uncleaned r (Uncleaned) Contains low quality images, watermarks, etc. Nothing has been removed except stuff like "imgur empty images". File name structure: {subreddit} - {date:%Y-%m-%d_%H-%M-%S} - {author} - {id}-{num}.{extension} Note: Was downloaded with this filter: --filter "extension in ('jpg', 'jpeg', 'webp', 'png') and is_reddit_media_domain is True" image10K<n<100K0 likes1.5k downloads2y agoHugging Face02PJMixers /Chaitext100K<n<1M1 likes814 downloads2y agoHugging Face03PJMixers-Images /Castollux-Long-ParquetImages captioned using Gemini API. Some with gemini-2.0-flash-thinking-exp-1219, but most with gemini-2.0-flash-thinking-exp-01-21. If you plan on using this for Text-to-Image training I highly suggest doing more filtering for things such as resolution and compression artifacts. I kept those images to improve Image-to-Text performance, but it would be undesirable in Text-to-Image. imageimage-to-text10K<n<100K0 likes655 downloads2y agoHugging Face04PJMixers-Dev /proxy-logs-ShareGPTAll the proxy logs I could find (lmk if there are more), converted to ShareGPT so it's all formatted the same way in one dataset. I didn't .strip() any of the turns, I only used ftfy.fix_text() on them, and skipped any empty turns. The original sample response is moved to be the final turn of conversations. If the last turn in the original sample prompt was a model turn, it is assumed that it was for doing prefill. So I moved this turn to response_prefill. sample["response_prefill"] and… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/proxy-logs-ShareGPT.text1M<n<10M0 likes580 downloads9mo agoHugging Face05PJMixers-Dev /LimaRP-G4-26B-KCPPPJMixers-Dev/lemonilia_LimaRP-Simple-CustomShareGPT with responses generated using non-thinking bartowski/google_gemma-4-26B-A4B-it-GGUF/google_gemma-4-26B-A4B-it-Q4_K_M.gguf. Used default recommended generation settings: temp=1, top_k=64, top_p=0.95. 0 likes319 downloads5mo agoHugging Face06PJMixers-Dev /ThinkyThinky-ShareGPTdataset_list = [ "PJMixers-Dev/Weyaxi_HelpSteer-filtered-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT", "PJMixers-Dev/lmsys_lmsys-chat-1m-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT", "PJMixers-Dev/allenai_WildChat-1M-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT", "PJMixers-Dev/WizardLMTeam_WizardLM_evol_instruct_70k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT"… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/ThinkyThinky-ShareGPT.text100K<n<1M0 likes299 downloads2y agoHugging Face07PJMixers-Dev /Fundus-AP-Newstext10K<n<100K0 likes183 downloads1y agoHugging Face08PJMixers-Dev /Fundus-CC-2.5M Fundus-CC-2.5M 2.5 million news articles from CC-NEWS using Fundus. text1M<n<10M0 likes165 downloads2y agoHugging Face09PJMixers /Anthropic_persuasion-ShareGPTAdded the prompts which were listed in the blog post. Skipped the control samples. text1K<n<10K1 likes135 downloads2y agoHugging Face10Arimancy /pjm-demand-weather PJM Hourly Electricity Demand + Population-Weighted Weather (2015 to 2026) A model ready hourly time series for forecasting load on the largest US grid: PJM RTO demand (MWh) pre joined to population weighted weather, with calendar, holiday, degree hour, irradiance, cloud cover, apparent temperature, and snowfall features. 97,850 hours from 2015-07-01 to 2026-08-29, no missing hours on the UTC spine, every imputed, interpolated, or preliminary value explicitly flagged, in Parquet… See the full description on the dataset page: https://huggingface.co/datasets/Arimancy/pjm-demand-weather.tabular10K<n<100K2 likes132 downloads25d agoHugging Face11PJMixers-Dev /Fundus-AP-News-Unwantedtext10K<n<100K0 likes125 downloads1y agoHugging Face12PJMixers /0-hero_Matter-0.1-final_set_cleaned-ShareGPTtext1M<n<10M0 likes120 downloads2y agoHugging Face13pj-mathematician /clef2025-bioasq-task13Btext10M<n<100M1 likes102 downloads1y agoHugging Face14PJMixers-Dev /proxy-logs-ReRolls-Minos non-refusal responses: 845,186 refusal responses: 38,285 https://gist.github.com/xzuyn/1d7f43db2750060a18a304eb84b396db Use a training prompt formatter like this: https://github.com/xzuyn/axolotl/blob/latest-formatters/src/axolotl/prompt_strategies/customllama3-regex-last-only-prefill-reroll.py tabular100K<n<1M0 likes95 downloads9mo agoHugging Face15PJMixers-Dev /open-thoughts_OpenThoughts-114k-CustomShareGPTtext100K<n<1M0 likes92 downloads2y agoHugging Face16PJMixers-Dev /cognitivecomputations_dolphin-r1-reasoning-deepseek-CustomShareGPTtext100K<n<1M0 likes91 downloads2y agoHugging Face17PJMixers-Dev /Tiefighter-13B-Fake-Distill-ShareGPTtext1K<n<10K1 likes90 downloads2y agoHugging Face18PJMixers-Dev /Neph0s_CoSER-SFTtext100K<n<1M2 likes90 downloads1y agoHugging Face19PJMixers-Dev /teatime-logstext10K<n<100K0 likes90 downloads9mo agoHugging Face20PJMixers-Dev /cute-logs-Oct-14-2023https://char-archive.evulid.cc/api/file/download?path=/logs/cute%20logs%20-%20Oct%2014%202023.7z cute logs - Oct 14 2023 fiz123321-cloudy.hf.space text10K<n<100K0 likes88 downloads9mo agoHugging Face21PJMixers-Dev /foundRP-G4-26B-KCPPPJMixers-Dev/foundRP with responses generated using non-thinking bartowski/google_gemma-4-26B-A4B-it-GGUF/google_gemma-4-26B-A4B-it-Q4_K_M.gguf. Used default recommended generation settings: temp=1, top_k=64, top_p=0.95. text1K<n<10K0 likes88 downloads5mo agoHugging Face22PJMixers-Dev /bigdata-pw_Lyrics1M-entext100K<n<1M1 likes86 downloads2y agoHugging Face23PJMixers-Images /Castollux-Long0 likes85 downloads2y agoHugging Face24PJMixers-Dev /proxy-logs-ReRollsDuplicate prompts combined into a single sample, with all responses in a list of dicts. I've also included some info like token count, and slop (though my slop list could use improvement). Use a training prompt formatter like this: https://github.com/xzuyn/axolotl/blob/84aec029dfa9eb9670b8a51d432a279be6c85871/src/axolotl/prompt_strategies/customllama3-regex-last-only-prefill-reroll.py Dataset creation script: https://gist.github.com/xzuyn/aa1f30b7394d2997766bef82edb67227 tabular100K<n<1M1 likes83 downloads9mo agoHugging Face25PJMixers /airoboros-3.2-splittext10K<n<100K0 likes82 downloads3y agoHugging Face26PJMixers-Dev /Fundus-CC-2.5M-Formattedtext1M<n<10M1 likes80 downloads1y agoHugging Face27PJMixers-Dev /dolphin-deepseek-1k-think-1k-response-filtered-ShareGPTtabular10K<n<100K0 likes75 downloads2y agoHugging Face28PJMixers-Dev /HailMary-v0.2-KTO-Public Details This only contains the sets which are not private. This is also an experiment, so don't expect anything that good. The idea is to just take existing datasets which seem high quality and then generate a bad response for every model turn. If you have suggestions for improving this idea, I'm all ears. Refer to the original linked datasets for licenses as I add no further restrictions to them. Rejected Generations… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/HailMary-v0.2-KTO-Public.textreinforcement-learning100K<n<1M0 likes73 downloads2y agoHugging Face29PJMixers-Dev /microsoft_EpiCoder-func-380k-ShareGPT-splittext100K<n<1M1 likes69 downloads2y agoHugging Face30PJMixers-Dev /KodCode_KodCode-V1-SFT-R1-4k-think-1k-response-ShareGPTtabular100K<n<1M1 likes69 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.