datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
r-uncleaned
r (Uncleaned)
Contains low quality images, watermarks, etc. Nothing has been removed except stuff like "imgur empty images".
File name structure: {subreddit} - {date:%Y-%m-%d_%H-%M-%S} - {author} - {id}-{num}.{extension}
Note: Was downloaded with this filter: --filter "extension in ('jpg', 'jpeg', 'webp', 'png') and is_reddit_media_domain is True"
ChaiCastollux-Long-ParquetImages captioned using Gemini API. Some with gemini-2.0-flash-thinking-exp-1219, but most with gemini-2.0-flash-thinking-exp-01-21.
If you plan on using this for Text-to-Image training I highly suggest doing more filtering for things such as resolution and compression artifacts. I kept those images to improve Image-to-Text performance, but it would be undesirable in Text-to-Image.
proxy-logs-ShareGPTAll the proxy logs I could find (lmk if there are more), converted to ShareGPT so it's all formatted the same way in one dataset. I didn't .strip() any of the turns, I only used ftfy.fix_text() on them, and skipped any empty turns.
The original sample response is moved to be the final turn of conversations. If the last turn in the original sample prompt was a model turn, it is assumed that it was for doing prefill. So I moved this turn to response_prefill.
sample["response_prefill"] and… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/proxy-logs-ShareGPT.LimaRP-G4-26B-KCPPPJMixers-Dev/lemonilia_LimaRP-Simple-CustomShareGPT with responses generated using non-thinking bartowski/google_gemma-4-26B-A4B-it-GGUF/google_gemma-4-26B-A4B-it-Q4_K_M.gguf.
Used default recommended generation settings: temp=1, top_k=64, top_p=0.95.
ThinkyThinky-ShareGPTdataset_list = [
"PJMixers-Dev/Weyaxi_HelpSteer-filtered-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT",
"PJMixers-Dev/lmsys_lmsys-chat-1m-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT",
"PJMixers-Dev/allenai_WildChat-1M-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT",
"PJMixers-Dev/WizardLMTeam_WizardLM_evol_instruct_70k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT"… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/ThinkyThinky-ShareGPT.Fundus-AP-NewsFundus-CC-2.5M
Fundus-CC-2.5M
2.5 million news articles from CC-NEWS using Fundus.
Anthropic_persuasion-ShareGPTAdded the prompts which were listed in the blog post. Skipped the control samples.
pjm-demand-weather
PJM Hourly Electricity Demand + Population-Weighted Weather (2015 to 2026)
A model ready hourly time series for forecasting load on the largest US
grid: PJM RTO demand (MWh) pre joined to population weighted weather,
with calendar, holiday, degree hour, irradiance, cloud cover, apparent
temperature, and snowfall features. 97,850 hours from 2015-07-01 to
2026-08-29, no missing hours on the UTC spine, every imputed,
interpolated, or preliminary value explicitly flagged, in Parquet… See the full description on the dataset page: https://huggingface.co/datasets/Arimancy/pjm-demand-weather.Fundus-AP-News-Unwanted0-hero_Matter-0.1-final_set_cleaned-ShareGPTclef2025-bioasq-task13Bproxy-logs-ReRolls-Minos
non-refusal responses: 845,186
refusal responses: 38,285
https://gist.github.com/xzuyn/1d7f43db2750060a18a304eb84b396db
Use a training prompt formatter like this: https://github.com/xzuyn/axolotl/blob/latest-formatters/src/axolotl/prompt_strategies/customllama3-regex-last-only-prefill-reroll.py
open-thoughts_OpenThoughts-114k-CustomShareGPTcognitivecomputations_dolphin-r1-reasoning-deepseek-CustomShareGPTTiefighter-13B-Fake-Distill-ShareGPTNeph0s_CoSER-SFTteatime-logscute-logs-Oct-14-2023https://char-archive.evulid.cc/api/file/download?path=/logs/cute%20logs%20-%20Oct%2014%202023.7z
cute logs - Oct 14 2023
fiz123321-cloudy.hf.space
foundRP-G4-26B-KCPPPJMixers-Dev/foundRP with responses generated using non-thinking bartowski/google_gemma-4-26B-A4B-it-GGUF/google_gemma-4-26B-A4B-it-Q4_K_M.gguf.
Used default recommended generation settings: temp=1, top_k=64, top_p=0.95.
bigdata-pw_Lyrics1M-enCastollux-Longproxy-logs-ReRollsDuplicate prompts combined into a single sample, with all responses in a list of dicts. I've also included some info like token count, and slop (though my slop list could use improvement).
Use a training prompt formatter like this: https://github.com/xzuyn/axolotl/blob/84aec029dfa9eb9670b8a51d432a279be6c85871/src/axolotl/prompt_strategies/customllama3-regex-last-only-prefill-reroll.py
Dataset creation script: https://gist.github.com/xzuyn/aa1f30b7394d2997766bef82edb67227
airoboros-3.2-splitFundus-CC-2.5M-Formatteddolphin-deepseek-1k-think-1k-response-filtered-ShareGPTHailMary-v0.2-KTO-Public
Details
This only contains the sets which are not private. This is also an experiment, so don't expect anything that good.
The idea is to just take existing datasets which seem high quality and then generate a bad response for every model turn. If you have suggestions for improving this idea, I'm all ears.
Refer to the original linked datasets for licenses as I add no further restrictions to them.
Rejected Generations… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/HailMary-v0.2-KTO-Public.microsoft_EpiCoder-func-380k-ShareGPT-splitKodCode_KodCode-V1-SFT-R1-4k-think-1k-response-ShareGPT
