pjm
Datasets
All datasets matching “pjm”r-uncleaned
r (Uncleaned)
Contains low quality images, watermarks, etc. Nothing has been removed except stuff like "imgur empty images".
File name structure: {subreddit} - {date:%Y-%m-%d_%H-%M-%S} - {author} - {id}-{num}.{extension}
Note: Was downloaded with this filter: --filter "extension in ('jpg', 'jpeg', 'webp', 'png') and is_reddit_media_domain is True"
ChaiCastollux-Long-ParquetImages captioned using Gemini API. Some with gemini-2.0-flash-thinking-exp-1219, but most with gemini-2.0-flash-thinking-exp-01-21.
If you plan on using this for Text-to-Image training I highly suggest doing more filtering for things such as resolution and compression artifacts. I kept those images to improve Image-to-Text performance, but it would be undesirable in Text-to-Image.
proxy-logs-ShareGPTAll the proxy logs I could find (lmk if there are more), converted to ShareGPT so it's all formatted the same way in one dataset. I didn't .strip() any of the turns, I only used ftfy.fix_text() on them, and skipped any empty turns.
The original sample response is moved to be the final turn of conversations. If the last turn in the original sample prompt was a model turn, it is assumed that it was for doing prefill. So I moved this turn to response_prefill.
sample["response_prefill"] and… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/proxy-logs-ShareGPT.LimaRP-G4-26B-KCPPPJMixers-Dev/lemonilia_LimaRP-Simple-CustomShareGPT with responses generated using non-thinking bartowski/google_gemma-4-26B-A4B-it-GGUF/google_gemma-4-26B-A4B-it-Q4_K_M.gguf.
Used default recommended generation settings: temp=1, top_k=64, top_p=0.95.
ThinkyThinky-ShareGPTdataset_list = [
"PJMixers-Dev/Weyaxi_HelpSteer-filtered-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT",
"PJMixers-Dev/lmsys_lmsys-chat-1m-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT",
"PJMixers-Dev/allenai_WildChat-1M-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT",
"PJMixers-Dev/WizardLMTeam_WizardLM_evol_instruct_70k-gemini-2.0-flash-thinking-exp-1219-CustomShareGPT"… See the full description on the dataset page: https://huggingface.co/datasets/PJMixers-Dev/ThinkyThinky-ShareGPT.
