CoolFace
Datasetpublic

shuttie/reddit-dadjokes

Reddit /r/DadJokes dataset Dataset is based on a semi-public Pushshift Reddit dataset. Source data The actual extract of all top-level posts on /r/DadJokes can be found in the src/reddit_dadjokes.csv.gz file. The raw post contents are very noisy, so we did a couple of steps to clean things up: Used a Llama-8B-instruct to split each dad joke into the intro and a punchline parts. See src/parse.py script for details. Removed too short and too long posts. Also… See the full description on the dataset page: https://huggingface.co/datasets/shuttie/reddit-dadjokes.

sourceHugging Faceapache-2.0updated 2y agoView on Hugging Face
1likes85downloads
settings

This repository belongs to shuttie on Hugging Face.

CoolFace never edits a repository it does not host. Visibility, licence, collaborators and gating are all managed at the source.

namereddit-dadjokes
visibilitypublic
licenceapache-2.0
gatedno
ownershuttie
Account settings
shuttie/reddit-dadjokes · CoolFace