datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
puzzlescript-gists
PuzzleScript Human-Authored Games (Full Gist Corpus)
35,705 human-authored PuzzleScript games — the
complete source text of each — collected from public GitHub gists.
This is the full corpus: every distinct gist is kept, and each row is tagged
with its deduplication cluster so you can reduce to a unique set with a one-line
filter. The deduplication is reproducible from the shipped dedup_master.json +
dedup_master.py; non-vanilla PuzzleScript-Plus files are excluded (listed in… See the full description on the dataset page: https://huggingface.co/datasets/smearle/puzzlescript-gists.StereoTales
Multilingual Story-Generation Bias Samples
A multilingual evaluation dataset for probing demographic biases in LLM
story generation. Each sample instructs a model to write a ~200-word story
about a character carrying a given demographic attribute value (age, gender,
ethnicity, religion, disability status, immigration status, ...) placed into a
specific life scenario, with the goal of surfacing socio-economic and
demographic biases in the generated narratives.… See the full description on the dataset page: https://huggingface.co/datasets/giskardai/StereoTales.neo_ara_v2neo_ara_v1
