datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
genz-slang-dataset
Dataset Details
This dataset contains a rich collection of popular slang terms and acronyms used primarily by Generation Z. It includes detailed descriptions of each term, its context of use, and practical examples that demonstrate how the slang is used in real-life conversations.
The dataset is designed to capture the unique and evolving language patterns of GenZ, reflecting their communication style in digital spaces such as social media, text messaging, and online forums. Each… See the full description on the dataset page: https://huggingface.co/datasets/MLBtrio/genz-slang-dataset.genz-to-english
GenZ-to-English Translation Dataset
A high-quality text-to-text dataset for translating Gen Z slang into clear, standard English.
The dataset is designed for training and evaluating language models that convert modern internet slang into natural, readable English while preserving the original meaning.
Overview
This dataset contains 300k++ curated translation pairs covering a wide range of contemporary internet slang.
It includes expressions commonly found across… See the full description on the dataset page: https://huggingface.co/datasets/Sankar-2910/genz-to-english.NeuroBio-GenZ-1K
NeuroBio GenZ 1K
Around 1000 neuroscience and biology questions, answered like your smartest friend is texting you back, not like a textbook is talking at you.
"Why does doomscrolling give me dopamine?" gets answered in three sentences, casual tone, real neuroscience terms (nucleus accumbens, not "reward center"), zero fluff.
Why did you make this?
Because there's genuinely not that much high quality neuroscience and biology data on Hugging Face that isn't either… See the full description on the dataset page: https://huggingface.co/datasets/luka0x12/NeuroBio-GenZ-1K.details_budecosystem__genz-13b-v2
Dataset Card for Evaluation run of budecosystem/genz-13b-v2
Dataset Summary
Dataset automatically created during the evaluation run of model budecosystem/genz-13b-v2 on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_budecosystem__genz-13b-v2.details_budecosystem__genz-70b
Dataset Card for Evaluation run of budecosystem/genz-70b
Dataset Summary
Dataset automatically created during the evaluation run of model budecosystem/genz-70b on the Open LLM Leaderboard.
The dataset is composed of 64 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_budecosystem__genz-70b.genz-slang-pairs-1k
Gen Z Slang Pairs Corpus (1 K)
The Gen Z Slang Pairs Corpus (1 K) contains 1,000 everyday English sentences alongside their Gen Z–style slang rewrites. This dataset is designed for style-transfer, informal-language generation, and paraphrasing research. Use it to train models that transform formal or neutral sentences into expressive, youth‑oriented slang.
Dataset Details
This dataset was generated programmatically using OpenAI GPT-4.1 Nano.
Language: English… See the full description on the dataset page: https://huggingface.co/datasets/Programmer-RD-AI/genz-slang-pairs-1k.genz-slang-pairs-1k
Gen Z Slang Pairs Corpus (1 K)
The Gen Z Slang Pairs Corpus (1 K) contains 1,000 everyday English sentences alongside their Gen Z–style slang rewrites. This dataset is designed for style-transfer, informal-language generation, and paraphrasing research. Use it to train models that transform formal or neutral sentences into expressive, youth‑oriented slang.
Dataset Details
This dataset was generated programmatically using OpenAI GPT-4.1 Nano.
Language:… See the full description on the dataset page: https://huggingface.co/datasets/JScharp/genz-slang-pairs-1k.Genz-Translationgenz-slang-dataset
Dataset Details
This dataset contains a rich collection of popular slang terms and acronyms used primarily by Generation Z. It includes detailed descriptions of each term, its context of use, and practical examples that demonstrate how the slang is used in real-life conversations.
The dataset is designed to capture the unique and evolving language patterns of GenZ, reflecting their communication style in digital spaces such as social media, text messaging, and online forums.… See the full description on the dataset page: https://huggingface.co/datasets/hihry/genz-slang-dataset.details_TheBloke__Genz-70b-GPTQ
Dataset Card for Evaluation run of TheBloke/Genz-70b-GPTQ
Dataset Summary
Dataset automatically created during the evaluation run of model TheBloke/Genz-70b-GPTQ on the Open LLM Leaderboard.
The dataset is composed of 61 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_TheBloke__Genz-70b-GPTQ.genz_preference
GenZ-Slang-ORPO Dataset v1.0
This dataset is designed for ORPO and DPO training, but with an additional style-control layer.
The system prompt encodes four knobs:
• <SLANG:0–3>
• <EMOJI:0–3>
• <HYPE:0–3>
• <FORMAL:0–3>
Each knob indicates intensity, where 0 = none and 3 = heavy usage.
📑 Schema
Each record has:
• prompt: serialized conversation history (system + user)
• chosen: preferred model reply
• rejected: less-preferred reply
• messages: full structured… See the full description on the dataset page: https://huggingface.co/datasets/biropost/genz_preference.Personal-Generate-GenZ-Students-Khmer
Personas Generate Genz Students Khmer
Welcome to the SeyhaLite collection. This dataset has been curated and cleaned to support the development of high-quality Khmer Language Models (LLMs) specifically focused on persona generation and synthetic character creation for the Gen Z demographic in Cambodia.
Project Vision
I hope this dataset helps your project succeed! Whether you are building a specialized chatbot, an AI assistant, or conducting social research, this data is… See the full description on the dataset page: https://huggingface.co/datasets/SeyhaLite/Personal-Generate-GenZ-Students-Khmer.genz-contextual-abusive-slang-v2-silver
Gen Z Contextual Abusive Slang Benchmark v2 Silver
This is a research-only silver candidate dataset for studying Gen Z / internet-slang abusive-language classification. It is designed to test whether models can distinguish slang from abuse, profanity from harassment, meme mockery from benign meme use, and identity mentions from identity attacks.
This is not a final gold benchmark. Labels are model-assisted silver labels generated with a three-pass LLM annotation pipeline and… See the full description on the dataset page: https://huggingface.co/datasets/AliceYin/genz-contextual-abusive-slang-v2-silver.gen_z_translationgenz-slang-instruction-datasetgenz-slang-dataset
Dataset Details
This dataset contains a rich collection of popular slang terms and acronyms used primarily by Generation Z. It includes detailed descriptions of each term, its context of use, and practical examples that demonstrate how the slang is used in real-life conversations.
The dataset is designed to capture the unique and evolving language patterns of GenZ, reflecting their communication style in digital spaces such as social media, text messaging, and online forums. Each… See the full description on the dataset page: https://huggingface.co/datasets/MariyaAnjum/genz-slang-dataset.genz-slang-completions
Gen Z Slang Chat Completions
This data is based on the genz-slang-dataset, with gpt-4o-mini generated questions for each response.
genz-slang-dataset
Dataset Details
This dataset contains a rich collection of popular slang terms and acronyms used primarily by Generation Z. It includes detailed descriptions of each term, its context of use, and practical examples that demonstrate how the slang is used in real-life conversations.
The dataset is designed to capture the unique and evolving language patterns of GenZ, reflecting their communication style in digital spaces such as social media, text messaging, and online forums. Each… See the full description on the dataset page: https://huggingface.co/datasets/averz/genz-slang-dataset.genz-sft-dataset
Dataset Card for Gen Z SFT Dataset
Synthetic chat data for supervised fine-tuning of models that should reply in casual internet / Gen Z slang.
Each row is a 3-turn conversation: a fixed system prompt, an everyday user message, and an in-character assistant reply.
Dataset Details
Dataset Description
This dataset is for style / persona SFT, not facts. User turns are short everyday chats (school, work, dating, food, roommates, social media, small… See the full description on the dataset page: https://huggingface.co/datasets/grenishrai/genz-sft-dataset.genz-alpha-slangsvietnamese-genz-datasetGenzTranscribe-higen-z-translationgen-z-slangs-translationGenZ-SlangGenZ_dataformatted_genz_normal_enggen-z-datasetgenz_harmonygenz
