datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
genz-slang-pairs-1k
Gen Z Slang Pairs Corpus (1 K)
The Gen Z Slang Pairs Corpus (1 K) contains 1,000 everyday English sentences alongside their Gen Z–style slang rewrites. This dataset is designed for style-transfer, informal-language generation, and paraphrasing research. Use it to train models that transform formal or neutral sentences into expressive, youth‑oriented slang.
Dataset Details
This dataset was generated programmatically using OpenAI GPT-4.1 Nano.
Language:… See the full description on the dataset page: https://huggingface.co/datasets/JScharp/genz-slang-pairs-1k.herman-json-mode
Herman: Indonesian Single-Turn JSON Mode
Herman is an Indonesian language dataset specifically designed
for training LLMs using a single-turn JSON mode. This dataset
is used in Supervised Fine-Tuning (SFT) to improve JSON parsing
capabilities in LLMs. Herman was obtained from Hermes and translated
into Indonesian for the purpose of training Indonesian language models.
Code used for constructing Herman can be found here.
Schema Format
The desired JSON schema can… See the full description on the dataset page: https://huggingface.co/datasets/SulthanAbiyyu/herman-json-mode.JS-Code-Solutions
Python Code Solutions
Features
1000k of JS Code Solutions for Text Generation and Question Answering
JS Coding Problems labelled by topic and difficulty
Recommendations
Train your Model on Logical Operations and Mathematical Problems Before Training it on this. This is optional for Fine Tuning 2B parameter + models.
Format the prompts in a orderly way when formatting data eg. {question} Solution: {solution} Topic: {topic}
Fraud_Case_Verdicts
The "Crime Facts" of "Offenses of Fraudulence" in Judicial Yuan Verdicts Dataset
This data set is based on the judgments of "Offenses of Fraudulence" cases published by the Judicial Yuan. The data range of the dataset is from January 1, 2011, to December 31, 2021. 74,823 pieces of original data (judgments and rulings) were collected. We only took the contents of the "criminal facts" field of the judgment. This dataset is divided into three parts. The training dataset has 59,858… See the full description on the dataset page: https://huggingface.co/datasets/jslin09/Fraud_Case_Verdicts.
