datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SynPerForm-synthetic-persian-formality-pairs
SynPerForm
SynPerForm is a paired Persian dataset for formality style
transfer. Each informal text is paired with a freely written formal rewrite
that preserves its meaning without requiring lexical or structural
equivalence. The formal rewrites were generated with OpenAI GPT-5.6 Luna.
Columns
Informal: the original informal Persian text.
Formal: a free formal rewrite that preserves the original meaning.
Intended uses
Formality style transfer:… See the full description on the dataset page: https://huggingface.co/datasets/shekar-ai/SynPerForm-synthetic-persian-formality-pairs.genz-slang-pairs-1k
Gen Z Slang Pairs Corpus (1 K)
The Gen Z Slang Pairs Corpus (1 K) contains 1,000 everyday English sentences alongside their Gen Z–style slang rewrites. This dataset is designed for style-transfer, informal-language generation, and paraphrasing research. Use it to train models that transform formal or neutral sentences into expressive, youth‑oriented slang.
Dataset Details
This dataset was generated programmatically using OpenAI GPT-4.1 Nano.
Language: English… See the full description on the dataset page: https://huggingface.co/datasets/Programmer-RD-AI/genz-slang-pairs-1k.genz-slang-pairs-1k
Gen Z Slang Pairs Corpus (1 K)
The Gen Z Slang Pairs Corpus (1 K) contains 1,000 everyday English sentences alongside their Gen Z–style slang rewrites. This dataset is designed for style-transfer, informal-language generation, and paraphrasing research. Use it to train models that transform formal or neutral sentences into expressive, youth‑oriented slang.
Dataset Details
This dataset was generated programmatically using OpenAI GPT-4.1 Nano.
Language:… See the full description on the dataset page: https://huggingface.co/datasets/JScharp/genz-slang-pairs-1k.film_qa_pairs_datasetgenz-slang-pairs-1k
Gen Z Slang Pairs Corpus (1 K)
The Gen Z Slang Pairs Corpus (1 K) contains 1,000 everyday English sentences alongside their Gen Z–style slang rewrites. This dataset is designed for style-transfer, informal-language generation, and paraphrasing research. Use it to train models that transform formal or neutral sentences into expressive, youth‑oriented slang.
Dataset Details
This dataset was generated programmatically using OpenAI GPT-4.1 Nano.
Language: English… See the full description on the dataset page: https://huggingface.co/datasets/kawaii4mano/genz-slang-pairs-1k.Film_QA_Pairsenglish_karakalpak_pairs_parallel_corpus_v2_8907
English-Karakalpak Parallel Corpus v2 (8.9K)
Dataset Description
English-Karakalpak Parallel Corpus v2 is a high-quality dataset containing 8,906 carefully aligned sentence pairs in English (en) and Karakalpak (kaa).
This dataset is designed to advance the representation and capability of the Karakalpak language in large-scale AI models (LLMs) and Neural Machine Translation (NMT) systems, enabling them to better understand and generate Karakalpak text. This resource… See the full description on the dataset page: https://huggingface.co/datasets/bekan/english_karakalpak_pairs_parallel_corpus_v2_8907.
