datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Persuasive-Pairs
Persuasive Pairs
The dataset consists of pairs of short-text; one from a news,debate or chat (see field 'source' to see where the text originates from), one rewritten by LLM to contain more or less persuasive language.
The pairs are judged on degrees of persuasive language by three annotators: the task is to select which text contains much persuasive language and how much more on an ordinary scale with 'marginally','moderately', or 'heavily' more.
Flatten out the score is a 6-point… See the full description on the dataset page: https://huggingface.co/datasets/APauli/Persuasive-Pairs.polish-question-passage-pairsQuora-Question-Pairs
Quora Question Pairs — canonical 2017 release
A verbatim mirror of Quora's January 2017 Question Pairs release, packaged as a single tab-delimited file. No rows added, removed, or reordered relative to the upstream quora_duplicate_questions.tsv — only the hosting moved.
Re-hosted under Heliosoph for ingestion-pipeline stability — Quora's original CDN at qim.fs.quoracdn.net has been intermittently unreachable since the Kaggle competition wrapped, and the file has no checksumed… See the full description on the dataset page: https://huggingface.co/datasets/Heliosoph/Quora-Question-Pairs.viral_news_pairsThis dataset consists of popular news articles from google+, linkedin and facebook
In additin, a label column has been added to show the virality of the respectice title.
cichlid-behavior-pairs
Cichlid Behavior Pairs
Video/annotation pairs of Astatotilapia burtoni (cichlid fish) reproductive and
courtship behavior, recorded during a PGF2a-induced spawning assay comparing
nose-occluded ("VetBond") vs. sham-treated ("Sham") females — a manipulation
of olfactory input to the mating interaction. Each pair consists of one
top-down video of a male/female tank trial and one point-annotated behavior
event table (exported from BORIS) marking the
timing of specific… See the full description on the dataset page: https://huggingface.co/datasets/bds062/cichlid-behavior-pairs.Agriculture-Soil-QA-Pairs-Dataset
Agriculture-Soil-QA-Pairs-Dataset
Released freely — support more experiments like it:
More: https://huggingface.co/YuvrajSingh9886
protein-pairs-uniprot-swissprot
Protein Pairs and Similarity
Selected protein similarities within training, test, and validation sets.
Each protein gets two similarities selected at random and (usually) proteins within the top and bottom quintiles for similarity.
The protein is represented by its UniProt ID and its amino acid sequence (using IUPAC-IUB codes where each amino acid maps to a letter of the alphabet, see: https://en.wikipedia.org/wiki/FASTA_format ).
The distance column is cosine distance (identical =… See the full description on the dataset page: https://huggingface.co/datasets/monsoon-nlp/protein-pairs-uniprot-swissprot.amharic-oromo_sentence-pairs
Amharic-Oromo_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Amharic-Oromo_Sentence-Pairs
Number of Rows: 109805
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/amharic-oromo_sentence-pairs.Dutch-QA-Pairs-RijksoverheidThis dataset originates from the Open Government Web Portal provided by the Dutch Government. It contains Dutch question-answer pairs (vraag-antwoord combinaties or VAC), offering insights into various governmental inquiries and corresponding responses.
Columns: [instruction, input, output, text]
Nr. of entries: 1940
Agriculture-Plan-Diseases-QA-Pairs-Dataset
Agriculture-Plan-Diseases-QA-Pairs-Dataset
Released freely — support more experiments like it:
More: https://huggingface.co/YuvrajSingh9886
USDT_pairs_1H_3_month_data
Dataset Card for Dataset Name
<This dataset include 538 usdt pairs. time 1 hour. Duration 3 month. It's compatible for techical analysis.
Dataset Details
Dataset Description
Curated by: [kozannilker@gmail.com]
Funded by [optional]: [More Information Needed]
Shared by [optional]: [kozannilker@gmail.com]
Language(s) (NLP): [More Information Needed]
License: [mit]
Dataset Sources [optional]
Repository: [More Information Needed]
Paper… See the full description on the dataset page: https://huggingface.co/datasets/kozannilker/USDT_pairs_1H_3_month_data.English-Swahili-Sentence-Pairsidris_doc_pairs_datasetAgriculture-Irrigation-QA-Pairs-Dataset
Agriculture-Irrigation-QA-Pairs-Dataset
Released freely — support more experiments like it:
More: https://huggingface.co/YuvrajSingh9886
Philosophical-STS-Text-Pairs
Philosophical-STS-Text-Pairs
Gemma 3 Generated Synthetic Text Pairs for Embedding Pre-training
This project introduces SEP-STS-Text-Pairs, a synthetic dataset specifically designed for pre-training and fine-tuning embedding models on Semantic Textual Similarity (STS) tasks. The core of the effort involves a Python script that leverages the Gemma 3 12b generative AI model to create high-quality, diverse pairs of texts along with numerical similarity scores.
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/johnnyboycurtis/Philosophical-STS-Text-Pairs.SynPerForm-synthetic-persian-formality-pairs
SynPerForm
SynPerForm is a paired Persian dataset for formality style
transfer. Each informal text is paired with a freely written formal rewrite
that preserves its meaning without requiring lexical or structural
equivalence. The formal rewrites were generated with OpenAI GPT-5.6 Luna.
Columns
Informal: the original informal Persian text.
Formal: a free formal rewrite that preserves the original meaning.
Intended uses
Formality style transfer:… See the full description on the dataset page: https://huggingface.co/datasets/shekar-ai/SynPerForm-synthetic-persian-formality-pairs.query-hard-pos-neg-doc-pairs-statictablegenz-slang-pairs-1k
Gen Z Slang Pairs Corpus (1 K)
The Gen Z Slang Pairs Corpus (1 K) contains 1,000 everyday English sentences alongside their Gen Z–style slang rewrites. This dataset is designed for style-transfer, informal-language generation, and paraphrasing research. Use it to train models that transform formal or neutral sentences into expressive, youth‑oriented slang.
Dataset Details
This dataset was generated programmatically using OpenAI GPT-4.1 Nano.
Language: English… See the full description on the dataset page: https://huggingface.co/datasets/Programmer-RD-AI/genz-slang-pairs-1k.molecular-depiction-pairs-20k
Molecular depiction pairs, 20K
39,891 synthetic depictions of 19,990 drug-like PubChem molecules, each paired with the molecular identity it was rendered from, plus precomputed embeddings and fingerprints.
This is the development-scale dataset from molecular-depiction-alignment, published so the experiments in that repository can be reproduced without standing up the generation environment.
Why this exists
Not because a synthetic depiction corpus is novel. It is… See the full description on the dataset page: https://huggingface.co/datasets/hheiden/molecular-depiction-pairs-20k.fulah-hausa_sentence-pairs
Fulah-Hausa_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fulah-Hausa_Sentence-Pairs
Number of Rows: 269337
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fulah-hausa_sentence-pairs.ewe-twi_sentence-pairs-200k
This dataset is shared under CC BY-NC 4.0, which means you are free to use, share, and adapt it for non-commercial research and educational purposes with attribution. You can read the full license at https://creativecommons.org/licenses/by-nc/4.0/.
This dataset is made available because of Ghana NLP's volunteer driven research work. Please consider contributing to any of our projects on Github
Ewe-Twi_Sentence-Pairs Dataset
This dataset contains sentence pairs for African… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ewe-twi_sentence-pairs-200k.kamba-lingala_sentence-pairs
Kamba-Lingala_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Kamba-Lingala_Sentence-Pairs
Number of Rows: 50317
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/kamba-lingala_sentence-pairs.igbo-kongo_sentence-pairs
Igbo-Kongo_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Igbo-Kongo_Sentence-Pairs
Number of Rows: 68066
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/igbo-kongo_sentence-pairs.ipa-phonemes-word-pairs
license: cc-by-sa 4.0
size: ~275k pairs, ~7mb (~4mb parquet)
generated using: phonemizer/espeak
check out openphonemizer for more details!
bemba-xhosa_sentence-pairs
Bemba-Xhosa_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Bemba-Xhosa_Sentence-Pairs
Number of Rows: 213174
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/bemba-xhosa_sentence-pairs.akan-umbundu_sentence-pairs
Akan-Umbundu_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Akan-Umbundu_Sentence-Pairs
Number of Rows: 22651
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/akan-umbundu_sentence-pairs.afrikaans-dinka_sentence-pairs
Afrikaans-Dinka_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Afrikaans-Dinka_Sentence-Pairs
Number of Rows: 113795
Number… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/afrikaans-dinka_sentence-pairs.correct-incorrect-spelling-pairsThis is a dataset containing correct and incorrect spelling pairs in Gujarati, created by us using artificial noise.
fulah-rundi_sentence-pairs
Fulah-Rundi_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Fulah-Rundi_Sentence-Pairs
Number of Rows: 80036
Number of… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/fulah-rundi_sentence-pairs.dinka-twi_sentence-pairs
Dinka-Twi_Sentence-Pairs Dataset
This dataset contains sentence pairs for African languages along with similarity scores. It can be used for machine translation, sentence alignment, or other natural language processing tasks.
This dataset is based on the NLLBv1 dataset, published on OPUS under an open-source initiative led by META. You can find more information here: OPUS - NLLB-v1
Metadata
File Name: Dinka-Twi_Sentence-Pairs
Number of Rows: 34672
Number of Columns:… See the full description on the dataset page: https://huggingface.co/datasets/michsethowusu/dinka-twi_sentence-pairs.
