CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01justintiensmith /Ordering_Constrained_ParaphrasesThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/Ordering_Constrained_Paraphrases.tabularrobotics100K<n<1M0 likes404 downloads3mo agoHugging Face02justintiensmith /Reorient_Block_ParaphrasesThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 30, "features": { "action": { "dtype": "float32", "names": [ "shoulder_pan.pos", "shoulder_lift.pos", "elbow_flex.pos", "wrist_flex.pos", "wrist_roll.pos", "gripper.pos" ], "shape": [ 6… See the full description on the dataset page: https://huggingface.co/datasets/justintiensmith/Reorient_Block_Paraphrases.tabularrobotics10K<n<100K0 likes260 downloads3mo agoHugging Face03Faless /harvest_apples_with_agilex_piper_sim_ee_paraphrases20This dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v3.0", "fps": 25, "features": { "observation.state": { "dtype": "float32", "fps": 25, "shape": [ 8 ], "names": [ "ee.x", "ee.y", "ee.z", "ee.roll", "ee.pitch", "ee.yaw"… See the full description on the dataset page: https://huggingface.co/datasets/Faless/harvest_apples_with_agilex_piper_sim_ee_paraphrases20.tabularrobotics1M<n<10M0 likes200 downloads4mo agoHugging Face04jpwahle /autoencoder-paraphrase-dataset Dataset Card for Machine Paraphrase Dataset (MPC) Dataset Summary The Autoencoder Paraphrase Corpus (APC) consists of ~200k examples of original, and paraphrases using three neural language models. It uses three models (BERT, RoBERTa, Longformer) on three source texts (Wikipedia, arXiv, student theses). The examples are aligned, i.e., we sample the same paragraphs for originals and paraphrased versions. How to use it You can load the dataset using the… See the full description on the dataset page: https://huggingface.co/datasets/jpwahle/autoencoder-paraphrase-dataset.tabulartext-classification1M<n<10M2 likes141 downloads1y agoHugging Face05cestwc /paraphrasetabular1M<n<10M5 likes106 downloads4y agoHugging Face06deutsche-telekom /ger-backtrans-paraphrase German Backtranslated Paraphrase Dataset This is a dataset of more than 21 million German paraphrases. These are text pairs that have the same meaning but are expressed with different words. The source of the paraphrases are different parallel German / English text corpora. The English texts were machine translated back into German to obtain the paraphrases. This dataset can be used for example to train semantic text embeddings. To do this, for example, SentenceTransformers and the… See the full description on the dataset page: https://huggingface.co/datasets/deutsche-telekom/ger-backtrans-paraphrase.tabularsentence-similarity10M<n<100M12 likes77 downloads2y agoHugging Face07bishalagrawal /medical_questions_paraphrasestabular1K<n<10K0 likes43 downloads2y agoHugging Face08imperialwarrior /open-australian-legal-qa-paraphrased-easy-geminitabularn<1K0 likes41 downloads3y agoHugging Face09open-llm-leaderboard /cluebbers__Llama-3.1-8B-paraphrase-type-generation-apty-sigmoid-detailsgated Dataset Card for Evaluation run of cluebbers/Llama-3.1-8B-paraphrase-type-generation-apty-sigmoid Dataset automatically created during the evaluation run of model cluebbers/Llama-3.1-8B-paraphrase-type-generation-apty-sigmoid The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/cluebbers__Llama-3.1-8B-paraphrase-type-generation-apty-sigmoid-details.tabular10K<n<100K0 likes41 downloads2y agoHugging Face10imperialwarrior /open-australian-legal-qa-paraphrased-hard-gemini-with-embtabularn<1K0 likes40 downloads3y agoHugging Face11imperialwarrior /open-australian-legal-qa-paraphrased-easy-gemini-with-embtabularn<1K0 likes39 downloads3y agoHugging Face12imperialwarrior /open-australian-legal-qa-paraphrased-hard-geminitabularn<1K0 likes37 downloads3y agoHugging Face13faisal4590aziz /bangla-health-related-paraphrased-dataset Dataset Card for "BanglaHealthParaphrase" BanglaHealthParaphrase is a Bengali paraphrasing dataset specifically curated for the health domain. It contains over 200,000 sentence pairs, where each pair consists of an original Bengali sentence and its paraphrased version. The dataset was created through a multi-step pipeline involving extraction of health-related content from Bengali news sources, English pivot-based paraphrasing, and back-translation to ensure linguistic diversity… See the full description on the dataset page: https://huggingface.co/datasets/faisal4590aziz/bangla-health-related-paraphrased-dataset.tabulartext-generation100K<n<1M2 likes33 downloads1y agoHugging Face14arjun10g /slop-paraphrase-pairs-v2tabular1K<n<10K0 likes28 downloads4mo agoHugging Face15skeskinen /books3_basic_sentenses_paraphrased Dataset Card for "books3_basic_sentenses_paraphrased" More Information needed tabular100K<n<1M1 likes24 downloads3y agoHugging Face16fyaronskiy /ru-paraphrase-NMT-Leipzig-cleaned Dataset Description The dataset is obtained by filtering dataset of russian paraphrases by David Dale with automatic metrics. The data structure is saved. Have been deleted: Paraphrases that have cosine LABSE similarity with source sentences < 0.75. Paraphrases that are more than 2.5 times longer than source sentences. (Most of them are looped errors of back translation) Paraphrases that are similar in spelling to the original texts (paraphrases that have ChrF++ similarity > 0.6… See the full description on the dataset page: https://huggingface.co/datasets/fyaronskiy/ru-paraphrase-NMT-Leipzig-cleaned.tabulartext-generation100K<n<1M2 likes23 downloads1y agoHugging Face17impresso-project /amr-true-paraphrases True Paraphrases Test Set The True Paraphrases sentence/phrase pairs derived from the AMR Annotation Guidelines. It was introduced as part of the PARAPHRASUS: A Comprehensive Benchmark for Evaluating Paraphrase Detection Models. For more details, refer to the original paper that was presented at COLING 2025. Citation If you use this dataset, please cite it using the following BibTeX entry: @inproceedings{michail-etal-2025-paraphrasus, title = "{PARAPHRASUS}: A… See the full description on the dataset page: https://huggingface.co/datasets/impresso-project/amr-true-paraphrases.tabulartext-classificationn<1K0 likes23 downloads2y agoHugging Face18matboz /alpaca_llama3.18b_em_paraphrased_divergence_ratiotabular10K<n<100K0 likes23 downloads10mo agoHugging Face19kenken6696 /FOLIO_by_paraphrased_gpt4tabular1K<n<10K0 likes22 downloads3y agoHugging Face20memyprokotow /popqa_full_w_paraphrasestabular10K<n<100K0 likes22 downloads7mo agoHugging Face21impresso-project /sts-h-paraphrase-detection STS-Hard Test Set The STS-Hard dataset is a paraphrase detection test set derived from the STSBenchmark dataset. It was introduced as part of the PARAPHRASUS: A Comprehensive Benchmark for Evaluating Paraphrase Detection Models. The test set includes the paraphrase label as well as individual annotation labels from two annotators: P1: The semanticist. P2: A student annotator. For more details, refer to the original paper that was presented at COLING 2025. Citation… See the full description on the dataset page: https://huggingface.co/datasets/impresso-project/sts-h-paraphrase-detection.tabulartext-classificationn<1K0 likes21 downloads2y agoHugging Face22lasha-nlp /bookmia_paraphrase_claudetabular1K<n<10K0 likes21 downloads2y agoHugging Face23Myashka /SO-Python_basics_QA-filtered-2023-T5_paraphrased-tanh_scoretabular100K<n<1M0 likes20 downloads3y agoHugging Face24Lidor-Mashiach /snli-training-paraphrase-augmentation SNLI Training Paraphrase Augmentation Purpose This dataset contains new paraphrases created for training paraphrase augmentation during Phase B of the research project. It was not used as an NLI evaluation set. It was not used for paraphrase consistency evaluation. It is separate from the published SNLI Paraphrase Bank used for evaluation. The generation records report zero collisions with that evaluation bank. The CSV contains only the new augmentation rows. It… See the full description on the dataset page: https://huggingface.co/datasets/Lidor-Mashiach/snli-training-paraphrase-augmentation.tabular1M<n<10M1 likes20 downloads2mo agoHugging Face25hf-future-backdoors /OpenHermes-paraphrased-headlines-2017-2019-eval-settabular1K<n<10K0 likes19 downloads2y agoHugging Face26nora-team /paul-paraphrased-questiontabularn<1K0 likes19 downloads1y agoHugging Face27Lidor-Mashiach /mnli-training-paraphrase-augmentation MNLI Training Paraphrase Augmentation Purpose This dataset contains new paraphrases created for training paraphrase augmentation during Phase B of the research project. It was not used as an NLI evaluation set. It was not used for paraphrase consistency evaluation. It is separate from the published MNLI Paraphrase Bank used for evaluation. The generation records report zero collisions with that evaluation bank. The CSV contains only the new augmentation rows. It… See the full description on the dataset page: https://huggingface.co/datasets/Lidor-Mashiach/mnli-training-paraphrase-augmentation.tabular1M<n<10M2 likes19 downloads2mo agoHugging Face28kenken6696 /FOLIO_by_paraphrased_gpt3.5tabular1K<n<10K0 likes18 downloads3y agoHugging Face29axay /test_paraphrasetabularn<1K0 likes18 downloads2y agoHugging Face30math-extraction-comp /cluebbers__Llama-3.1-8B-paraphrase-type-generation-apty-sigmoidtabular1K<n<10K0 likes18 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.