CoolFace
7 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hpprc /paraphrase-qa日本語Wikipedia中のテキストを元に言い換えを生成し、その言い換えを元にクエリと回答をLLMに生成させたデータセットです。 出力にライセンス的な制約があるモデルを利用していないことと、元データとして日本語Wikipediaを利用していることから、CC-BY-SA 4.0ライセンスのもとでの配布とします。 textquestion-answering10M<n<100M2 likes184 downloads2y agoHugging Face02DomainLLM /gerlayqa-bgb-paraphrased GerLayQA-BGB Paraphrased 🇩🇪⚖️ Dataset Description This is a paraphrased and restructured version of the GerLayQA BGB (Bürgerliches Gesetzbuch / German Civil Code) dataset, specifically prepared for fine-tuning large language models on German civil law question-answering tasks. Key Features 5,255 high-quality QA pairs about German Civil Law (BGB) Paraphrased questions to remove plagiarism while maintaining legal accuracy Structured 7-section answers following… See the full description on the dataset page: https://huggingface.co/datasets/DomainLLM/gerlayqa-bgb-paraphrased.textquestion-answering1K<n<10K0 likes32 downloads1y agoHugging Face03DomainLLM /gerlayqa-combined-paraphrased GerLayQA Combined Paraphrased 🇩🇪⚖️ Dataset Description This is a combined, shuffled dataset merging both the BGB (civil law) and StGB (criminal law) paraphrased German legal QA datasets. All examples are paraphrased and restructured by GPT-5 for fine-tuning large language models on German legal question-answering tasks. Key Features 6,462 high-quality QA pairs covering both German Civil and Criminal Law Combined coverage: BGB (Bürgerliches Gesetzbuch) + StGB… See the full description on the dataset page: https://huggingface.co/datasets/DomainLLM/gerlayqa-combined-paraphrased.textquestion-answering1K<n<10K0 likes30 downloads1y agoHugging Face04DomainLLM /gerlayqa-stgb-paraphrased GerLayQA-StGB Paraphrased 🇩🇪⚖️ Dataset Description This is a paraphrased and restructured version of the GerLayQA StGB (Strafgesetzbuch / German Criminal Code) dataset, specifically prepared for fine-tuning large language models on German criminal law question-answering tasks. Key Features 1,207 high-quality QA pairs about German Criminal Law (StGB) Paraphrased questions to remove plagiarism while maintaining legal accuracy Structured 7-section answers… See the full description on the dataset page: https://huggingface.co/datasets/DomainLLM/gerlayqa-stgb-paraphrased.textquestion-answering1K<n<10K0 likes25 downloads1y agoHugging Face05onnookk /gsm8k-paraphrase-deepseek GSM8K paraphrase variants (DeepSeek-v4-flash) Paraphrase ablations over the GSM8K-test split, formatted to match the openai/gsm8k schema: a question string and an answer string ending with #### N. The first three configs cover the same 660-item subsample used to contaminate the m2_mixed model (random.Random(42).sample(range(1319), 660)). The fourth config, pq_complement, covers the remaining 659 held-out items (the complement of that subsample), so pq ∪ pq_complement… See the full description on the dataset page: https://huggingface.co/datasets/onnookk/gsm8k-paraphrase-deepseek.texttext-generation1K<n<10K0 likes22 downloads4mo agoHugging Face06prithivMLmods /GPT-Paraphrases GPT-Paraphrases dataset This dataset contains text passages and their paraphrases generated using the GPT-3 language model. The paraphrases are designed to be semantically equivalent to the original text, but with different wording and structure. The dataset includes text formatted in JSON and is in English. Dataset Statistics Number of text passages: Not specified in the information you provided. Source of text passages: Not specified in the information you provided.… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/GPT-Paraphrases.texttext-generation100K<n<1M1 likes20 downloads2y agoHugging Face07nlpctx /legal-faq-paraphrased Legal FAQ Query Reformulation Benchmark Dataset Summary This dataset is derived from the Legal-FAQ dataset and is designed for evaluating the robustness of retrieval and embedding models under query reformulation. Each original FAQ question has been expanded into multiple query styles: Formal Casual Search Hard Reformulation The objective is to measure how well retrieval systems can retrieve the correct answer despite substantial changes in wording.… See the full description on the dataset page: https://huggingface.co/datasets/nlpctx/legal-faq-paraphrased.texttext-retrieval1K<n<10K0 likes12 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.