datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Japanese-RAG-Generator-Benchmark
Japanese RAG Generator Benchmark: 日本語 RAG における Generator 評価ベンチマーク
Japanese RAG Generator Benchmark (J-RAGBench) は日本語RAGにおけるGeneratorに用いるLLMの評価データセットを提供する。
実運用時のRAGに求められる多様な評価カテゴリを同一条件下で評価可能であり、複数の評価カテゴリが同時に出現する問題が含まれるQAデータセットを人手および、補助的にOpenAI API(gpt-4.1-2025-04-14)を用いて構築した。
J-RAGBenchの評価カテゴリ
Integration: 2~3文書程度の複数の情報源から適切な根拠を抽出・統合して回答を導く
Reasoning: 抽出された情報を踏まえて多段階の推論や数値計算などを実行する
Logical: 質問・関連文書間での語彙や表現の差異を解釈し、適切な回答を導く
Table:… See the full description on the dataset page: https://huggingface.co/datasets/neoai-inc/Japanese-RAG-Generator-Benchmark.Prompt-Generator_TestStory to Prompt Conversion Dataset or Prompt to Story
Dataset Summary
The "Story to Prompt Conversion Dataset" is designed to convert narrative text into concise and relevant prompts.
This is still a work in progress, so it currently only contains 100 rows.I have around 6.1K rows of non-public data. Please get in touch with me if you need the rest of the data.
Dataset is for short story generation about 300 to 400 Words.
Supported Tasks and Leaderboards
Text… See the full description on the dataset page: https://huggingface.co/datasets/SayedAmaan/Prompt-Generator_Test.QA-Dataset-Generator
RAG Scientific QA Dataset (Generated)
Dataset Description
This dataset contains 711 high-quality Question-Answering pairs synthetically generated from ArXiv scientific papers. It is specifically designed to fine-tune Large Language Models (LLMs) for Retrieval-Augmented Generation (RAG) tasks.
Source Data: 200 ArXiv papers (Computer Science: AI, CL, LG, IR).
Generation Method: Generated using gpt-4o-mini with strict rules to prevent hallucination.
Language:… See the full description on the dataset page: https://huggingface.co/datasets/xunnhi/QA-Dataset-Generator.question-generator-dataset
Dataset Card for question-generator-dataset
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/rjt1221/question-generator-dataset/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/rjt1221/question-generator-dataset.cognitive-question-generator-v1
Cognitive Question Generator Dataset
Dataset for fine-tuning an expert analysis and question generation model. Contains 5,637 prompt-response pairs capturing expert reasoning patterns for technology transactions and product counseling.
Dataset Description
This dataset was generated from the CognitiveTrainer platform's Mode 1 (Expert Analysis) system, capturing:
Initial scenario analysis
Claim validation with chain-of-trust
Multi-turn expert dialogue
Final synthesis… See the full description on the dataset page: https://huggingface.co/datasets/KevinKeller/cognitive-question-generator-v1.Automated_Model_Generator_Q_and_AAutomated_Model_Generator_Prompt
