shanaka95/enterprise-rag-questions
shanaka95/enterprise-rag-questions A focused training corpus for domain-adapted embedding models targeting retrieval over enterprise documents. Each row is a (doc_id, question) pair where the question is synthetically generated by an instruction-tuned LLM (Gemma-4-12B-it-qat) from a corresponding source document. The dataset is derived from the public benchmark onyx-dot-app/EnterpriseRAG-Bench. For each document in the source corpus we generate five diverse, answerable training… See the full description on the dataset page: https://huggingface.co/datasets/shanaka95/enterprise-rag-questions.
065
Add dataset card
Initial upload: 2.56M (doc_id, question) pairs from EnterpriseRAG-Bench documents
initial commit
