shanaka95/enterprise-rag-questions
shanaka95/enterprise-rag-questions A focused training corpus for domain-adapted embedding models targeting retrieval over enterprise documents. Each row is a (doc_id, question) pair where the question is synthetically generated by an instruction-tuned LLM (Gemma-4-12B-it-qat) from a corresponding source document. The dataset is derived from the public benchmark onyx-dot-app/EnterpriseRAG-Bench. For each document in the source corpus we generate five diverse, answerable training… See the full description on the dataset page: https://huggingface.co/datasets/shanaka95/enterprise-rag-questions.
Conversations for this repository live on Hugging Face.
CoolFace shows imported repositories read-only. Posting into someone else’s repository from here would need an authorised integration and the account holder’s consent, so the link goes to the source instead.
Open discussions on Hugging Face