CoolFace
Datasetpublic

chauben/stevens-qa-finetuning-87k

Stevens Institute of Technology — Q&A Fine-Tuning Dataset 87,782 context-grounded question-answer pairs scraped and generated from the official Stevens Institute of Technology website, used to fine-tune LLaMA-2-7B for domain-specific academic advising. Dataset Structure Each entry has three fields: context: Source URL + scraped web content question: Natural language question answerable from context answer: Concise, grounded answer Stats Total… See the full description on the dataset page: https://huggingface.co/datasets/chauben/stevens-qa-finetuning-87k.

sourceHugging Faceupdated 6mo agoView on Hugging Face
0likes23downloads
Dataset Card

Stevens Institute of Technology — Q&A Fine-Tuning Dataset

87,782 context-grounded question-answer pairs scraped and generated from the official Stevens Institute of Technology website, used to fine-tune LLaMA-2-7B for domain-specific academic advising.

Dataset Structure

Each entry has three fields:

  • —context: Source URL + scraped web content
  • —question: Natural language question answerable from context
  • —answer: Concise, grounded answer

Stats

  • —Total pairs: 87,782
  • —Format: JSONL
  • —Train/Eval split: 90% / 10%
  • —Sources: Stevens programs, course catalog, faculty, admissions, campus

Used In

Citation

Chaube N., Jadhav P., Sompura K. (2025). AdvisorAI: A Retrieval-Augmented Generation System with Fine-Tuned LLaMA for Domain-Specific Academic Advising. Stevens Institute of Technology.

chauben/stevens-qa-finetuning-87k · CoolFace