chauben/stevens-qa-finetuning-87k
Stevens Institute of Technology — Q&A Fine-Tuning Dataset 87,782 context-grounded question-answer pairs scraped and generated from the official Stevens Institute of Technology website, used to fine-tune LLaMA-2-7B for domain-specific academic advising. Dataset Structure Each entry has three fields: context: Source URL + scraped web content question: Natural language question answerable from context answer: Concise, grounded answer Stats Total… See the full description on the dataset page: https://huggingface.co/datasets/chauben/stevens-qa-finetuning-87k.
Stevens Institute of Technology — Q&A Fine-Tuning Dataset
87,782 context-grounded question-answer pairs scraped and generated from the official Stevens Institute of Technology website, used to fine-tune LLaMA-2-7B for domain-specific academic advising.
Dataset Structure
Each entry has three fields:
context: Source URL + scraped web contentquestion: Natural language question answerable from contextanswer: Concise, grounded answer
Stats
- Total pairs: 87,782
- Format: JSONL
- Train/Eval split: 90% / 10%
- Sources: Stevens programs, course catalog, faculty, admissions, campus
Used In
- AdvisorAI — production chatbot
- StevensDomainFineTunedLM — fine-tuning pipeline
Citation
Chaube N., Jadhav P., Sompura K. (2025). AdvisorAI: A Retrieval-Augmented Generation System with Fine-Tuned LLaMA for Domain-Specific Academic Advising. Stevens Institute of Technology.
