samuelb98/ds_final_p
CourseMate Synthetic QA 10k Dataset of synthetic question–answer pairs grounded in short context chunks extracted from course PDFs (Intro to Data Science).Each row contains: question, answer, context + metadata (topic, doc_name, page, chunk_id, difficulty, generation info). Quick stats Rows: 10000 Topics (lectures): 11 Documents: 11 Distinct pages: 33 Distinct chunks (doc_name, chunk_id): 374 Question unique rate: 0.0655 Distributions… See the full description on the dataset page: https://huggingface.co/datasets/samuelb98/ds_final_p.
CourseMate Synthetic QA 10k
Dataset of synthetic question–answer pairs grounded in short context chunks extracted from course PDFs (Intro to Data Science). Each row contains: question, answer, context + metadata (topic, doc_name, page, chunk_id, difficulty, generation info).
Quick stats
- Rows: 10000
- Topics (lectures): 11
- Documents: 11
- Distinct pages: 33
- Distinct chunks (docname, chunkid): 374
- Question unique rate: 0.0655
Distributions
Topics
Difficulty
Context length
Text length (chars)
question_len
- min: 24, p50: 51, p95: 95, max: 136, mean: 55.81
answer_len
- min: 27, p50: 76, p95: 100, max: 119, mean: 76.24
context_len
- min: 83, p50: 253, p95: 513, max: 519, mean: 272.18
Columns
id, topic, docname, page, chunkid, context, answer, question, difficulty, generatormodel, generatortype, questionlen, answerlen, context_len
Notes
topicis the lecture label.doc_nameandpagepoint back to the source PDF page.chunk_ididentifies the chunk within the document.
