CoolFace
Datasetpublic

samuelb98/ds_final_p

CourseMate Synthetic QA 10k Dataset of synthetic question–answer pairs grounded in short context chunks extracted from course PDFs (Intro to Data Science).Each row contains: question, answer, context + metadata (topic, doc_name, page, chunk_id, difficulty, generation info). Quick stats Rows: 10000 Topics (lectures): 11 Documents: 11 Distinct pages: 33 Distinct chunks (doc_name, chunk_id): 374 Question unique rate: 0.0655 Distributions… See the full description on the dataset page: https://huggingface.co/datasets/samuelb98/ds_final_p.

sourceHugging Faceotherupdated 8mo agoView on Hugging Face
0likes23downloads
Dataset Card

CourseMate Synthetic QA 10k

Dataset of synthetic question–answer pairs grounded in short context chunks extracted from course PDFs (Intro to Data Science). Each row contains: question, answer, context + metadata (topic, doc_name, page, chunk_id, difficulty, generation info).

Quick stats

  • —Rows: 10000
  • —Topics (lectures): 11
  • —Documents: 11
  • —Distinct pages: 33
  • —Distinct chunks (docname, chunkid): 374
  • —Question unique rate: 0.0655

Distributions

Topics

[image]

Difficulty

[image]

Context length

[image]

Text length (chars)

question_len

  • —min: 24, p50: 51, p95: 95, max: 136, mean: 55.81

answer_len

  • —min: 27, p50: 76, p95: 100, max: 119, mean: 76.24

context_len

  • —min: 83, p50: 253, p95: 513, max: 519, mean: 272.18

Columns

id, topic, docname, page, chunkid, context, answer, question, difficulty, generatormodel, generatortype, questionlen, answerlen, context_len

Notes

  • —topic is the lecture label.
  • —doc_name and page point back to the source PDF page.
  • —chunk_id identifies the chunk within the document.