CoolFace
Datasetpublic

vaadewoyin/arxiv-ml-qa-dataset

ArXiv ML Q&A Dataset Dataset Description 951 high-quality Q&A pairs generated from ArXiv machine learning paper abstracts. Built for fine-tuning language models on technical ML questions. Repository: GitHub Repository - Source code for dataset generation, cleaning, and filtering. How It Was Built Scraped 2000+ ArXiv ML papers via ArXiv API (cat:cs.LG) Cleaned — deduplication, length filter (50–400 words) Generated — Llama-3-8B-Instruct via… See the full description on the dataset page: https://huggingface.co/datasets/vaadewoyin/arxiv-ml-qa-dataset.

sourceHugging Facemitupdated 4mo agoView on Hugging Face
0likes9downloads
7 commits on main
93263034mo ago

Upload dataset

vaadewoyin
4c9d49b4mo ago

Update README.md

vaadewoyin
d9acec44mo ago

Update README.md

vaadewoyin
bf24fd14mo ago

Update README.md

vaadewoyin
16b752f4mo ago

Update README.md

vaadewoyin
d7ff97c4mo ago

Upload dataset

vaadewoyin
e0d7f884mo ago

initial commit

vaadewoyin