vaadewoyin/arxiv-ml-qa-dataset
ArXiv ML Q&A Dataset Dataset Description 951 high-quality Q&A pairs generated from ArXiv machine learning paper abstracts. Built for fine-tuning language models on technical ML questions. Repository: GitHub Repository - Source code for dataset generation, cleaning, and filtering. How It Was Built Scraped 2000+ ArXiv ML papers via ArXiv API (cat:cs.LG) Cleaned — deduplication, length filter (50–400 words) Generated — Llama-3-8B-Instruct via… See the full description on the dataset page: https://huggingface.co/datasets/vaadewoyin/arxiv-ml-qa-dataset.
09
Upload dataset
Update README.md
Update README.md
Update README.md
Update README.md
Upload dataset
initial commit
