datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Computer-Science-Conversational-Dataset-IndicComputer_Science_25k
CS_Archon_25k (Master Scholar)
CS_Archon_25k is a 25,000-example dataset intended to train models toward master-scholar capability across
advanced computer science and modern computer technology: algorithms, data structures, theory of computation,
operating systems and performance engineering, distributed systems, networking, databases, compilers/programming languages,
ML systems engineering, security (defensive), HCI/product experimentation, and software… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Computer_Science_25k.aalen_university_faculty_computer_science
Dataset Card
This dataset contains question-answer pairs from all study programmes of the Faculty of Computer Science at the University of Aalen, Germany. The training dataset is automatically generated by ChatGPT. The validation dataset was manually created.
It was collected to train an answer-Q&A chatbot based on LLM fine-tuning. All used scripts and examples can be found in the linked GitHub repository (https://github.com/pattplatt/llm_dataset_creation_and_finetuning).… See the full description on the dataset page: https://huggingface.co/datasets/Puidii/aalen_university_faculty_computer_science.ComputerSciencemmlu-college-computer-sciencebenchmark-vidore-v3-computer-sciencemmlu-high-school-computer-scienceSciTrust2-ComputerScienceQAcomputer_science_non_ai_search_queries
