CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01vidore /vidore_v3_computer_scienceViDoRe V3 : Computer Science This dataset, Computer Science, is a corpus of textbooks from the openstacks website, intended for long-document understanding tasks. It is one of the 10 corpora comprising the ViDoRe v3 Benchmark. About ViDoRe v3 ViDoRe V3 is our latest benchmark for RAG evaluation on visually-rich documents from real-world applications. It features 10 datasets with, in total, 26,000 pages and 3099 queries, translated into 6 languages. Each query comes with… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_computer_science.documentvisual-document-retrieval1K<n<10K6 likes2.2k downloads8mo agoHugging Face02vidore /vidore_v3_computer_science_mteb_format Vidore3ComputerScienceRetrieval An MTEB dataset Massive Text Embedding Benchmark Retrieve associated pages according to questions. Task category t2i Domains Academic Reference https://huggingface.co/blog/QuentinJG/introducing-vidore-v3 Source datasets: vidore/vidore_v3_computer_science How to evaluate on this task You can evaluate an embedding model on this dataset using the following code: import mteb task =… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_computer_science_mteb_format.imagevisual-document-retrieval10K<n<100K0 likes956 downloads11mo agoHugging Face03oss-codes /Computer-Science-Conversational-Dataset-Indictext10K<n<100K0 likes718 downloads1y agoHugging Face04ComputerScienceHouse /GroceryInContextimagen<1K0 likes283 downloads2y agoHugging Face05Lots-of-LoRAs /task701_mmmlu_answer_generation_high_school_computer_science Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task701_mmmlu_answer_generation_high_school_computer_science Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task701_mmmlu_answer_generation_high_school_computer_science.texttext-generationn<1K1 likes80 downloads2y agoHugging Face06Kaeyze /computer-science-synthetic-datasettext10K<n<100K12 likes74 downloads2y agoHugging Face07pritamdeb68 /Computer-Science-Pretrainingtext100K<n<1M1 likes72 downloads1y agoHugging Face08WenxingZhu /vidore_v3_computer_science_embeddingNOTE ViDoRe V3: Computer Science dataset ColQwen2 Embeddings This dataset contains pre-computed embeddings for the ViDoRe V3 : Computer Science dataset using the ColQwen2 model. ViDoRe V3 : Computer Science This dataset, Computer Science, is a corpus of textbooks from the openstacks website, intended for long-document understanding tasks. It is one of the 10 corpora comprising the ViDoRe v3 Benchmark. About ViDoRe v3 ViDoRe V3 is our latest benchmark for RAG evaluation on… See the full description on the dataset page: https://huggingface.co/datasets/WenxingZhu/vidore_v3_computer_science_embedding.document1K<n<10K0 likes56 downloads11mo agoHugging Face09WithinUsAI /Computer_Science_25k CS_Archon_25k (Master Scholar) CS_Archon_25k is a 25,000-example dataset intended to train models toward master-scholar capability across advanced computer science and modern computer technology: algorithms, data structures, theory of computation, operating systems and performance engineering, distributed systems, networking, databases, compilers/programming languages, ML systems engineering, security (defensive), HCI/product experimentation, and software… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Computer_Science_25k.text10K<n<100K4 likes44 downloads9mo agoHugging Face10gabrieljimenez /epfl-computer-science-mcqatabularn<1K0 likes40 downloads1y agoHugging Face11robworks-software /k12-computer-science-standards K-12 Computer Science Standards 696 generated learning-objective records organized around K-12 computer science concept areas, including computing systems, networks, data, algorithms, programming, AI/ML, cybersecurity, data science, and robotics. How this was built (read this first) These records are programmatically generated, not transcribed from official standards documents. A generator took a standards taxonomy - codes, grade levels, domains, and similar… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-computer-science-standards.textn<1K0 likes39 downloads2mo agoHugging Face12Puidii /aalen_university_faculty_computer_science Dataset Card This dataset contains question-answer pairs from all study programmes of the Faculty of Computer Science at the University of Aalen, Germany. The training dataset is automatically generated by ChatGPT. The validation dataset was manually created. It was collected to train an answer-Q&A chatbot based on LLM fine-tuning. All used scripts and examples can be found in the linked GitHub repository (https://github.com/pattplatt/llm_dataset_creation_and_finetuning).… See the full description on the dataset page: https://huggingface.co/datasets/Puidii/aalen_university_faculty_computer_science.textquestion-answering1K<n<10K0 likes31 downloads2y agoHugging Face13theprint /ComputerSciencetext1K<n<10K0 likes30 downloads4d agoHugging Face14joey234 /mmlu-college_computer_science Dataset Card for "mmlu-college_computer_science" More Information needed textn<1K4 likes26 downloads3y agoHugging Face15masoudc /mmlu-college-computer-science-compilers Extensions to the MMLU Computer Science Datasets for specialization in compilers This dataset contains data specialized in the compilers domain. ** Dataset Details ** Number of rows: 95 Columns: topic, context, question, options, correct_options_literal, correct_options, correct_options_idx ** Usage ** To load this dataset: python from datasets import load_dataset dataset = load_dataset("masoudc/mmlu-college-computer-science-compilers") textn<1K1 likes26 downloads2y agoHugging Face16yourbench /reproduction_qwen235b_computersciencetext1K<n<10K0 likes26 downloads1y agoHugging Face17Lots-of-LoRAs /task688_mmmlu_answer_generation_college_computer_science Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task688_mmmlu_answer_generation_college_computer_science Additional Information Citation Information The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it: @misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions, title={Super-NaturalInstructions: Generalization via Declarative Instructions… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task688_mmmlu_answer_generation_college_computer_science.texttext-generationn<1K1 likes25 downloads2y agoHugging Face18yourbench /reproduction_o4mini_computersciencetext1K<n<10K0 likes23 downloads1y agoHugging Face19Jerichog0731 /vidore_v3_computer_scienceViDoRe V3 : Computer Science This dataset, Computer Science, is a corpus of textbooks from the openstacks website, intended for long-document understanding tasks. It is one of the 10 corpora comprising the ViDoRe v3 Benchmark. About ViDoRe v3 ViDoRe V3 is our latest benchmark for RAG evaluation on visually-rich documents from real-world applications. It features 10 datasets with, in total, 26,000 pages and 3099 queries, translated into 6 languages. Each query comes with… See the full description on the dataset page: https://huggingface.co/datasets/Jerichog0731/vidore_v3_computer_science.documentvisual-document-retrieval1K<n<10K0 likes23 downloads3mo agoHugging Face20masoudc /mmlu-college-computer-science-distribution-parallelism Extensions to the MMLU Computer Science Datasets for specialization in distribution-parallelism This dataset contains data specialized in the distribution-parallelism domain. ** Dataset Details ** Number of rows: 182 Columns: topic, context, question, options, correct_options_literal, correct_options, correct_options_idx ** Usage ** To load this dataset: python from datasets import load_dataset dataset = load_dataset("masoudc/mmlu-college-computer-science-distribution-parallelism") textn<1K1 likes22 downloads2y agoHugging Face21yourbench /yourbench_reproduction_o4mini_computersciencetext1K<n<10K0 likes22 downloads1y agoHugging Face22AlaaElhilo /Wikipedia_ComputerSciencetext1K<n<10K3 likes21 downloads2y agoHugging Face23brucewlee1 /mmlu-college-computer-sciencetextn<1K0 likes19 downloads3y agoHugging Face24joey234 /mmlu-high_school_computer_science-neg-answer Dataset Card for "mmlu-high_school_computer_science-neg-answer" More Information needed textn<1K1 likes18 downloads3y agoHugging Face25joey234 /mmlu-college_computer_science-neg-prepend-verbal Dataset Card for "mmlu-college_computer_science-neg-prepend-verbal" More Information needed textn<1K1 likes16 downloads3y agoHugging Face26yourbench /reproduction_g3_mini_computersciencetext1K<n<10K0 likes16 downloads1y agoHugging Face27harisarang /benchmark-vidore-v3-computer-sciencetext1K<n<10K0 likes16 downloads6mo agoHugging Face28joey234 /mmlu-college_computer_science-neg Dataset Card for "mmlu-college_computer_science-neg" More Information needed textn<1K4 likes15 downloads3y agoHugging Face295CD-AI /Viet-ComputerScience-VQAgated Dataset Overview This dataset is was created from 6899 Vietnamese 🇻🇳 Computer Science books. Each image has been analyzed and annotated using advanced Visual Question Answering (VQA) techniques to produce a comprehensive dataset. There is a set of 40,000 detailed descriptions, and query-based questions and answers generated by the Gemini 1.5 Flash model, currently Google's leading model on the WildVision Arena Leaderboard. This results in a richly annotated dataset, ideal for… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Viet-ComputerScience-VQA.image1K<n<10K2 likes15 downloads2y agoHugging Face30yourbench /reproduction_deepseekr1_computersciencetext1K<n<10K0 likes15 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.