datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vidore_v3_computer_scienceViDoRe V3 : Computer Science
This dataset, Computer Science, is a corpus of textbooks from the openstacks website, intended for long-document understanding tasks. It is one of the 10 corpora comprising the ViDoRe v3 Benchmark.
About ViDoRe v3
ViDoRe V3 is our latest benchmark for RAG evaluation on visually-rich documents from real-world applications. It features 10 datasets with, in total, 26,000 pages and 3099 queries, translated into 6 languages. Each query comes with… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_computer_science.Computer-Science
מאגר הנתונים CS26 HIT — מדעי המחשב
מאגר זה משמש לאחסון מרכזי של נתוני לימוד ומשאבים אקדמיים עבור סטודנטים למדעי המחשב במכון הטכנולוגי חולון. המידע המצוי כאן מונגש בצורה נוחה באמצעות פורטל גישה נפרד המאפשר ניווט ויזואלי וחיפוש יעיל בתוך התיקיות השונות.
קישורים וגישה
ניתן להשתמש בפורטל בכתובת https://cs26-cs26-portal.hf.space/
תנאי שימוש וזכויות יוצרים
כל חומרי הלימוד והתכנים המופיעים במאגר זה פתוחים וחופשיים לשימוש לצורכי למידה בלבד. ניתן לקחת את… See the full description on the dataset page: https://huggingface.co/datasets/CS26/Computer-Science.vidore_v3_computer_science_mteb_format
Vidore3ComputerScienceRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
Retrieve associated pages according to questions.
Task category
t2i
Domains
Academic
Reference
https://huggingface.co/blog/QuentinJG/introducing-vidore-v3
Source datasets:
vidore/vidore_v3_computer_science
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task =… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_computer_science_mteb_format.Computer-Science-Conversational-Dataset-IndicComputer-Science-Parallel-Dataset-IndicHandwritten-Computer-Science-Notes-Dataset
English Handwritten Computer Science Notes Dataset
This dataset contains high-resolution images of handwritten computer science notes written in English. It includes algorithm explanations, code snippets, flowcharts, theoretical content, and annotations. The dataset is designed to support AI research in handwriting recognition, OCR, and document understanding specifically for computer science education.
Contact
For queries or collaborations related to this dataset… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Handwritten-Computer-Science-Notes-Dataset.GroceryInContextUniversity-level_Mathematics_Physics_Chemistry_Computer_Science_Reasoning_Corpus
Title
University-level Mathematics, Physics, Chemistry, Computer Science Reasoning Corpus
Size
200,000+ text+ multimodal university-level problems, each with step-by-step solutions and final answers
Format
Natural language explanations with multimodal samples include images (graphs, diagrams, etc.)
Subject
Mathematics, Physics, Chemistry, Computer Science
Labeling Details
Question ID/Question Stem (Full text/content) /Subject/Question Type… See the full description on the dataset page: https://huggingface.co/datasets/DataoceanAI/University-level_Mathematics_Physics_Chemistry_Computer_Science_Reasoning_Corpus.task701_mmmlu_answer_generation_high_school_computer_science
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task701_mmmlu_answer_generation_high_school_computer_science
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task701_mmmlu_answer_generation_high_school_computer_science.computer-science-synthetic-datasetComputer-Science-Pretrainingvidore_v3_computer_science_embeddingNOTE
ViDoRe V3: Computer Science dataset ColQwen2 Embeddings
This dataset contains pre-computed embeddings for the ViDoRe V3 : Computer Science dataset using the ColQwen2 model.
ViDoRe V3 : Computer Science
This dataset, Computer Science, is a corpus of textbooks from the openstacks website, intended for long-document understanding tasks. It is one of the 10 corpora comprising the ViDoRe v3 Benchmark.
About ViDoRe v3
ViDoRe V3 is our latest benchmark for RAG evaluation on… See the full description on the dataset page: https://huggingface.co/datasets/WenxingZhu/vidore_v3_computer_science_embedding.k12-computer-science-standards
K-12 Computer Science Standards
696 generated learning-objective records organized around K-12 computer science concept
areas, including computing systems, networks, data, algorithms, programming, AI/ML,
cybersecurity, data science, and robotics.
How this was built (read this first)
These records are programmatically generated, not transcribed from official standards
documents. A generator took a standards taxonomy - codes, grade levels, domains, and
similar… See the full description on the dataset page: https://huggingface.co/datasets/robworks-software/k12-computer-science-standards.epfl-computer-science-mcqaaalen_university_faculty_computer_science
Dataset Card
This dataset contains question-answer pairs from all study programmes of the Faculty of Computer Science at the University of Aalen, Germany. The training dataset is automatically generated by ChatGPT. The validation dataset was manually created.
It was collected to train an answer-Q&A chatbot based on LLM fine-tuning. All used scripts and examples can be found in the linked GitHub repository (https://github.com/pattplatt/llm_dataset_creation_and_finetuning).… See the full description on the dataset page: https://huggingface.co/datasets/Puidii/aalen_university_faculty_computer_science.task688_mmmlu_answer_generation_college_computer_science
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task688_mmmlu_answer_generation_college_computer_science
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task688_mmmlu_answer_generation_college_computer_science.ARxiv_Metadata_ComputerScienceComputer_Science_25k
CS_Archon_25k (Master Scholar)
CS_Archon_25k is a 25,000-example dataset intended to train models toward master-scholar capability across
advanced computer science and modern computer technology: algorithms, data structures, theory of computation,
operating systems and performance engineering, distributed systems, networking, databases, compilers/programming languages,
ML systems engineering, security (defensive), HCI/product experimentation, and software… See the full description on the dataset page: https://huggingface.co/datasets/WithinUsAI/Computer_Science_25k.mmlu-college_computer_science
Dataset Card for "mmlu-college_computer_science"
More Information needed
reproduction_qwen235b_computersciencemmlu-college-computer-science-compilers
Extensions to the MMLU Computer Science Datasets for specialization in compilers
This dataset contains data specialized in the compilers domain.
** Dataset Details **
Number of rows: 95
Columns: topic, context, question, options, correct_options_literal, correct_options, correct_options_idx
** Usage **
To load this dataset:
python
from datasets import load_dataset
dataset = load_dataset("masoudc/mmlu-college-computer-science-compilers")
mmlu-college-computer-science-distribution-parallelism
Extensions to the MMLU Computer Science Datasets for specialization in distribution-parallelism
This dataset contains data specialized in the distribution-parallelism domain.
** Dataset Details **
Number of rows: 182
Columns: topic, context, question, options, correct_options_literal, correct_options, correct_options_idx
** Usage **
To load this dataset:
python
from datasets import load_dataset
dataset = load_dataset("masoudc/mmlu-college-computer-science-distribution-parallelism")
yourbench_reproduction_o4mini_computersciencemmlu-high_school_computer_science-neg-answer
Dataset Card for "mmlu-high_school_computer_science-neg-answer"
More Information needed
Wikipedia_ComputerSciencereproduction_o4mini_computersciencemmlu-college_computer_science-neg
Dataset Card for "mmlu-college_computer_science-neg"
More Information needed
mmlu-college-computer-sciencemmlu-college_computer_science-neg-prepend
Dataset Card for "mmlu-college_computer_science-neg-prepend"
More Information needed
mmlu-college_computer_science-neg-answer
Dataset Card for "mmlu-college_computer_science-neg-answer"
More Information needed
