datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vidore_v3_computer_scienceViDoRe V3 : Computer Science
This dataset, Computer Science, is a corpus of textbooks from the openstacks website, intended for long-document understanding tasks. It is one of the 10 corpora comprising the ViDoRe v3 Benchmark.
About ViDoRe v3
ViDoRe V3 is our latest benchmark for RAG evaluation on visually-rich documents from real-world applications. It features 10 datasets with, in total, 26,000 pages and 3099 queries, translated into 6 languages. Each query comes with… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_computer_science.vidore_v3_computer_science_mteb_format
Vidore3ComputerScienceRetrieval
An MTEB dataset
Massive Text Embedding Benchmark
Retrieve associated pages according to questions.
Task category
t2i
Domains
Academic
Reference
https://huggingface.co/blog/QuentinJG/introducing-vidore-v3
Source datasets:
vidore/vidore_v3_computer_science
How to evaluate on this task
You can evaluate an embedding model on this dataset using the following code:
import mteb
task =… See the full description on the dataset page: https://huggingface.co/datasets/vidore/vidore_v3_computer_science_mteb_format.Handwritten-Computer-Science-Notes-Dataset
English Handwritten Computer Science Notes Dataset
This dataset contains high-resolution images of handwritten computer science notes written in English. It includes algorithm explanations, code snippets, flowcharts, theoretical content, and annotations. The dataset is designed to support AI research in handwriting recognition, OCR, and document understanding specifically for computer science education.
Contact
For queries or collaborations related to this dataset… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/Handwritten-Computer-Science-Notes-Dataset.GroceryInContextvidore_v3_computer_science_embeddingNOTE
ViDoRe V3: Computer Science dataset ColQwen2 Embeddings
This dataset contains pre-computed embeddings for the ViDoRe V3 : Computer Science dataset using the ColQwen2 model.
ViDoRe V3 : Computer Science
This dataset, Computer Science, is a corpus of textbooks from the openstacks website, intended for long-document understanding tasks. It is one of the 10 corpora comprising the ViDoRe v3 Benchmark.
About ViDoRe v3
ViDoRe V3 is our latest benchmark for RAG evaluation on… See the full description on the dataset page: https://huggingface.co/datasets/WenxingZhu/vidore_v3_computer_science_embedding.Viet-ComputerScience-VQA
Dataset Overview
This dataset is was created from 6899 Vietnamese 🇻🇳 Computer Science books. Each image has been analyzed and annotated using advanced Visual Question Answering (VQA) techniques to produce a comprehensive dataset.
There is a set of 40,000 detailed descriptions, and query-based questions and answers generated by the Gemini 1.5 Flash model, currently Google's leading model on the WildVision Arena Leaderboard. This results in a richly annotated dataset, ideal for… See the full description on the dataset page: https://huggingface.co/datasets/5CD-AI/Viet-ComputerScience-VQA.vidore_v3_computer_science_english_open-ended_Chartvidore3_computerscience_neomme_260m_li
vidore3_computerscience_neomme_260m_li
Multi-vector (late-interaction) embeddings of ViDoRe computerscience (vidore/computerscience), encoded with
Hcompany/NeoMME-260M-Retriever-ST-late at revision 023be2a8ab9d797f5aa76f5bf8b5dde78d819659.
Source data: Hugging Face dataset vidore/vidore_v3_computer_science at revision d5cc75883d92e294f0c0fc2662551c9708a06ebc, configs corpus / queries / qrels, split test, loaded with datasets. Document, query and qrel ids are the source's own ids… See the full description on the dataset page: https://huggingface.co/datasets/robro612/vidore3_computerscience_neomme_260m_li.vidore_v3_computer_science_english_instructionvidore_v3_computer_science_english_boolean_Othervidore_v3_computer_science_english_compare-contrast_Othervidore_v3_computer_science_english_open-ended_TableComputer-sciencevidore_v3_computer_science_english_keywordvidore_v3_computer_science_english_booleanvidore_v3_computer_science_english_enumerativevidore_v3_computer_science_english_extractivevidore_v3_computer_science_english_Textvidore_v3_computer_science_english_compare-contrast_Infographicvidore_v3_computer_science_english_compare-contrastvidore_v3_computer_science_english_Imagevidore_v3_computer_science_english_compare-contrast_Chartvidore_v3_computer_science_english_enumerative_Chartvidore_v3_computer_science_english_compare-contrast_Imagevidore_v3_computer_science_english_extractive_Imagevidore_v3_computer_science_english_open-ended_Imagevidore_v3_computer_science_english_boolean_Infographicvidore_v3_computer_science_english_enumerative_Infographicvidore_v3_computer_science_english_extractive_Infographicvidore_v3_computer_science_english_enumerative_Other
