datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Multi-Source-Video-Captioning
Multi-source Video Captioning (MSVC) Dataset Card
Dataset details
Dataset type:
MSVC is a set of collected video captioning data. It is constructed to ensure a robust and thorough evaluation of Video-LLMs' video-captioning capabilities.
Dataset detail:
MSVC is introduced to address limitations in existing video caption benchmarks, MSVC samples a total of 1,500 videos with human-annotated captions from MSVD, MSRVTT, and VATEX, ensuring diverse scenarios and domains.… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/Multi-Source-Video-Captioning.mhlc-training-qwen3vl-qwen3_vl_2b_thinking_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3-VL 2B Thinking hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_2b_thinking_hard_mixed_sources_120k.mhlc-training-qwen3vl-qwen3-vl-32b-instruct_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3-VL 32B Instruct hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3-vl-32b-instruct_hard_mixed_sources_120k.mhlc-training-qwen3vl-qwen3_vl_4b_instruct_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3-VL 4B Instruct hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_4b_instruct_hard_mixed_sources_120k.mhlc-training-qwen3vl-qwen3_vl_4b_thinking_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3-VL 4B Thinking hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_4b_thinking_hard_mixed_sources_120k.mhlc-training-qwen3vl-qwen3_vl_2b_instruct_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3-VL 2B Instruct hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_2b_instruct_hard_mixed_sources_120k.mhlc-training-qwen3.5-qwen3_5_4b_think_off_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3.5 4B think off hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3.5-qwen3_5_4b_think_off_hard_mixed_sources_120k.mhlc-training-gemma4-gemma4_e4b_it_think_on_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Gemma 4 E4B it think on hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-gemma4-gemma4_e4b_it_think_on_hard_mixed_sources_120k.mhlc-training-qwen3.5-qwen3_5_9b_think_off_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Qwen3.5 9B think off hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3.5-qwen3_5_9b_think_off_hard_mixed_sources_120k.mhlc-training-gemma4-gemma4_e4b_it_think_off_hard_mixed_sources_120k
Multi Head Latent Control Training Data - Gemma 4 E4B it think off hard Mixed Sources 120k
Dataset Description
This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection.
Paper
https://arxiv.org/abs/2607.14277
Code
https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control
Dataset Summary
Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-gemma4-gemma4_e4b_it_think_off_hard_mixed_sources_120k.olympiad-books-open-source
olympiad-books-open-source
Chunked content from 12 open-source mathematics textbooks, suitable for retrieval (RAG), embedding, and math reasoning research.
Source code: github.com/yoonholee/olympiad-books-open-source-pipeline
Books
Book
Author(s)
License
Source
An Infinitely Large Napkin
Evan Chen
CC BY-SA 4.0 / GPL v3
GitHub
Mathematical Reasoning: Writing and Proof
Ted Sundstrom
CC BY-NC-SA 3.0
GitHub
Exploring Combinatorial Mathematics
Richard Grassl… See the full description on the dataset page: https://huggingface.co/datasets/yoonholee/olympiad-books-open-source.biosciences-sources
Biosciences RAG Source Documents
Dataset Description
This dataset contains 140 page-level document chunks extracted from 10 biomedical research papers. The documents form the knowledge base for a Retrieval-Augmented Generation (RAG) system focused on biosciences topics including knowledge graphs, LLM applications in biomedicine, and protein interaction databases.
Dataset Summary
Total Documents: 140 pages from 10 research papers
Domain: Biomedical NLP… See the full description on the dataset page: https://huggingface.co/datasets/open-biosciences/biosciences-sources.RLVR-Env-Retrieval-Source-Retrieval-Synthetic-NVDocs-v1
RLVR-Env-Retrieval-Source-Retrieval-Synthetic-NVDocs-v1
RLVR-ready retrieval environment derived from nvidia/Retrieval-Synthetic-NVDocs-v1.
Author: Aman Priyanshu
What Is This
A 100k-row retrieval QA dataset where each row contains a question, ground-truth chunks, and pre-mined distractor chunks (random + semantically similar). Designed for training and evaluating retrieval agents in an RLVR (Reinforcement Learning with Verifiable Rewards) setup — the agent searches… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/RLVR-Env-Retrieval-Source-Retrieval-Synthetic-NVDocs-v1.RLVR-Env-Retrieval-Source-code-search-net-javascript
RLVR-Env-Retrieval-Source-code-search-net-javascript
RLVR-ready retrieval environment derived from Nan-Do/code-search-net-javascript.
Author: Aman Priyanshu
What Is This
A 100k-row retrieval QA dataset where each row contains a question, ground-truth chunks, and pre-mined distractor chunks (random + semantically similar). Designed for training and evaluating retrieval agents in an RLVR (Reinforcement Learning with Verifiable Rewards) setup — the agent searches through… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/RLVR-Env-Retrieval-Source-code-search-net-javascript.RLVR-Env-Retrieval-Source-code-search-net-python
RLVR-Env-Retrieval-Source-code-search-net-python
RLVR-ready retrieval environment derived from Nan-Do/code-search-net-python.
Author: Aman Priyanshu
What Is This
A 100k-row retrieval QA dataset where each row contains a question, ground-truth chunks, and pre-mined distractor chunks (random + semantically similar). Designed for training and evaluating retrieval agents in an RLVR (Reinforcement Learning with Verifiable Rewards) setup — the agent searches through distractors… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/RLVR-Env-Retrieval-Source-code-search-net-python.gdelt-rag-sources
GDELT RAG Source Documents
Dataset Description
This dataset contains source documents extracted from the research paper "Talking to GDELT Through Knowledge Graphs"
(arXiv:2503.07584v3). The documents are used as the knowledge base for a Retrieval-Augmented Generation (RAG) system
focused on GDELT (Global Database of Events, Language, and Tone) analysis.
Dataset Summary
Total Documents: 38 pages
Source: Research paper on GDELT Knowledge Graphs
Format: PDF pages… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-sources.river_source
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/antonixe/river_source.hudanet-sources
HUDA-Net Sources
Hajj and Umrah Jurisprudence Corpus
مصادر شبكة هدى لفقه الحج والعمرة
HUDA-Net Sources is the private, version-controlled source repository used to build the HUDA-Net bilingual question-answering, semantic retrieval, and text-ranking system for Hajj and Umrah jurisprudence.
This repository preserves the original book datasets, cleaned source files, provenance metadata, and integrity manifests required to reproduce and audit the HUDA-Net corpus.… See the full description on the dataset page: https://huggingface.co/datasets/dakheel/hudanet-sources.gdelt-rag-sources-v2
GDELT RAG Source Documents
Dataset Description
This dataset contains source documents extracted from the research paper "Talking to GDELT Through Knowledge Graphs"
(arXiv:2503.07584v3). The documents are used as the knowledge base for a Retrieval-Augmented Generation (RAG) system
focused on GDELT (Global Database of Events, Language, and Tone) analysis.
Dataset Summary
Total Documents: 38 pages
Source: Research paper on GDELT Knowledge Graphs
Format: PDF pages… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-sources-v2.gdelt-rag-sources-v4
GDELT RAG Source Documents
Dataset Description
This dataset contains source documents extracted from the research paper "Talking to GDELT Through Knowledge Graphs"
(arXiv:2503.07584v3). The documents are used as the knowledge base for a Retrieval-Augmented Generation (RAG) system
focused on GDELT (Global Database of Events, Language, and Tone) analysis.
Dataset Summary
Total Documents: 38 pages
Source: Research paper on GDELT Knowledge Graphs
Format: PDF pages… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-sources-v4.gdelt-rag-sources-v3
GDELT RAG Source Documents
Dataset Description
This dataset contains source documents extracted from the research paper "Talking to GDELT Through Knowledge Graphs"
(arXiv:2503.07584v3). The documents are used as the knowledge base for a Retrieval-Augmented Generation (RAG) system
focused on GDELT (Global Database of Events, Language, and Tone) analysis.
Dataset Summary
Total Documents: 38 pages
Source: Research paper on GDELT Knowledge Graphs
Format: PDF pages… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-sources-v3.tigle-source-code
TIGLE
The interface is built as a prototype based on Dzogchen, Atiyoga teachings available in English and sourced, compiled by a practitioner exploring how Dharma language and current global AI could intersect. The architecture, the pipeline works. The answers are useful for orientation — learning key terms, lineages, main practices, understanding the view.
This is why it is accessible as repository rather than a product:
Digital Bardo - is the current state of samsara.… See the full description on the dataset page: https://huggingface.co/datasets/Tigle/tigle-source-code.
