CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01DAMO-NLP-SG /Multi-Source-Video-Captioning Multi-source Video Captioning (MSVC) Dataset Card Dataset details Dataset type: MSVC is a set of collected video captioning data. It is constructed to ensure a robust and thorough evaluation of Video-LLMs' video-captioning capabilities. Dataset detail: MSVC is introduced to address limitations in existing video caption benchmarks, MSVC samples a total of 1,500 videos with human-annotated captions from MSVD, MSRVTT, and VATEX, ensuring diverse scenarios and domains.… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/Multi-Source-Video-Captioning.textvisual-question-answering1K<n<10K7 likes316 downloads2y agoHugging Face02AmirhoseinGH /mhlc-training-qwen3vl-qwen3_vl_2b_thinking_hard_mixed_sources_120k Multi Head Latent Control Training Data - Qwen3-VL 2B Thinking hard Mixed Sources 120k Dataset Description This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection. Paper https://arxiv.org/abs/2607.14277 Code https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control Dataset Summary Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_2b_thinking_hard_mixed_sources_120k.imagequestion-answering100K<n<1M0 likes314 downloads2mo agoHugging Face03AmirhoseinGH /mhlc-training-qwen3vl-qwen3-vl-32b-instruct_hard_mixed_sources_120k Multi Head Latent Control Training Data - Qwen3-VL 32B Instruct hard Mixed Sources 120k Dataset Description This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection. Paper https://arxiv.org/abs/2607.14277 Code https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control Dataset Summary Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3-vl-32b-instruct_hard_mixed_sources_120k.imagequestion-answering10K<n<100K0 likes289 downloads2mo agoHugging Face04AmirhoseinGH /mhlc-training-qwen3vl-qwen3_vl_4b_instruct_hard_mixed_sources_120k Multi Head Latent Control Training Data - Qwen3-VL 4B Instruct hard Mixed Sources 120k Dataset Description This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection. Paper https://arxiv.org/abs/2607.14277 Code https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control Dataset Summary Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_4b_instruct_hard_mixed_sources_120k.imagequestion-answering100K<n<1M0 likes272 downloads2mo agoHugging Face05AmirhoseinGH /mhlc-training-qwen3vl-qwen3_vl_4b_thinking_hard_mixed_sources_120k Multi Head Latent Control Training Data - Qwen3-VL 4B Thinking hard Mixed Sources 120k Dataset Description This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection. Paper https://arxiv.org/abs/2607.14277 Code https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control Dataset Summary Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_4b_thinking_hard_mixed_sources_120k.imagequestion-answering100K<n<1M0 likes272 downloads2mo agoHugging Face06AmirhoseinGH /mhlc-training-qwen3vl-qwen3_vl_2b_instruct_hard_mixed_sources_120k Multi Head Latent Control Training Data - Qwen3-VL 2B Instruct hard Mixed Sources 120k Dataset Description This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection. Paper https://arxiv.org/abs/2607.14277 Code https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control Dataset Summary Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3vl-qwen3_vl_2b_instruct_hard_mixed_sources_120k.imagequestion-answering100K<n<1M0 likes243 downloads2mo agoHugging Face07AmirhoseinGH /mhlc-training-qwen3.5-qwen3_5_4b_think_off_hard_mixed_sources_120k Multi Head Latent Control Training Data - Qwen3.5 4B think off hard Mixed Sources 120k Dataset Description This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection. Paper https://arxiv.org/abs/2607.14277 Code https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control Dataset Summary Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3.5-qwen3_5_4b_think_off_hard_mixed_sources_120k.imagequestion-answering100K<n<1M0 likes220 downloads2mo agoHugging Face08AmirhoseinGH /mhlc-training-gemma4-gemma4_e4b_it_think_on_hard_mixed_sources_120k Multi Head Latent Control Training Data - Gemma 4 E4B it think on hard Mixed Sources 120k Dataset Description This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection. Paper https://arxiv.org/abs/2607.14277 Code https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control Dataset Summary Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-gemma4-gemma4_e4b_it_think_on_hard_mixed_sources_120k.imagequestion-answering100K<n<1M0 likes194 downloads2mo agoHugging Face09AmirhoseinGH /mhlc-training-qwen3.5-qwen3_5_9b_think_off_hard_mixed_sources_120k Multi Head Latent Control Training Data - Qwen3.5 9B think off hard Mixed Sources 120k Dataset Description This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection. Paper https://arxiv.org/abs/2607.14277 Code https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control Dataset Summary Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-qwen3.5-qwen3_5_9b_think_off_hard_mixed_sources_120k.imagequestion-answering100K<n<1M0 likes179 downloads2mo agoHugging Face10AmirhoseinGH /mhlc-training-gemma4-gemma4_e4b_it_think_off_hard_mixed_sources_120k Multi Head Latent Control Training Data - Gemma 4 E4B it think off hard Mixed Sources 120k Dataset Description This repository contains verified training data for the Multi Head Latent Control paper release. It is part of the Multi Head Latent Control training data Hugging Face collection. Paper https://arxiv.org/abs/2607.14277 Code https://github.com/Amirhosein-gh98/Multi-Head-Latent-Control Dataset Summary Field… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/mhlc-training-gemma4-gemma4_e4b_it_think_off_hard_mixed_sources_120k.imagequestion-answering100K<n<1M0 likes154 downloads2mo agoHugging Face11yoonholee /olympiad-books-open-source olympiad-books-open-source Chunked content from 12 open-source mathematics textbooks, suitable for retrieval (RAG), embedding, and math reasoning research. Source code: github.com/yoonholee/olympiad-books-open-source-pipeline Books Book Author(s) License Source An Infinitely Large Napkin Evan Chen CC BY-SA 4.0 / GPL v3 GitHub Mathematical Reasoning: Writing and Proof Ted Sundstrom CC BY-NC-SA 3.0 GitHub Exploring Combinatorial Mathematics Richard Grassl… See the full description on the dataset page: https://huggingface.co/datasets/yoonholee/olympiad-books-open-source.tabulartext-retrieval1K<n<10K1 likes34 downloads7mo agoHugging Face12open-biosciences /biosciences-sources Biosciences RAG Source Documents Dataset Description This dataset contains 140 page-level document chunks extracted from 10 biomedical research papers. The documents form the knowledge base for a Retrieval-Augmented Generation (RAG) system focused on biosciences topics including knowledge graphs, LLM applications in biomedicine, and protein interaction databases. Dataset Summary Total Documents: 140 pages from 10 research papers Domain: Biomedical NLP… See the full description on the dataset page: https://huggingface.co/datasets/open-biosciences/biosciences-sources.texttext-retrievaln<1K0 likes29 downloads7mo agoHugging Face13AmanPriyanshu /RLVR-Env-Retrieval-Source-Retrieval-Synthetic-NVDocs-v1 RLVR-Env-Retrieval-Source-Retrieval-Synthetic-NVDocs-v1 RLVR-ready retrieval environment derived from nvidia/Retrieval-Synthetic-NVDocs-v1. Author: Aman Priyanshu What Is This A 100k-row retrieval QA dataset where each row contains a question, ground-truth chunks, and pre-mined distractor chunks (random + semantically similar). Designed for training and evaluating retrieval agents in an RLVR (Reinforcement Learning with Verifiable Rewards) setup — the agent searches… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/RLVR-Env-Retrieval-Source-Retrieval-Synthetic-NVDocs-v1.texttext-retrieval100K<n<1M0 likes29 downloads6mo agoHugging Face14AmanPriyanshu /RLVR-Env-Retrieval-Source-code-search-net-javascript RLVR-Env-Retrieval-Source-code-search-net-javascript RLVR-ready retrieval environment derived from Nan-Do/code-search-net-javascript. Author: Aman Priyanshu What Is This A 100k-row retrieval QA dataset where each row contains a question, ground-truth chunks, and pre-mined distractor chunks (random + semantically similar). Designed for training and evaluating retrieval agents in an RLVR (Reinforcement Learning with Verifiable Rewards) setup — the agent searches through… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/RLVR-Env-Retrieval-Source-code-search-net-javascript.texttext-retrieval100K<n<1M0 likes27 downloads7mo agoHugging Face15AmanPriyanshu /RLVR-Env-Retrieval-Source-code-search-net-python RLVR-Env-Retrieval-Source-code-search-net-python RLVR-ready retrieval environment derived from Nan-Do/code-search-net-python. Author: Aman Priyanshu What Is This A 100k-row retrieval QA dataset where each row contains a question, ground-truth chunks, and pre-mined distractor chunks (random + semantically similar). Designed for training and evaluating retrieval agents in an RLVR (Reinforcement Learning with Verifiable Rewards) setup — the agent searches through distractors… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/RLVR-Env-Retrieval-Source-code-search-net-python.texttext-retrieval100K<n<1M0 likes23 downloads7mo agoHugging Face16dwb2023 /gdelt-rag-sources GDELT RAG Source Documents Dataset Description This dataset contains source documents extracted from the research paper "Talking to GDELT Through Knowledge Graphs" (arXiv:2503.07584v3). The documents are used as the knowledge base for a Retrieval-Augmented Generation (RAG) system focused on GDELT (Global Database of Events, Language, and Tone) analysis. Dataset Summary Total Documents: 38 pages Source: Research paper on GDELT Knowledge Graphs Format: PDF pages… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-sources.texttext-retrievaln<1K0 likes19 downloads7mo agoHugging Face17antonixe /river_source Dataset Card for Dataset Name Dataset Summary This dataset card aims to be a base template for new datasets. It has been generated using this raw template. Supported Tasks and Leaderboards [More Information Needed] Languages [More Information Needed] Dataset Structure Data Instances [More Information Needed] Data Fields [More Information Needed] Data Splits [More Information Needed] Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/antonixe/river_source.textquestion-answeringn<1K0 likes16 downloads3y agoHugging Face18dakheel /hudanet-sourcesgated HUDA-Net Sources Hajj and Umrah Jurisprudence Corpus مصادر شبكة هدى لفقه الحج والعمرة HUDA-Net Sources is the private, version-controlled source repository used to build the HUDA-Net bilingual question-answering, semantic retrieval, and text-ranking system for Hajj and Umrah jurisprudence. This repository preserves the original book datasets, cleaned source files, provenance metadata, and integrity manifests required to reproduce and audit the HUDA-Net corpus.… See the full description on the dataset page: https://huggingface.co/datasets/dakheel/hudanet-sources.textquestion-answering1K<n<10K0 likes16 downloads2mo agoHugging Face19dwb2023 /gdelt-rag-sources-v2 GDELT RAG Source Documents Dataset Description This dataset contains source documents extracted from the research paper "Talking to GDELT Through Knowledge Graphs" (arXiv:2503.07584v3). The documents are used as the knowledge base for a Retrieval-Augmented Generation (RAG) system focused on GDELT (Global Database of Events, Language, and Tone) analysis. Dataset Summary Total Documents: 38 pages Source: Research paper on GDELT Knowledge Graphs Format: PDF pages… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-sources-v2.texttext-retrievaln<1K0 likes14 downloads11mo agoHugging Face20dwb2023 /gdelt-rag-sources-v4 GDELT RAG Source Documents Dataset Description This dataset contains source documents extracted from the research paper "Talking to GDELT Through Knowledge Graphs" (arXiv:2503.07584v3). The documents are used as the knowledge base for a Retrieval-Augmented Generation (RAG) system focused on GDELT (Global Database of Events, Language, and Tone) analysis. Dataset Summary Total Documents: 38 pages Source: Research paper on GDELT Knowledge Graphs Format: PDF pages… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-sources-v4.texttext-retrievaln<1K0 likes11 downloads11mo agoHugging Face21dwb2023 /gdelt-rag-sources-v3 GDELT RAG Source Documents Dataset Description This dataset contains source documents extracted from the research paper "Talking to GDELT Through Knowledge Graphs" (arXiv:2503.07584v3). The documents are used as the knowledge base for a Retrieval-Augmented Generation (RAG) system focused on GDELT (Global Database of Events, Language, and Tone) analysis. Dataset Summary Total Documents: 38 pages Source: Research paper on GDELT Knowledge Graphs Format: PDF pages… See the full description on the dataset page: https://huggingface.co/datasets/dwb2023/gdelt-rag-sources-v3.texttext-retrievaln<1K0 likes9 downloads11mo agoHugging Face22Tigle /tigle-source-codegated TIGLE The interface is built as a prototype based on Dzogchen, Atiyoga teachings available in English and sourced, compiled by a practitioner exploring how Dharma language and current global AI could intersect. The architecture, the pipeline works. The answers are useful for orientation — learning key terms, lineages, main practices, understanding the view. This is why it is accessible as repository rather than a product: Digital Bardo - is the current state of samsara.… See the full description on the dataset page: https://huggingface.co/datasets/Tigle/tigle-source-code.texttext-generationn<1K0 likes8 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.