datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Multi-Source-Video-Captioning
Multi-source Video Captioning (MSVC) Dataset Card
Dataset details
Dataset type:
MSVC is a set of collected video captioning data. It is constructed to ensure a robust and thorough evaluation of Video-LLMs' video-captioning capabilities.
Dataset detail:
MSVC is introduced to address limitations in existing video caption benchmarks, MSVC samples a total of 1,500 videos with human-annotated captions from MSVD, MSRVTT, and VATEX, ensuring diverse scenarios and domains.… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/Multi-Source-Video-Captioning.multisource-membench
Multi-Source Memory Benchmark
Status — public release.
A diagnostic testbed for selective question-answering (ANSWER / SKIP) over conflicting multi-source personal memory.
Each persona has five evidence streams projected from a single latent event table with known, controlled per-source distortions (bias direction, dropout rate, granularity), allowing methods to be measured against the latent ground truth rather than against any single source.
The benchmark accompanies the… See the full description on the dataset page: https://huggingface.co/datasets/ytc1997/multisource-membench.multisource-memory-benchmark
Multi-Source Memory Benchmark
Status — anonymous artefact for double-blind review (NeurIPS 2026 Evaluations & Datasets Track).
Author identities, organisations, and funders are intentionally withheld until the review period concludes.
A diagnostic testbed for selective question-answering (ANSWER / SKIP) over conflicting multi-source personal memory.
Each persona has five evidence streams projected from a single latent event table with known, controlled per-source distortions… See the full description on the dataset page: https://huggingface.co/datasets/anon-neuripsed26/multisource-memory-benchmark.lemonseed-multisource-cogen
lemonseed-multisource-cogen
LemonSeed — multi-source teacher-co-gen training.
Contents
intelligent_multisource_train.jsonl (464 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
LLM-teacher co-generated instruction/chat data for LemonSeed fine-tuning.
lemonseed-multisource-cogen-corrective
lemonseed-multisource-cogen-corrective
LemonSeed — multi-source corrective teacher-co-gen training (v3).
Contents
intelligent_multisource_corrective_train_v3.jsonl (522 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
LLM-teacher co-generated instruction/chat data for LemonSeed fine-tuning.
lemonseed-multisource-cogen-diversified
lemonseed-multisource-cogen-diversified
LemonSeed — multi-source diversified teacher-co-gen training (v2).
Contents
intelligent_multisource_diversified_train_v2.jsonl (464 rows)
Format
JSON Lines (.jsonl), one example per line.
Provenance
LLM-teacher co-generated instruction/chat data for LemonSeed fine-tuning.
artificial-intelligence-multi-source-datasetheterogeneous-multisource-kg-benchmark
Heterogeneous Multi-Source Knowledge Graph Benchmark
A benchmark for evaluating multi-source knowledge graph construction and cross-source question answering systems. The benchmark comprises 50 questions designed to assess systems' ability to reason across heterogeneous data sources (structured SQL, semi-structured JSON, and unstructured text).
Dataset Overview
Property
Value
Questions
50
Question Categories
Source Attribution, Entity Integration, Conflict… See the full description on the dataset page: https://huggingface.co/datasets/Cool-EdwardH/heterogeneous-multisource-kg-benchmark.biomedical-multi-source-finetunemultisource
