datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CI-VID
📄 CI-VID: A Coherent Interleaved Text-Video Dataset
CI-VID is a large-scale dataset designed to advance coherent multi-clip video generation. Unlike traditional text-to-video (T2V) datasets with isolated clip-caption pairs, CI-VID supports text-and-video-to-video (TV2V) generation by providing over 340,000 interleaved sequences of video clips and rich captions. It enables models to learn both intra-clip content and inter-clip transitions, fostering story-driven generation with… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/CI-VID.circleci-test-resultscitations
Common Crawl Citations Overview
This dataset contains citations referencing Common Crawl Foundation and its datasets, pulled from Google Scholar.
Please note that these citations are not curated, so they will include some false positives. An annotated subset of these citations with additional fields can be found at cc-citations.
Dr-CiK
Dr-CiK: A Testbed for Foresight-Driven Agents
Dr-CiK is a benchmark for evaluating whether agents can retrieve
forecasting-relevant context from a noisy document corpus, filter out
distractors, distill the retrieved context into forecast-useful evidence, and
produce forecasts grounded in that evidence.
Real-world time-series forecasting often depends not only on historical
observations but also on external context that must be actively discovered
from heterogeneous, noisy… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow/Dr-CiK.AusLaw-Citation-BenchmarkThis is the dataset proposed in the paper: Methods for Legal Citation Prediction in the Age of LLMs: An Australian Law Case Study.
conll2003-ciparsed_datacia-world-factbook-snapshotssquad-cicivil-code-phil
Civilex — Philippine Legal RAG & SFT Dataset
Retrieval corpus and supervised fine-tuning (SFT) data for a retrieval-augmented generation (RAG) pipeline over Philippine law: the Civil Code (Republic Act No. 386) and Supreme Court jurisprudence. Produced by the civilex-thesis research pipeline.
Contents: 11k+ Supreme Court jurisprudence cases spanning 1949–2025, and 2,270 articles from the Civil Code (Republic Act No. 386).
Dataset structure
.
├── README.md
├──… See the full description on the dataset page: https://huggingface.co/datasets/renzzyyy1028/civil-code-phil.cinepile_10knetryx-new-york-city-13km
Nyc-Core-Usethis 13km
Pre-computed MegaLoc index for Netryx Drishti geolocation.
Coverage
Center: 40.713200, -74.002500
Radius: 13.0 km
Panoramas: 663,084
Index entries: 2,652,336
Descriptor model: MegaLoc
Descriptor dim: 1024 (PCA from 8448)
Usage
from netryx_hub import NetryxHub
hub = NetryxHub()
hub.download("nyc-core-usethis-13km", output_dir="./netryx_data/index")
# Now open Netryx and search!
Or download manually and use Import Index in… See the full description on the dataset page: https://huggingface.co/datasets/samsepiol4/netryx-new-york-city-13km.ObjaversePlusPlus
Objaverse++: Curated 3D Object Dataset with Quality Annotations
Paper
Code
Chendi Lin,
Heshan Liu,
Qunshu Lin,
Zachary Bright,
Shitao Tang,
Yihui He,
Minghao Liu,
Ling Zhu,
Cindy Le
We cleaned the Objaverse dataset so you don't have to. In this work, we meticulously curated a collection of Objaverse objects and developed an effective classifier capable of scoring the entire Objaverse. Our extensive annotation system considers geometric structure and texture… See the full description on the dataset page: https://huggingface.co/datasets/cindyxl/ObjaversePlusPlus.CineBoard3D-plus
🎬 CineBoard3D++: Dynamic 3D Story World Dataset
📊 Dataset Summary
CineBoard3D++ is a collection of editable, movie-inspired 3D story worlds built with StoryBlender for narrative-grounded camera planning and world visual attention. It brings together story scripts, animated characters, scene geometry, and shot-level configurations in native Blender projects.
The benchmark covers 50 stories, 457 scenes, 1,585 shots, and 3,197 3D assets (836 plot-related and 2,361… See the full description on the dataset page: https://huggingface.co/datasets/EngineeringAI-LAB/CineBoard3D-plus.wildtrace
WildTrace strict481
WildTrace is a source-internal long-context multi-hop reasoning benchmark
built from natural evidence trails. Unlike reverse-synthetic QA, its tasks are
mined in situ from long source documents before questions are written. The
strict481 release contains 481 locked tasks over 214 public long-form sources,
with full-document, evidence-withheld evaluation. The model under test receives
only the source document and the public question; evidence spans, clue… See the full description on the dataset page: https://huggingface.co/datasets/CinderD/wildtrace.floorplans-cityscapes
Dataset Summary
This is a curated collection of floorplan images sourced from across the internet. It is intended for research in architectural AI, layout generation, and urban scene understanding.
Data format: Image files with associated integer labels.
Sources: Publicly available images from various web sources (This dataset is one unified collections).
Purpose: Educational and research use.
Dataset Structure
The dataset follows the standard Hugging Face Image… See the full description on the dataset page: https://huggingface.co/datasets/wheres-my-python/floorplans-cityscapes.L-CiteEval
L-CITEEVAL: DO LONG-CONTEXT MODELS TRULY LEVERAGE CONTEXT FOR RESPONDING?
Paper Github Zhihu
Benchmark Quickview
L-CiteEval is a multi-task long-context understanding with citation benchmark, covering 5 task categories, including single-document question answering, multi-document question answering, summarization, dialogue understanding, and synthetic tasks, encompassing 11 different long-context tasks. The context lengths for these tasks range from 8K to 48K.… See the full description on the dataset page: https://huggingface.co/datasets/Jonaszky123/L-CiteEval.CiteVQA
CiteVQA
English | 简体中文
CiteVQA is a document visual question answering benchmark for faithful evidence attribution. Unlike conventional DocVQA datasets that only score the final answer, CiteVQA requires a model to answer a question with evidence grounded in the source document at the element level. The benchmark is designed to evaluate whether a system can not only answer correctly, but also cite the right supporting region in long, real-world PDFs.
The dataset contains 1,897… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/CiteVQA.CityCube-Benchworld-cities-geoDataset containing city, country, region, and continents alongside their longitude and latitude co-ordinates. Cartesian coordinates are provided in x, y, z features.
s2orc-citation-pairs-translated-nlThis is a Dutch version of the S2ORC: The Semantic Scholar Open Research Corpus. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation.
cinematic-world-stills
Council of AI — cinematic stills
Cinematic stills produced for Council of AI surfaces. metadata.jsonl gives the
Hub image viewer a caption per file. These are illustrations — they carry no measurement and back no slot.
The live board is the authority
GET https://councilof.ai/api/gspc — quote totals.public_count. This Hub card is a printer of that GET, never a second
engine. If the fetch fails the honest answer is UNCHECKABLE — never a fabricated 0.000.
Status… See the full description on the dataset page: https://huggingface.co/datasets/csoai/cinematic-world-stills.llm-cipher-reasoning
llm-cipher-reasoning — data, eval results and full research ledger
Everything except the weights from a research run asking: can an LLM be trained to reason in a
more compact "language" than English, and does that actually save tokens?
Two linked lines of work on Qwen/Qwen3-4B-Instruct-2507:
Cipher invention / cross-model communication — cold-decoding tests, negotiated cipher
collusion between model pairs, a cipher-hardening arms race, and GEPA prompt optimization to get
a… See the full description on the dataset page: https://huggingface.co/datasets/AlexWortega/llm-cipher-reasoning.cicero
Dataset Card for CICERO
Description
Homepage: https://declare-lab.net/CICERO/
Repository: https://github.com/declare-lab/CICERO
Paper: https://aclanthology.org/2022.acl-long.344/
arXiv: https://arxiv.org/abs/2203.13926
Summary
CICERO is a new dataset for dialogue reasoning with contextualized commonsense inference. It contains 53K inferences for five commonsense dimensions – cause, subsequent event, prerequisite, motivation, and emotional reaction… See the full description on the dataset page: https://huggingface.co/datasets/declare-lab/cicero.false-citation-bench
False Citation Bench
False Citation Bench is a compact evaluation and inspection dataset for false or misleading case citations in legal documents. It contains 26 source documents, their PDFs, and manually reviewed citation annotations grounded in the local text extraction.
Dataset contents
The repository has one matching document in each directory:
documents_txt/{index}__{case-name}__{filing}.txt
documents_pdf/{index}__{case-name}__{filing}.pdf… See the full description on the dataset page: https://huggingface.co/datasets/gt-csse/false-citation-bench.Nemotron-RL-Instruction-Following-Citation-Formatting-v1
Dataset Description:
Teaches the model to cite specific document parts using reference markers like [ref:1], ref:3, etc. Supports single-reference, multi-reference, and inline citations.
This dataset is ready for commercial/non-commercial uses.
Dataset Owner(s):
NVIDIA Corporation
Dataset Creation Date:
Created on: April 10, 2026
Last Modified on: April 10, 2026
Version:
Nemotron-RL-Instruction-Following-CitationFormatting-v1… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-Citation-Formatting-v1.CipherBank
CipherBank Benchmark
Benchmark description
CipherBank, a comprehensive benchmark designed to evaluate the reasoning capabilities of LLMs in cryptographic decryption tasks.
CipherBank comprises 2,358 meticulously crafted problems, covering 262 unique plaintexts across 5 domains and 14 subdomains, with a focus on privacy-sensitive and real-world scenarios that necessitate encryption. From a cryptographic perspective, CipherBank incorporates 3 major categories of encryption… See the full description on the dataset page: https://huggingface.co/datasets/yu0226/CipherBank.CIMD
CIMD
[[中文]] | [[English]]
CSGHub Dataset Page | Hugging Face | OpenCSG Community
中文说明
数据集概述
CIMD 是一个面向文档智能任务的跨来源、多语言 JSONL 语料库。当前公开快照包含 111,308 条解析记录,覆盖制度参考、学术与长文档资料、机构分析、企业运营、公共讨论和市场相关材料等来源家族。每条记录都把正文与来源类型、语言、时间、关键词、授权标签和来源字段放在同一个结构里,用户拿到数据后可以直接做检索、抽样、审计和数据治理。
公开数据已转换为统一字段,并按来源家族拆分为可单独加载的子集;它不是原始文件夹的简单打包。用户可以只读取制度参考、学术长文档或公共讨论记录,也可以合并多个子集构建检索库、抽取训练候选样本、构造评测样本池,并按来源、语言和时间字段继续筛选。
CIMD 和通用网页语料的差别在于记录级元数据。它不只提供可索引文本,还提供… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/CIMD.Legal_Phantom_Citation
Legal Phantom Citation Benchmark
LePhamtomCite is a benchmarking dataset for evaluating AI systems on legal citation hallucination detection.
Dataset description
Legal citation hallucinations (fabricated or misrepresented case citations in court filings) are a growing problem as attorneys, judges, and pro se litigants increasingly rely on LLMs to draft legal documents. LePhamtomCite provides a structured benchmark for evaluating automated citation verification… See the full description on the dataset page: https://huggingface.co/datasets/ai-law-society-lab/Legal_Phantom_Citation.civil-engineering-gemma-data
