CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01BAAI /CI-VID 📄 CI-VID: A Coherent Interleaved Text-Video Dataset CI-VID is a large-scale dataset designed to advance coherent multi-clip video generation. Unlike traditional text-to-video (T2V) datasets with isolated clip-caption pairs, CI-VID supports text-and-video-to-video (TV2V) generation by providing over 340,000 interleaved sequences of video clips and rich captions. It enables models to learn both intra-clip content and inter-clip transitions, fostering story-driven generation with… See the full description on the dataset page: https://huggingface.co/datasets/BAAI/CI-VID.text100K<n<1M7 likes4k downloads10mo agoHugging Face02transformers-community /circleci-test-resultstextn<1K4 likes2k downloads3mo agoHugging Face03commoncrawl /citations Common Crawl Citations Overview This dataset contains citations referencing Common Crawl Foundation and its datasets, pulled from Google Scholar. Please note that these citations are not curated, so they will include some false positives. An annotated subset of these citations with additional fields can be found at cc-citations. text1K<n<10K5 likes1.2k downloads6mo agoHugging Face04ServiceNow /Dr-CiK Dr-CiK: A Testbed for Foresight-Driven Agents Dr-CiK is a benchmark for evaluating whether agents can retrieve forecasting-relevant context from a noisy document corpus, filter out distractors, distill the retrieved context into forecast-useful evidence, and produce forecasts grounded in that evidence. Real-world time-series forecasting often depends not only on historical observations but also on external context that must be actively discovered from heterogeneous, noisy… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow/Dr-CiK.tabulartime-series-forecasting10K<n<100K3 likes1k downloads3mo agoHugging Face05auslawbench /AusLaw-Citation-BenchmarkThis is the dataset proposed in the paper: Methods for Legal Citation Prediction in the Age of LLMs: An Australian Law Case Study. text10K<n<100K3 likes887 downloads1y agoHugging Face06evaluate /conll2003-citextn<1K0 likes802 downloads4y agoHugging Face07cia-tools /parsed_datatext1K<n<10K0 likes682 downloads1y agoHugging Face08olm /cia-world-factbook-snapshotstext1K<n<10K1 likes624 downloads4y agoHugging Face09evaluate /squad-citextn<1K0 likes606 downloads4y agoHugging Face10renzzyyy1028 /civil-code-phil Civilex — Philippine Legal RAG & SFT Dataset Retrieval corpus and supervised fine-tuning (SFT) data for a retrieval-augmented generation (RAG) pipeline over Philippine law: the Civil Code (Republic Act No. 386) and Supreme Court jurisprudence. Produced by the civilex-thesis research pipeline. Contents: 11k+ Supreme Court jurisprudence cases spanning 1949–2025, and 2,270 articles from the Civil Code (Republic Act No. 386). Dataset structure . ├── README.md ├──… See the full description on the dataset page: https://huggingface.co/datasets/renzzyyy1028/civil-code-phil.textquestion-answering100K<n<1M1 likes550 downloads1d agoHugging Face11KerryMe /cinepile_10ktabular10K<n<100K0 likes482 downloads3mo agoHugging Face12samsepiol4 /netryx-new-york-city-13km Nyc-Core-Usethis 13km Pre-computed MegaLoc index for Netryx Drishti geolocation. Coverage Center: 40.713200, -74.002500 Radius: 13.0 km Panoramas: 663,084 Index entries: 2,652,336 Descriptor model: MegaLoc Descriptor dim: 1024 (PCA from 8448) Usage from netryx_hub import NetryxHub hub = NetryxHub() hub.download("nyc-core-usethis-13km", output_dir="./netryx_data/index") # Now open Netryx and search! Or download manually and use Import Index in… See the full description on the dataset page: https://huggingface.co/datasets/samsepiol4/netryx-new-york-city-13km.tabularn<1K0 likes460 downloads2mo agoHugging Face13cindyxl /ObjaversePlusPlus Objaverse++: Curated 3D Object Dataset with Quality Annotations Paper Code Chendi Lin, Heshan Liu, Qunshu Lin, Zachary Bright, Shitao Tang, Yihui He, Minghao Liu, Ling Zhu, Cindy Le We cleaned the Objaverse dataset so you don't have to. In this work, we meticulously curated a collection of Objaverse objects and developed an effective classifier capable of scoring the entire Objaverse. Our extensive annotation system considers geometric structure and texture… See the full description on the dataset page: https://huggingface.co/datasets/cindyxl/ObjaversePlusPlus.texttext-to-3d100K<n<1M23 likes440 downloads10mo agoHugging Face14EngineeringAI-LAB /CineBoard3D-plus 🎬 CineBoard3D++: Dynamic 3D Story World Dataset 📊 Dataset Summary CineBoard3D++ is a collection of editable, movie-inspired 3D story worlds built with StoryBlender for narrative-grounded camera planning and world visual attention. It brings together story scripts, animated characters, scene geometry, and shot-level configurations in native Blender projects. The benchmark covers 50 stories, 457 scenes, 1,585 shots, and 3,197 3D assets (836 plot-related and 2,361… See the full description on the dataset page: https://huggingface.co/datasets/EngineeringAI-LAB/CineBoard3D-plus.3dn<1K0 likes374 downloads14d agoHugging Face15CinderD /wildtrace WildTrace strict481 WildTrace is a source-internal long-context multi-hop reasoning benchmark built from natural evidence trails. Unlike reverse-synthetic QA, its tasks are mined in situ from long source documents before questions are written. The strict481 release contains 481 locked tasks over 214 public long-form sources, with full-document, evidence-withheld evaluation. The model under test receives only the source document and the public question; evidence spans, clue… See the full description on the dataset page: https://huggingface.co/datasets/CinderD/wildtrace.textquestion-answeringn<1K1 likes322 downloads1mo agoHugging Face16wheres-my-python /floorplans-cityscapes Dataset Summary This is a curated collection of floorplan images sourced from across the internet. It is intended for research in architectural AI, layout generation, and urban scene understanding. Data format: Image files with associated integer labels. Sources: Publicly available images from various web sources (This dataset is one unified collections). Purpose: Educational and research use. Dataset Structure The dataset follows the standard Hugging Face Image… See the full description on the dataset page: https://huggingface.co/datasets/wheres-my-python/floorplans-cityscapes.imagefeature-extraction1K<n<10K1 likes304 downloads6mo agoHugging Face17Jonaszky123 /L-CiteEval L-CITEEVAL: DO LONG-CONTEXT MODELS TRULY LEVERAGE CONTEXT FOR RESPONDING? Paper   Github   Zhihu Benchmark Quickview L-CiteEval is a multi-task long-context understanding with citation benchmark, covering 5 task categories, including single-document question answering, multi-document question answering, summarization, dialogue understanding, and synthetic tasks, encompassing 11 different long-context tasks. The context lengths for these tasks range from 8K to 48K.… See the full description on the dataset page: https://huggingface.co/datasets/Jonaszky123/L-CiteEval.tabularquestion-answering1K<n<10K3 likes297 downloads2y agoHugging Face18opendatalab /CiteVQA CiteVQA English | 简体中文 CiteVQA is a document visual question answering benchmark for faithful evidence attribution. Unlike conventional DocVQA datasets that only score the final answer, CiteVQA requires a model to answer a question with evidence grounded in the source document at the element level. The benchmark is designed to evaluate whether a system can not only answer correctly, but also cite the right supporting region in long, real-world PDFs. The dataset contains 1,897… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/CiteVQA.textvisual-question-answering1K<n<10K11 likes295 downloads4mo agoHugging Face19BigBro23 /CityCube-Benchimagequestion-answering1K<n<10K3 likes285 downloads9mo agoHugging Face20jamescalam /world-cities-geoDataset containing city, country, region, and continents alongside their longitude and latitude co-ordinates. Cartesian coordinates are provided in x, y, z features. tabular1K<n<10K15 likes260 downloads4y agoHugging Face21NetherlandsForensicInstitute /s2orc-citation-pairs-translated-nlThis is a Dutch version of the S2ORC: The Semantic Scholar Open Research Corpus. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation. textsentence-similarity10M<n<100M0 likes260 downloads2y agoHugging Face22csoai /cinematic-world-stills Council of AI — cinematic stills Cinematic stills produced for Council of AI surfaces. metadata.jsonl gives the Hub image viewer a caption per file. These are illustrations — they carry no measurement and back no slot. The live board is the authority GET https://councilof.ai/api/gspc — quote totals.public_count. This Hub card is a printer of that GET, never a second engine. If the fetch fails the honest answer is UNCHECKABLE — never a fabricated 0.000. Status… See the full description on the dataset page: https://huggingface.co/datasets/csoai/cinematic-world-stills.imageothern<1K0 likes237 downloads10d agoHugging Face23AlexWortega /llm-cipher-reasoning llm-cipher-reasoning — data, eval results and full research ledger Everything except the weights from a research run asking: can an LLM be trained to reason in a more compact "language" than English, and does that actually save tokens? Two linked lines of work on Qwen/Qwen3-4B-Instruct-2507: Cipher invention / cross-model communication — cold-decoding tests, negotiated cipher collusion between model pairs, a cipher-hardening arms race, and GEPA prompt optimization to get a… See the full description on the dataset page: https://huggingface.co/datasets/AlexWortega/llm-cipher-reasoning.texttext-generation10K<n<100K0 likes221 downloads22d agoHugging Face24declare-lab /cicero Dataset Card for CICERO Description Homepage: https://declare-lab.net/CICERO/ Repository: https://github.com/declare-lab/CICERO Paper: https://aclanthology.org/2022.acl-long.344/ arXiv: https://arxiv.org/abs/2203.13926 Summary CICERO is a new dataset for dialogue reasoning with contextualized commonsense inference. It contains 53K inferences for five commonsense dimensions – cause, subsequent event, prerequisite, motivation, and emotional reaction… See the full description on the dataset page: https://huggingface.co/datasets/declare-lab/cicero.text10K<n<100K1 likes211 downloads4y agoHugging Face25gt-csse /false-citation-bench False Citation Bench False Citation Bench is a compact evaluation and inspection dataset for false or misleading case citations in legal documents. It contains 26 source documents, their PDFs, and manually reviewed citation annotations grounded in the local text extraction. Dataset contents The repository has one matching document in each directory: documents_txt/{index}__{case-name}__{filing}.txt documents_pdf/{index}__{case-name}__{filing}.pdf… See the full description on the dataset page: https://huggingface.co/datasets/gt-csse/false-citation-bench.documentn<1K2 likes194 downloads1mo agoHugging Face26nvidia /Nemotron-RL-Instruction-Following-Citation-Formatting-v1 Dataset Description: Teaches the model to cite specific document parts using reference markers like [ref:1], ref:3, etc. Supports single-reference, multi-reference, and inline citations. This dataset is ready for commercial/non-commercial uses. Dataset Owner(s): NVIDIA Corporation Dataset Creation Date: Created on: April 10, 2026 Last Modified on: April 10, 2026 Version: Nemotron-RL-Instruction-Following-CitationFormatting-v1… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Instruction-Following-Citation-Formatting-v1.texttext-generation1K<n<10K2 likes179 downloads4mo agoHugging Face27yu0226 /CipherBank CipherBank Benchmark Benchmark description CipherBank, a comprehensive benchmark designed to evaluate the reasoning capabilities of LLMs in cryptographic decryption tasks. CipherBank comprises 2,358 meticulously crafted problems, covering 262 unique plaintexts across 5 domains and 14 subdomains, with a focus on privacy-sensitive and real-world scenarios that necessitate encryption. From a cryptographic perspective, CipherBank incorporates 3 major categories of encryption… See the full description on the dataset page: https://huggingface.co/datasets/yu0226/CipherBank.textquestion-answering1K<n<10K3 likes177 downloads1y agoHugging Face28opencsg /CIMD CIMD [[中文]] | [[English]] CSGHub Dataset Page | Hugging Face | OpenCSG Community 中文说明 数据集概述 CIMD 是一个面向文档智能任务的跨来源、多语言 JSONL 语料库。当前公开快照包含 111,308 条解析记录,覆盖制度参考、学术与长文档资料、机构分析、企业运营、公共讨论和市场相关材料等来源家族。每条记录都把正文与来源类型、语言、时间、关键词、授权标签和来源字段放在同一个结构里,用户拿到数据后可以直接做检索、抽样、审计和数据治理。 公开数据已转换为统一字段,并按来源家族拆分为可单独加载的子集;它不是原始文件夹的简单打包。用户可以只读取制度参考、学术长文档或公共讨论记录,也可以合并多个子集构建检索库、抽取训练候选样本、构造评测样本池,并按来源、语言和时间字段继续筛选。 CIMD 和通用网页语料的差别在于记录级元数据。它不只提供可索引文本,还提供… See the full description on the dataset page: https://huggingface.co/datasets/opencsg/CIMD.text10K<n<100K7 likes176 downloads4mo agoHugging Face29ai-law-society-lab /Legal_Phantom_Citation Legal Phantom Citation Benchmark LePhamtomCite is a benchmarking dataset for evaluating AI systems on legal citation hallucination detection. Dataset description Legal citation hallucinations (fabricated or misrepresented case citations in court filings) are a growing problem as attorneys, judges, and pro se litigants increasingly rely on LLMs to draft legal documents. LePhamtomCite provides a structured benchmark for evaluating automated citation verification… See the full description on the dataset page: https://huggingface.co/datasets/ai-law-society-lab/Legal_Phantom_Citation.text1K<n<10K1 likes153 downloads3mo agoHugging Face30karanjaWakaba /civil-engineering-gemma-datatext10K<n<100K0 likes141 downloads8mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.