CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01hf-internal-testing /transformers_circleci_workflow_runs5 likes271k downloads2y agoHugging Face02ruggsea /infini-news-corpus INFINI-NEWS Corpus 🔎 Search this corpus online: query it with sub-second full-text search and n-gram counts — in the browser or via a public, keyless REST API, no download required — at infini-news.uni-graz.at (API reference). A multilingual news corpus extracted from Common Crawl CC-News WARC files. One row per article, with body text extracted via trafilatura, WARC provenance, and derived metadata (publish date, language, topic, byte hashes) in a single flat schema. Covers… See the full description on the dataset page: https://huggingface.co/datasets/ruggsea/infini-news-corpus.tabulartext-generation1B<n<10B39 likes45k downloads5d agoHugging Face03RukawaY /gs_scenes A High-Fidelity Navigation Simulator with Dynamic Gaussian SplattingECCV 2026 Ziyuan Xia • Jingyi Xu • Chong Cui • Yuanhong Yu • Jiazhao Zhang • Qingsong Yan • Tao Ni Junbo Chen • Xiaowei Zhou • Hujun Bao • Ruizhen Hu • Sida Peng 🤗 About This Dataset This is the official GS dataset for Habitat-GS, a high-fidelity embodied navigation simulator built on 3D Gaussian Splatting and dynamic gaussian avatars. The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/RukawaY/gs_scenes.roboticsn<1K7 likes28k downloads4d agoHugging Face04rubend18 /ChatGPT-Jailbreak-Prompts Dataset Card for Dataset Name Name ChatGPT Jailbreak Prompts Dataset Summary ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT. Languages [English] tabularquestion-answeringn<1K274 likes27k downloads3y agoHugging Face05RuoliuYang /ulvr_subset ULVR stage-0 subsets (latent + source) Curated, nested subsets of the Unified Visual Latent Reasoning (ULVR) stage-0 training data. Each subset folder is self-contained and ships both: latent/ — pre-computed teacher latents, identical schema to RuoliuYang/step0-all source/ — the matching source samples (images + question/answer + messages), identical schema to RuoliuYang/ULVR_v2_clean Latents and source rows are joinable by sample_id (within a category). Folder… See the full description on the dataset page: https://huggingface.co/datasets/RuoliuYang/ulvr_subset.textvisual-question-answering100K<n<1M0 likes26k downloads3mo agoHugging Face06jake-mercor /cherrl-runs3 likes24k downloads1mo agoHugging Face07rugds /ditec-wdn-- Dataset Card for DiTEC-WDN Dataset Summary DiTEC-WDN Dataset consists of 36 Water Distribution Networks (WDNs). Each network has unique 1,000 scenarios with distinct characteristics. Scenario represents a timeseries of directed shared-topology graphs, referred to as states or snapshots. In terms of graph-ml, it can be seen as a spatiotemporal graph where nodes and edges are multivariate time series. A node can represent a reservoir, junction, or tank, while an edge… See the full description on the dataset page: https://huggingface.co/datasets/rugds/ditec-wdn.tabulargraph-ml1B<n<10B4 likes24k downloads10mo agoHugging Face08dsaddsaf /sn80-data-run10 likes19k downloads2mo agoHugging Face09mackelab /benchmarking_sbi_runs Benchmarking SBI Runs This dataset contains the raw, per-run results underlying the manuscript "Benchmarking Simulation-Based Inference" (Lueckmann, Boelts, Greenberg, Goncalves & Macke, AISTATS 2021). It is a direct migration of the Git LFS data from mackelab/benchmarking_sbi_runs on GitHub. For compiled, ready-to-use dataframes built from these raw results (and the code that produced them), see the companion repository:… See the full description on the dataset page: https://huggingface.co/datasets/mackelab/benchmarking_sbi_runs.100K<n<1M0 likes19k downloads2mo agoHugging Face10RUC-NLPIR /FlashRAG_datasets ⚡FlashRAG: A Python Toolkit for Efficient RAG Research FlashRAG is a Python toolkit for the reproduction and development of Retrieval Augmented Generation (RAG) research. Our toolkit includes 36 pre-processed benchmark RAG datasets and 16 state-of-the-art RAG algorithms. With FlashRAG and provided resources, you can effortlessly reproduce existing SOTA works in the RAG domain or implement your custom RAG processes and components. For more information, please view our GitHub repo… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/FlashRAG_datasets.textquestion-answering1M<n<10M94 likes18k downloads1y agoHugging Face11NCSpeech /YO-CPT-ru YO-CPT-ru YouTube-Oriented dataset for Continual Pre-Training (Russian). A large, heavily quality-filtered corpus of Russian speech mined from YouTube (via YODAS2) and processed into clean, single-speaker, TTS-grade utterances. Every utterance ships with an ensemble-verified transcription, a punctuated/denormalized and stress-marked text variant, word-level forced alignment, within- and cross-video speaker identities, an audio-quality (MOS) score, and a speaker persona built… See the full description on the dataset page: https://huggingface.co/datasets/NCSpeech/YO-CPT-ru.audiotext-to-speech1M<n<10M15 likes13k downloads2mo agoHugging Face12andlyu /Public-YAM-runs Public-YAM-runs Physical bimanual YAM episodes recorded by the BluPe operator station. Each run adds an episode to this repository. Failed, interrupted, stopped and timed-out runs are retained and labeled; these are not all successful demonstrations. A model saying done is not independently verified task success. Loading from datasets import load_dataset runs = load_dataset("andlyu/Public-YAM-runs", split="train") usable = runs.filter(lambda row:… See the full description on the dataset page: https://huggingface.co/datasets/andlyu/Public-YAM-runs.image100K<n<1M2 likes10k downloads1d agoHugging Face13RuoliuYang /ULVR_v2_clean ULVR_v2_clean Universal Latent Visual Reasoning training data, cleaned. 8 categories (subsets); each has train + validation splits. Every sample: input image + question -> assistant produces <abs_vis_token> + intermediate visual step(s) + \boxed{answer}. subset train validation text_cot 333,911 3,533 bbox_highlight 229,237 2,558 bbox_crop 229,237 2,558 depth 40,000 25 edge 40,000 14 segmentation 40,000 326 helper_interleaved 340,210 3,544 scene_graph 40… See the full description on the dataset page: https://huggingface.co/datasets/RuoliuYang/ULVR_v2_clean.imagevisual-question-answering1M<n<10M1 likes9.4k downloads3mo agoHugging Face14ruggsea /infini-news-index INFINI-NEWS FM-Index 🔎 Live search API: these FM-indexes power a public search service — full-text search, n-gram counts, and document retrieval in the browser or via a keyless REST API, without building the index yourself — at infini-news.uni-graz.at (API reference). Pre-built FM-indexes (Burrows–Wheeler Transform + suffix array, built with infini-gram-mini, Liu et al. 2025) over the ruggsea/infini-news-corpus parquets. Enables exact, byte-level substring count and document… See the full description on the dataset page: https://huggingface.co/datasets/ruggsea/infini-news-index.text-retrieval2 likes8.7k downloads5d agoHugging Face15simonjegou /rulertext10K<n<100K2 likes8.5k downloads2y agoHugging Face16oguzhanmeteozturk /flame-runs0 likes8.3k downloads3h agoHugging Face17cl-nagoya /ruri-dataset-v2-ptWIP: 正式公開準備中 各データセットのライセンスは元データセットに従います。 text100M<n<1B5 likes8.3k downloads2y agoHugging Face18ruggsea /social-sim-bench-genstext1K<n<10K0 likes8k downloads3mo agoHugging Face19RUC-AIBOX /Evo-BenchEvo-Bench: Can Language Models Improve Agent Harness? A benchmark for measuring the intrinsic harness-evolving capability of language models. Overview of the Evo-Bench evaluation pipeline. ✨ Highlights 608 harness-sensitive tasks from five established benchmarks, spanning Search, Office, and General agent domains with disjoint validation and evaluation suites. Harness-guided benchmark construction selects tasks that respond to harness improvements… See the full description on the dataset page: https://huggingface.co/datasets/RUC-AIBOX/Evo-Bench.documentothern<1K2 likes7.6k downloads1mo agoHugging Face20runorunoruno /GheoLei_BeamNG.drive_Modsimage1K<n<10K0 likes7.2k downloads2d agoHugging Face21dsaddsaf /sn80-data-run30 likes6.6k downloads2mo agoHugging Face22veerlosar /rule-ling-concepts0 likes6.6k downloads2h agoHugging Face23gavinlaw /rl-run-archive-2026 RL run archive 2026 Archived raw run artifacts (rollout trajectories, rendered frames, policy and optimizer checkpoints, configs, logs) from simulation reinforcement-learning experiments, published for long-term preservation and reproducibility. Layout mirrors the verified backup trees they were copied from: tilde/20260915-102000/ and taurus/20260915-085631/: batched tar archives. Every archive carries a per-file SHA-256 manifest inside it; the batch inventories (9998.json.gz… See the full description on the dataset page: https://huggingface.co/datasets/gavinlaw/rl-run-archive-2026.tabularn<1K0 likes6.5k downloads17h agoHugging Face24dsaddsaf /sn80-data-run20 likes6.5k downloads2mo agoHugging Face25garak-llm /rubygems-20230301text100K<n<1M1 likes6.2k downloads2y agoHugging Face26garak-llm /rubygems-20241031text100K<n<1M0 likes6.2k downloads2y agoHugging Face27chukfinley /gavel-runs0 likes5.9k downloads1d agoHugging Face28UkrainianCatholicUniversity /rukopys RUKOPYS: Ukrainian Handwritten Text Recognition Dataset RUKOPYS (Ukrainian: рукопис — manuscript) is the first large-scale open dataset for Ukrainian handwritten text recognition (HTR). It spans over a century of Ukrainian handwriting — from 1920s archival documents to present-day school homework — and is designed for end-to-end document understanding: region detection, type classification, and text transcription. Ukrainian is among the largest Slavic languages (45M+ native… See the full description on the dataset page: https://huggingface.co/datasets/UkrainianCatholicUniversity/rukopys.imageobject-detection10K<n<100K22 likes5.9k downloads2mo agoHugging Face29ruikle123 /SPIN-UV SPIN-UV SPIN-UV is a multimodal dataset for unstructured scene understanding in dense urban villages. It was collected from a motor-driven, human-steered single-track vehicle and pairs front-facing visual observations with frame-anchored riding-state signals. The dataset is intended to support semantic segmentation, RGB-D perception, state-conditioned traversability, temporal consistency, and motion-aware scene understanding in narrow, weakly structured urban-village corridors.… See the full description on the dataset page: https://huggingface.co/datasets/ruikle123/SPIN-UV.imageimage-segmentationn<1K0 likes5.5k downloads2mo agoHugging Face30EssentialAI /reflection_model_outputs_run1 Reflection Model Outputs This repository contains model output results from various LLMs across multiple tasks and configurations. 📂 Dataset Structure We have 3 runs of data, and all files are organized under the main directory: EssentialAI/reflection_model_outputs_run1/ EssentialAI/reflection_model_outputs_run2/ EssentialAI/reflection_model_outputs_run3/ Within this, you will find results grouped by model architecture and checkpoint size, including: OLMo-2 7B OLMo-2… See the full description on the dataset page: https://huggingface.co/datasets/EssentialAI/reflection_model_outputs_run1.0 likes5.3k downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.