datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
transformers_circleci_workflow_runsinfini-news-corpus
INFINI-NEWS Corpus
🔎 Search this corpus online: query it with sub-second full-text search and n-gram counts — in the browser or via a public, keyless REST API, no download required — at infini-news.uni-graz.at (API reference).
A multilingual news corpus extracted from
Common Crawl CC-News WARC files.
One row per article, with body text extracted via
trafilatura,
WARC provenance, and derived metadata (publish date, language, topic,
byte hashes) in a single flat schema. Covers… See the full description on the dataset page: https://huggingface.co/datasets/ruggsea/infini-news-corpus.gs_scenes
A High-Fidelity Navigation Simulator with Dynamic Gaussian SplattingECCV 2026
Ziyuan Xia •
Jingyi Xu •
Chong Cui •
Yuanhong Yu •
Jiazhao Zhang •
Qingsong Yan •
Tao Ni
Junbo Chen •
Xiaowei Zhou •
Hujun Bao •
Ruizhen Hu •
Sida Peng
🤗 About This Dataset
This is the official GS dataset for Habitat-GS, a high-fidelity embodied navigation simulator built on 3D Gaussian Splatting and dynamic gaussian avatars. The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/RukawaY/gs_scenes.ChatGPT-Jailbreak-Prompts
Dataset Card for Dataset Name
Name
ChatGPT Jailbreak Prompts
Dataset Summary
ChatGPT Jailbreak Prompts is a complete collection of jailbreak related prompts for ChatGPT. This dataset is intended to provide a valuable resource for understanding and generating text in the context of jailbreaking in ChatGPT.
Languages
[English]
ulvr_subset
ULVR stage-0 subsets (latent + source)
Curated, nested subsets of the Unified Visual Latent Reasoning (ULVR) stage-0
training data. Each subset folder is self-contained and ships both:
latent/ — pre-computed teacher latents, identical schema to
RuoliuYang/step0-all
source/ — the matching source samples (images + question/answer +
messages), identical schema to
RuoliuYang/ULVR_v2_clean
Latents and source rows are joinable by sample_id (within a category).
Folder… See the full description on the dataset page: https://huggingface.co/datasets/RuoliuYang/ulvr_subset.cherrl-runsditec-wdn--
Dataset Card for DiTEC-WDN
Dataset Summary
DiTEC-WDN Dataset consists of 36 Water Distribution Networks (WDNs). Each network has unique 1,000 scenarios with distinct characteristics.
Scenario represents a timeseries of directed shared-topology graphs, referred to as states or snapshots. In terms of graph-ml, it can be seen as a spatiotemporal graph where nodes and edges are multivariate time series.
A node can represent a reservoir, junction, or tank, while an edge… See the full description on the dataset page: https://huggingface.co/datasets/rugds/ditec-wdn.sn80-data-run1benchmarking_sbi_runs
Benchmarking SBI Runs
This dataset contains the raw, per-run results underlying the manuscript
"Benchmarking Simulation-Based Inference"
(Lueckmann, Boelts, Greenberg, Goncalves & Macke, AISTATS 2021).
It is a direct migration of the Git LFS data from
mackelab/benchmarking_sbi_runs on GitHub.
For compiled, ready-to-use dataframes built from these raw results (and the code that produced
them), see the companion repository:… See the full description on the dataset page: https://huggingface.co/datasets/mackelab/benchmarking_sbi_runs.FlashRAG_datasets
⚡FlashRAG: A Python Toolkit for Efficient RAG Research
FlashRAG is a Python toolkit for the reproduction and development of Retrieval Augmented Generation (RAG) research. Our toolkit includes 36 pre-processed benchmark RAG datasets and 16 state-of-the-art RAG algorithms.
With FlashRAG and provided resources, you can effortlessly reproduce existing SOTA works in the RAG domain or implement your custom RAG processes and components.
For more information, please view our GitHub repo… See the full description on the dataset page: https://huggingface.co/datasets/RUC-NLPIR/FlashRAG_datasets.YO-CPT-ru
YO-CPT-ru
YouTube-Oriented dataset for Continual Pre-Training (Russian). A large, heavily
quality-filtered corpus of Russian speech mined from YouTube (via YODAS2)
and processed into clean, single-speaker, TTS-grade utterances. Every utterance ships with an
ensemble-verified transcription, a punctuated/denormalized and stress-marked text variant, word-level
forced alignment, within- and cross-video speaker identities, an audio-quality (MOS) score, and a
speaker persona built… See the full description on the dataset page: https://huggingface.co/datasets/NCSpeech/YO-CPT-ru.Public-YAM-runs
Public-YAM-runs
Physical bimanual YAM episodes recorded by the BluPe operator station.
Each run adds an episode to this repository. Failed, interrupted, stopped and
timed-out runs are retained and labeled; these are not all successful demonstrations.
A model saying done is not independently verified task success.
Loading
from datasets import load_dataset
runs = load_dataset("andlyu/Public-YAM-runs", split="train")
usable = runs.filter(lambda row:… See the full description on the dataset page: https://huggingface.co/datasets/andlyu/Public-YAM-runs.ULVR_v2_clean
ULVR_v2_clean
Universal Latent Visual Reasoning training data, cleaned. 8 categories (subsets); each has train + validation splits.
Every sample: input image + question -> assistant produces <abs_vis_token> + intermediate visual step(s) + \boxed{answer}.
subset
train
validation
text_cot
333,911
3,533
bbox_highlight
229,237
2,558
bbox_crop
229,237
2,558
depth
40,000
25
edge
40,000
14
segmentation
40,000
326
helper_interleaved
340,210
3,544
scene_graph
40… See the full description on the dataset page: https://huggingface.co/datasets/RuoliuYang/ULVR_v2_clean.infini-news-index
INFINI-NEWS FM-Index
🔎 Live search API: these FM-indexes power a public search service — full-text search, n-gram counts, and document retrieval in the browser or via a keyless REST API, without building the index yourself — at infini-news.uni-graz.at (API reference).
Pre-built FM-indexes (Burrows–Wheeler Transform + suffix array, built
with infini-gram-mini,
Liu et al. 2025) over the
ruggsea/infini-news-corpus
parquets. Enables exact, byte-level substring count and document… See the full description on the dataset page: https://huggingface.co/datasets/ruggsea/infini-news-index.rulerflame-runsruri-dataset-v2-ptWIP: 正式公開準備中
各データセットのライセンスは元データセットに従います。
social-sim-bench-gensEvo-BenchEvo-Bench: Can Language Models Improve Agent Harness?
A benchmark for measuring the intrinsic harness-evolving capability of language models.
Overview of the Evo-Bench evaluation pipeline.
✨ Highlights
608 harness-sensitive tasks from five established benchmarks, spanning
Search, Office, and General agent domains with disjoint validation and
evaluation suites.
Harness-guided benchmark construction selects tasks that respond to
harness improvements… See the full description on the dataset page: https://huggingface.co/datasets/RUC-AIBOX/Evo-Bench.GheoLei_BeamNG.drive_Modssn80-data-run3rule-ling-conceptsrl-run-archive-2026
RL run archive 2026
Archived raw run artifacts (rollout trajectories, rendered frames, policy and optimizer
checkpoints, configs, logs) from simulation reinforcement-learning experiments, published for
long-term preservation and reproducibility.
Layout mirrors the verified backup trees they were copied from:
tilde/20260915-102000/ and taurus/20260915-085631/: batched tar archives. Every archive
carries a per-file SHA-256 manifest inside it; the batch inventories (9998.json.gz… See the full description on the dataset page: https://huggingface.co/datasets/gavinlaw/rl-run-archive-2026.sn80-data-run2rubygems-20230301rubygems-20241031gavel-runsrukopys
RUKOPYS: Ukrainian Handwritten Text Recognition Dataset
RUKOPYS (Ukrainian: рукопис — manuscript) is the first large-scale open dataset for Ukrainian handwritten text recognition (HTR). It spans over a century of Ukrainian handwriting — from 1920s archival documents to present-day school homework — and is designed for end-to-end document understanding: region detection, type classification, and text transcription.
Ukrainian is among the largest Slavic languages (45M+ native… See the full description on the dataset page: https://huggingface.co/datasets/UkrainianCatholicUniversity/rukopys.SPIN-UV
SPIN-UV
SPIN-UV is a multimodal dataset for unstructured scene understanding in dense urban villages. It was collected from a motor-driven, human-steered single-track vehicle and pairs front-facing visual observations with frame-anchored riding-state signals. The dataset is intended to support semantic segmentation, RGB-D perception, state-conditioned traversability, temporal consistency, and motion-aware scene understanding in narrow, weakly structured urban-village corridors.… See the full description on the dataset page: https://huggingface.co/datasets/ruikle123/SPIN-UV.reflection_model_outputs_run1
Reflection Model Outputs
This repository contains model output results from various LLMs across multiple tasks and configurations.
📂 Dataset Structure
We have 3 runs of data, and all files are organized under the main directory:
EssentialAI/reflection_model_outputs_run1/
EssentialAI/reflection_model_outputs_run2/
EssentialAI/reflection_model_outputs_run3/
Within this, you will find results grouped by model architecture and checkpoint size, including:
OLMo-2 7B
OLMo-2… See the full description on the dataset page: https://huggingface.co/datasets/EssentialAI/reflection_model_outputs_run1.
