datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fish_datasets_real_electrodyn_expertsys_twodim_fourierGroundCUA
GroundCUA: Grounding Computer Use Agents on Human Demonstrations
🌐 Website |
📑 Paper |
🤗 Dataset |
🤖 Models
GroundCUA Dataset
GroundCUA is a large and diverse dataset of real UI screenshots paired with structured annotations for building multimodal computer use agents. It covers 87 software platforms across productivity tools, browsers, creative tools, communication apps, development environments, and system utilities. GroundCUA is designed for research on GUI… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow/GroundCUA.server
2026-05-20 Non-distance Refresh Bundle
This bundle contains refreshed paper-facing outputs copied from the server-side
TabQueryBench/code_snapshot/Evaluation/... tree after the v2 refresh runs.
Included:
subgroup_breakdown/final
conditional_breakdown/final
conditional_locality_support_final
missingness_breakdown/final
missingness_regime_diagnostic
tail_breakdown/final
tail_support_diagnostics_final
tail_threshold_final
compare_updated_figures.pdf
compare_updated_figures.tex… See the full description on the dataset page: https://huggingface.co/datasets/TabQueryBench2026/server.SERWorkArena-Instances
ServiceNow Instances for WorkArena
This repository provides access to the ServiceNow instances used for the WorkArena benchmark.
Access is restricted.Please complete the form above to request access.
Usage Scope
Instances are provided exclusively for benchmarking, evaluation, and research. They must not be used for training, production workloads, or storing sensitive, proprietary, or personally identifiable information.
Usage Monitoring
Use of the… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow/WorkArena-Instances.Time-Series-Library
Time-Series-Library (TSLib)
TSLib is an open-source library for deep learning researchers, especially for deep time series analysis.
We provide a neat code base to evaluate advanced deep time series models or develop your model, which covers five mainstream tasks: long- and short-term forecasting, imputation, anomaly detection, and classification.
This benchmark collection is designed to evaluate and develop advanced deep time-series models. For an in-depth exploration of current… See the full description on the dataset page: https://huggingface.co/datasets/thuml/Time-Series-Library.EnterpriseOps-Gym
EnterpriseOps-Gym: Environments and Evaluations for Stateful Agentic Planning and Tool Use in Enterprise Settings
EnterpriseOps-Gym is a containerized, resettable enterprise simulation benchmark for evaluating LLM agents on stateful, multi-step planning and tool use across realistic enterprise workflows
About
EnterpriseOps-Gym is a large-scale benchmark for evaluating the agentic planning and tool-use capabilities of LLM agents across enterprise operations. It… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow-AI/EnterpriseOps-Gym.turkish-raw-text-cleaned
Turkish Raw Text Cleaned
turkish-raw-text-cleaned, Türkçe dil modeli çalışmaları için hazırlanmış temizlenmiş ham metin veri kümesidir. Veri kümesi, turkish-nlp-suite çatısı altında yayımlanan Türkçe metin kaynaklarının temizlenmesi, filtrelenmesi ve model eğitimine daha uygun hale getirilmesiyle oluşturulmuştur.
Bu çalışma özellikle Türkçe LLM ön-eğitimi, continual pre-training (CPT), tokenizer analizi, embedding modeli eğitimi, alan bağımsız Türkçe metin modelleme ve veri… See the full description on the dataset page: https://huggingface.co/datasets/serda-dev/turkish-raw-text-cleaned.test-mcp-logs(Put queries first as heuristics don't detect when there are no logs)
Disaster-tweet-jailbreakingHere the link to the paper: https://link.springer.com/chapter/10.1007/978-3-031-85240-4_14
repliqa
RepLiQA - Repository of Likely Question-Answer for benchmarking
NeurIPS Datasets presentation
Dataset Summary
RepLiQA is an evaluation dataset that contains Context-Question-Answer triplets, where contexts are non-factual but natural-looking documents about made up entities such as people or places that do not exist in reality. RepLiQA is human-created, and designed to test for the ability of Large Language Models (LLMs) to find and use contextual information in provided… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow/repliqa.cybergym-serversplitted_PretrainGiftEvalseraphim-drone-detection-dataset
Seraphim Drone Detection Dataset
Dataset Overview
This is a comprehensive drone image dataset curated from 23 open-source datasets and processed through a custom cleaning pipeline. The dataset is designed for training object detection models to identify drones in various environments and conditions. The majority of images feature rotary-wing (multi-rotor) unmanned aerial vehicles (UAVs), with a smaller portion representing fixed-wing and hybrid.… See the full description on the dataset page: https://huggingface.co/datasets/lgrzybowski/seraphim-drone-detection-dataset.fish_datasets_real_fowlers_expertsys_twodim_fourier_v2R1-Distill-SFT
🔉 𝗦𝗟𝗔𝗠 𝗹𝗮𝗯 - 𝗥𝟭-𝗗𝗶𝘀𝘁𝗶𝗹𝗹-𝗦𝗙𝗧 Dataset
Lewis Tunstall, Ed Beeching, Loubna Ben Allal, Clem Delangue 🤗 and others at Hugging Face announced today that they are - 𝗼𝗽𝗲𝗻𝗹𝘆 𝗿𝗲𝗽𝗿𝗼𝗱𝘂𝗰𝗶𝗻𝗴 𝗥𝟭 🔥
We at 𝗦𝗟𝗔𝗠 𝗹𝗮𝗯 (ServiceNow Language Models) have been cooking up something as well.
Inspired by Open-r1, we have decided to open source the data stage-by-stage to support the open source community.
𝗕𝗼𝗼𝗸𝗺𝗮𝗿𝗸 this page!
KEY DETAILS:
⚗️ Distilled… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow-AI/R1-Distill-SFT.cybergym-server-binarycybergym-serverMultimodal-Dataset-Image_Text_Table_TimeSeries-for-Financial-Time-Series-ForecastingThe sp500stock_data_description.csv file provides detailed information on the existence of four modalities (text, image, time series, and table) for 4,213 S&P 500 stocks.
The hs300stock_data_description.csv file provides detailed information on the existence of four modalities (text, image, time series, and table) for 858 HS 300 stocks.
If you find our research helpful, please cite our paper:
@article{xu2025finmultitime,
title={FinMultiTime: A Four-Modal Bilingual Dataset for… See the full description on the dataset page: https://huggingface.co/datasets/Wenyan0110/Multimodal-Dataset-Image_Text_Table_TimeSeries-for-Financial-Time-Series-Forecasting.tabularbenchui-vision
UI-Vision: A Desktop-centric GUI Benchmark for Visual Perception and Interaction
Introduction
Autonomous agents that navigate Graphical User Interfaces (GUIs) to automate tasks like document editing and file management can greatly enhance computer workflows. While existing research focuses on online settings, desktop environments, critical for many professional and everyday tasks, remain underexplored due to data collection challenges… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow/ui-vision.seriguela-results
Seriguela Evaluation Results
This dataset contains evaluation results for symbolic regression models trained in the Seriguela project.
Structure
quality/ - Generation quality evaluation results (valid rate, diversity, etc.)
benchmark/ - Benchmark evaluation results (R² scores on Nguyen benchmarks)
Usage
from datasets import load_dataset
# Load all results
ds = load_dataset("augustocsc/seriguela-results")
# Or load specific files
from huggingface_hub import… See the full description on the dataset page: https://huggingface.co/datasets/augustocsc/seriguela-results.serrobotwin_scan_object_place_dual_shoes_serialized_200multiview-pouring
MultiView Pouring Dataset, v1.0.
by Pierre Sermanet, Corey Lynch, Jasmine Hsu and Eric Jang
License
This data is licensed by Google Inc. under a Creative Commons Attribution 4.0 International License.
Downloading
Because of some downloading issues for a specific file, the file was split in two, call https://huggingface.co/datasets/sermanet/multiview-pouring/blob/main/tfrecords/test/whiteorange_to_clear1_real_combining.sh to recombine the parts.… See the full description on the dataset page: https://huggingface.co/datasets/sermanet/multiview-pouring.sergio1091ipfs_serbia_laws
Laws of Serbia
Research snapshot of official legislation collected from PIS + parlament.gov.rs zakoni PDFs.
Not legal advice. Official gazettes / government portals prevail over this corpus.
Snapshot
Field
Value
Snapshot date
2026-09-21
Coverage
catalog-backed incomplete
Source
PIS + parlament.gov.rs zakoni PDFs
Collector
scrapers/collect_rs.py
Laws / instruments
6288
Articles
60111
Language
sr
Jurisdiction
Serbia
License… See the full description on the dataset page: https://huggingface.co/datasets/endomorphosis/ipfs_serbia_laws.BigDocs-7.5M
BigDocs-7.5M
Training data for the paper: BigDocs: An Open and Permissively-Licensed Dataset for Training Multimodal Models on Document and Code Tasks
🌐 Homepage | 📖 arXiv
Guide on Data Loading
Some parts of BigDocs-7.5M are distributed without their "image" column, and instead have an "img_id" column. The file get_bigdocs_75m.py, part of this repository, provides tooling to substitutes such images back in.
from get_bigdocs_75m import get_bigdocs_75m… See the full description on the dataset page: https://huggingface.co/datasets/ServiceNow/BigDocs-7.5M.LabDex
LabDex: A Hierarchical Benchmark for Dexterous Manipulation in Laboratories
Project Page
|
Paper
|
Code
Overview
LabDex is a large-scale dataset and benchmark for dexterous manipulation in chemistry laboratories, designed to support the training and systematic evaluation of robotic policies across different levels of task complexity.
LabDex provides unified real-world and simulation platforms and organizes laboratory… See the full description on the dataset page: https://huggingface.co/datasets/serandt/LabDex.dsb_audio_corpus
Acknowledgements
Thanks to all speakers that contributed to this dataset!
Thanks to "Ludowe Nakładnistwo Domowina" and "Rěčny Centrum WITAJ" for donation of their recordings!
