datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
deception-probes-activations
Deception Probes Activations
Pre-extracted residual-stream activations for training and evaluating deception
detection probes on LLMs. Each example contains per-token hidden states from a
specific transformer layer, saved in bfloat16 safetensors format.
License
This dataset contains activations derived from multiple sources with different licenses.
See the LICENSE file for full details.
Component
Source
License
Apollo Probe Pairs (statements)
Azaria & Mitchell… See the full description on the dataset page: https://huggingface.co/datasets/xycoord/deception-probes-activations.xlam-function-calling-60k
APIGen Function-Calling Datasets
Paper | Website | Models
This repo contains 60,000 data collected by APIGen, an automated data generation pipeline designed to produce verifiable high-quality datasets for function-calling applications. Each data in our dataset is verified through three hierarchical stages: format checking, actual function executions, and semantic verification, ensuring its reliability and correctness.
We conducted human evaluation over 600 sampled data points, and… See the full description on the dataset page: https://huggingface.co/datasets/lockon/xlam-function-calling-60k.xlam-function-calling-60k
APIGen Function-Calling Datasets
Paper | Website | Models
This repo contains 60,000 data collected by APIGen, an automated data generation pipeline designed to produce verifiable high-quality datasets for function-calling applications. Each data in our dataset is verified through three hierarchical stages: format checking, actual function executions, and semantic verification, ensuring its reliability and correctness.
We conducted human evaluation over 600 sampled data points… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/xlam-function-calling-60k.xfield-radar-dataset-20260915
XField radar dataset — formal snapshot, 2026-09-15
Upload status: COMPLETE — every shard verified against its remote SHA-256 and byte size. See UPLOAD_COMPLETE.json.
This public research snapshot preserves the currently admitted XField/GRT simulation dataset: radar inputs, existing GT, available raw sensor products, provenance, and the frozen N141 train/validation split. It is not a claim that historical labels meet the newly repaired independent dense-GT pipeline.… See the full description on the dataset page: https://huggingface.co/datasets/KAS2003/xfield-radar-dataset-20260915.xlam-function-calling-60kDS-1000 DS-1000 in simplified format
🔥 Check the leaderboard from Eval-Arena on our project page.
See testing code and more information (also the original fill-in-the-middle/Insertion format) in the DS-1000 repo.
Reformatting credits: Yuhang Lai, Sida Wang
the-stack-smol-xl
Dataset Description
A small subset of the-stack dataset, with 87 programming languages, each has 10,000 random samples from the original dataset.
Languages
The dataset contains 87 programming languages:
'ada', 'agda', 'alloy', 'antlr', 'applescript', 'assembly', 'augeas', 'awk', 'batchfile', 'bison', 'bluespec', 'c',
'c++', 'c-sharp', 'clojure', 'cmake', 'coffeescript', 'common-lisp', 'css', 'cuda', 'dart', 'dockerfile', 'elixir',
'elm', 'emacs-lisp','erlang'… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-smol-xl.scibench
SciBench
SciBench is a novel benchmark for college-level scientific problems sourced from instructional textbooks. The benchmark is designed to evaluate the complex reasoning capabilities,
strong domain knowledge, and advanced calculation skills of LLMs.
Please refer to our paper or website for full description: SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models
.
Citation
If you find our paper useful, please cite our… See the full description on the dataset page: https://huggingface.co/datasets/xw27/scibench.data_4
SpatialEncoder WDS release (in progress)
This repository contains a partition of spatialencoder-wds-native-v1, released
as uncompressed WebDataset tar shards, normally about 1 GiB. All five
repositories are parts of the same release; consult each manifest.json.
The manifest lists only uploaded shards whose remote size and SHA-256 have
been verified. An incomplete manifest is not a complete dataset.
New uploads use bucketed paths such as… See the full description on the dataset page: https://huggingface.co/datasets/xxxspatialencoderwds4/data_4.Gastric-X
Gastric-X
Multi-phase abdominal CT cohort paired with structured laboratory panels
and free-text radiology reports, in proficient medical English with
the original Simplified Chinese preserved alongside.
Changelog
2026-06-26
Added per-phase organ masks (<phase>_organ_mask.nii.gz) — CADS
multi-organ segmentation on each phase's CT grid (e.g. label 6 = stomach);
all 4897 phases.
Added per-phase gastric tumor masks (<phase>_tumor_mask.nii.gz,
binary) — a patient's… See the full description on the dataset page: https://huggingface.co/datasets/HaoChen2/Gastric-X.data_3
SpatialEncoder WDS release (in progress)
This repository contains a partition of spatialencoder-wds-native-v1, released
as uncompressed WebDataset tar shards, normally about 1 GiB. All five
repositories are parts of the same release; consult each manifest.json.
The manifest lists only uploaded shards whose remote size and SHA-256 have
been verified. An incomplete manifest is not a complete dataset.
New uploads use bucketed paths such as… See the full description on the dataset page: https://huggingface.co/datasets/xxxspatialencoderwds3/data_3.data_2
SpatialEncoder WDS release (in progress)
This repository contains a partition of spatialencoder-wds-native-v1, released
as uncompressed WebDataset tar shards, normally about 1 GiB. All five
repositories are parts of the same release; consult each manifest.json.
The manifest lists only uploaded shards whose remote size and SHA-256 have
been verified. An incomplete manifest is not a complete dataset.
New uploads use bucketed paths such as… See the full description on the dataset page: https://huggingface.co/datasets/xxxspatialencoderwds2/data_2.DRIM-VisualReasonHardThis repository contains the RL training datasets used in the paper Deep But Reliable: Advancing Multi-turn Reasoning for Thinking with Images
osworld_v2_tasks
OSWorld V2 Task Classes
This gated dataset contains the official root-level task_*.py Python task classes for OSWorld V2.
The public GitHub repository keeps the task loader, helper utilities, and documentation. The task implementations are gated to reduce benchmark leakage and to help prevent evaluated agents from finding task answers, setup logic, or evaluator details online while executing a task.
Download from the public repository root with:
uvx --from huggingface_hub hf… See the full description on the dataset page: https://huggingface.co/datasets/xlangai/osworld_v2_tasks.data_1
SpatialEncoder WDS release (in progress)
This repository contains a partition of spatialencoder-wds-native-v1, released
as uncompressed WebDataset tar shards, normally about 1 GiB. All five
repositories are parts of the same release; consult each manifest.json.
The manifest lists only uploaded shards whose remote size and SHA-256 have
been verified. An incomplete manifest is not a complete dataset.
New uploads use bucketed paths such as… See the full description on the dataset page: https://huggingface.co/datasets/xxxspatialencoderwds1/data_1.OXE-slice1-xintongLego-RL-2699
SWE-Lego-RL-2699
2,699 executable, difficulty-filtered SWE tasks for agentic RL, shipped in two
parallel views of the same instances:
View
Path
What it is
Official OpenSWE records
openswe_official_2699/
The original upstream GAIR/OpenSWE rows for exactly these 2,699 instances
Harbor RL environments
openswe_harbor_2699/
The same instances converted into ready-to-run task directories (+ the training index)
Both views cover the identical 2,699 instance_ids. The… See the full description on the dataset page: https://huggingface.co/datasets/Lego-X/Lego-RL-2699.xauusd-gold-price-historical-data-2004-2025
XAUUSD Gold Price Historical Data 2004-2025
This dataset contains historical price data for XAUUSD (Gold vs US Dollar) from 2004 to 2025.
Source: Kaggle dataset "novandraanugrah/xauusd-gold-price-historical-data-2004-2024"
Content:
The dataset includes CSV files with different time granularities (e.g., 1 minute, 5 minutes, 1 hour, 1 day). Each file typically contains the following columns:
Date
Open
High
Low
Close
Volume
Usage:
This dataset can be used for analyzing historical… See the full description on the dataset page: https://huggingface.co/datasets/ZombitX64/xauusd-gold-price-historical-data-2004-2025.xlam-function-calling-60k-shareGPTShareGPT converted version of Salesforce/xlam-function-calling-60k
XLRS-Bench-lite_VLM
🐙GitHub
Information or evaluatation on this dataset can be found in this repo: https://github.com/AI9Stars/XLRS-Bench
📜Dataset License
Annotations of this dataset is released under a Creative Commons Attribution-NonCommercial 4.0 International License. For images from:
DOTARGB images from Google Earth and CycloMedia (for academic use only; commercial use is prohibited, and Google Earth terms of use apply).
ITCVDLicensed under CC-BY-NC-SA-4.0.
MiniFrance… See the full description on the dataset page: https://huggingface.co/datasets/initiacms/XLRS-Bench-lite_VLM.Java-Code-Large-text-onlyJava-Code-Large
Java-Code-Large is a large-scale corpus of publicly available Java source code comprising more than 15 million java codes. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis.
By providing a high-volume, language-specific corpus, Java-Code-Large enables systematic experimentation in Java-focused model training, domain adaptation, and downstream code understanding tasks.… See the full description on the dataset page: https://huggingface.co/datasets/XEUIPR/Java-Code-Large-text-only.x402-bazaar-conformance
x402 Bazaar conformance census
3,520 distinct hosts, every one listed in a public x402 Bazaar
(Coinbase CDP or PayAI), probed 2026-09-05.
hosts
share
conform (402 + PAYMENT-REQUIRED + x402Version 2 + extensions.bazaar)
394
11.19%
answer 402 at all
1,536
43.64%
carry the PAYMENT-REQUIRED header
1,395
declare x402Version 2
538
carry the bazaar extension an indexer reads
413
do not respond at all
1,027
29.18%
Being listed is not the same as… See the full description on the dataset page: https://huggingface.co/datasets/csoai/x402-bazaar-conformance.LLaVA-CoT-100k
Dataset Card for LLaVA-CoT
The LLaVA-CoT-100k dataset is introduced in the paper LLaVA-CoT: Let Vision Language Models Reason Step-by-Step. This dataset is designed to enable Vision-Language Models (VLMs) to perform autonomous multistage reasoning, integrating samples from various visual question-answering sources with structured reasoning annotations. It aims to address the challenges VLMs face in systematic and structured reasoning for complex visual question-answering tasks.… See the full description on the dataset page: https://huggingface.co/datasets/Xkev/LLaVA-CoT-100k.xlam-irrelevance-7.5k
xlam-irrelevance-7.5k
Overview
The xlam-irrelevance-7.5k is a specialized dataset designed to activate the ability of irrelevant function detection for large language models (LLMs).
Source and Construction
This dataset is built upon xlam-function-calling-60k dataset, from which we random sampled 7.5k instances, removed the ground truth function from the provided tool list, and relabel them as irrelevant. For more details, please refer to Hammer: Robust… See the full description on the dataset page: https://huggingface.co/datasets/MadeAgents/xlam-irrelevance-7.5k.gero-research-evidence-2026-09
GERO research evidence — 120 publications
This dataset contains 120 distinct report, case-study, experiment, preprint and research-map records, with individual Markdown pages. All previous 119 corpus rows, including the Collatz map, are preserved byte for byte. The newest addition is the actuarialmath continuous-annuity selection-duration audit, with verified developer issue7 and explicit limitations. Report counts are not independent-defect counts.
Latest numerical… See the full description on the dataset page: https://huggingface.co/datasets/XamitK/gero-research-evidence-2026-09.opengpt-x_mmluxThis is a copy of the translations from openGPT-X/mmlux, but the repo is
modified so it doesn't require trusting remote code.
Citation Information
If you find benchmarks useful in your research, please consider citing the test and also the MMLU dataset it draws from:
@misc{thellmann2024crosslingual,
title={Towards Cross-Lingual LLM Evaluation for European Languages},
author={Klaudia Thellmann and Bernhard Stadler and Michael Fromm and Jasper Schulze Buschhoff and Alex… See the full description on the dataset page: https://huggingface.co/datasets/LumiOpen/opengpt-x_mmlux.TIR-Bench
TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning
Introduction:
TIR-Bench is a comprehensive benchmark designed to evaluate the "thinking-with-images" capabilities of Multimodal Large Language Models (MLLMs), addressing a gap left by existing benchmarks like Visual Search which only test basic operations. As models like OpenAI o3 begin to intelligently create and operate tools to transform images for problem-solving, TIR-Bench provides 13… See the full description on the dataset page: https://huggingface.co/datasets/Agents-X/TIR-Bench.xent-tasks-gpt2MechVQA
Dataset Card for MechVQA VQA SFT
MechVQA VQA SFT is a bilingual visual question answering dataset for
supervised fine-tuning on mechanical engineering drawings. This public release
contains 13,515 question-answer records paired with 3,371 unique,
content-addressed images. Assistant targets use
<think>...</think><answer>...</answer> formatting.
This repository is the VQA-only SFT train/validation release associated
with the MechVQA project. The public evaluation benchmark is… See the full description on the dataset page: https://huggingface.co/datasets/XiaofengAlg/MechVQA.EventBench
EventBench Dataset Access Instructions
EventBench 数据集访问说明
Notice / 通知
From 2026.5.20 to 2026.7.10, we are hosting EventBench competitions at @ECCV. To ensure fairness, the standard answers have been hidden during this period. We will make the standard answers available again after the competitions end.
If you would like to obtain evaluation results, please submit your predictions directly to the competition servers.
2026.5.20 至 2026.7.10 期间,我们在 @ECCV 部署了… See the full description on the dataset page: https://huggingface.co/datasets/XduSyL/EventBench.
