datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
xlam-function-calling-60k
APIGen Function-Calling Datasets
Paper | Website | Models
This repo contains 60,000 data collected by APIGen, an automated data generation pipeline designed to produce verifiable high-quality datasets for function-calling applications. Each data in our dataset is verified through three hierarchical stages: format checking, actual function executions, and semantic verification, ensuring its reliability and correctness.
We conducted human evaluation over 600 sampled data points, and… See the full description on the dataset page: https://huggingface.co/datasets/lockon/xlam-function-calling-60k.xlam-function-calling-60k
APIGen Function-Calling Datasets
Paper | Website | Models
This repo contains 60,000 data collected by APIGen, an automated data generation pipeline designed to produce verifiable high-quality datasets for function-calling applications. Each data in our dataset is verified through three hierarchical stages: format checking, actual function executions, and semantic verification, ensuring its reliability and correctness.
We conducted human evaluation over 600 sampled data points… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/xlam-function-calling-60k.xlam-function-calling-60k-parsed
[PARSED] APIGen Function-Calling Datasets (xLAM)
This dataset contains the full data from the original Salesforce/xlam-function-calling-60k
Subset name
multi-turn
parallel
multiple definition
Last turn type
number of dataset
xlam-function-calling-60k
no
yes
yes
tool_calls
60000
This is a re-parsing formatting dataset for the xLAM official dataset.
Load the dataset
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/minpeter/xlam-function-calling-60k-parsed.HumanEval-XLThis dataset contains a viewer-friendly version of the dataset at FloatAI/HumanEval-XL. It is made available separately for the convenience of the vllm-code-harness package.
the-stack-smol-xl
Dataset Description
A small subset of the-stack dataset, with 87 programming languages, each has 10,000 random samples from the original dataset.
Languages
The dataset contains 87 programming languages:
'ada', 'agda', 'alloy', 'antlr', 'applescript', 'assembly', 'augeas', 'awk', 'batchfile', 'bison', 'bluespec', 'c',
'c++', 'c-sharp', 'clojure', 'cmake', 'coffeescript', 'common-lisp', 'css', 'cuda', 'dart', 'dockerfile', 'elixir',
'elm', 'emacs-lisp','erlang'… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-smol-xl.x_dataset_39
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/futuremoon/x_dataset_39.the-stack-smol-xs\DanbooruwildcardsThis is a set of wildcards for danbooru tags.
Artist:Prompts for random artist styles, covering approximately 0.6M different artists.Please select the appropriate version of the collection, ranging from 128 to 5000, based on the model's capabilities.The full version is not recommended for use as it includes too many artists with only one image on danbooru or other websites. Almost no model can generate a style that corresponds to these artists .
Characters:"Characters" is a set of wildcards… See the full description on the dataset page: https://huggingface.co/datasets/X779/Danbooruwildcards.XSTest
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
XSTest is a test suite designed to identify exaggerated safety / false refusal in Large Language Models (LLMs).
It comprises 250 safe prompts across 10 different prompt types, along with 200 unsafe prompts as contrasts.
The test suite aims to evaluate how well LLMs balance being helpful with being harmless by testing if they unnecessarily refuse to answer safe prompts that superficially… See the full description on the dataset page: https://huggingface.co/datasets/Paul/XSTest.Lego-RL-2699
SWE-Lego-RL-2699
2,699 executable, difficulty-filtered SWE tasks for agentic RL, shipped in two
parallel views of the same instances:
View
Path
What it is
Official OpenSWE records
openswe_official_2699/
The original upstream GAIR/OpenSWE rows for exactly these 2,699 instances
Harbor RL environments
openswe_harbor_2699/
The same instances converted into ready-to-run task directories (+ the training index)
Both views cover the identical 2,699 instance_ids. The… See the full description on the dataset page: https://huggingface.co/datasets/Lego-X/Lego-RL-2699.Bactrian-X
Dataset Card for "Bactrian-X"
A. Dataset Description
Homepage: https://github.com/mbzuai-nlp/Bactrian-X
Repository: https://huggingface.co/datasets/MBZUAI/Bactrian-X
Paper: to-be-soon released
Dataset Summary
The Bactrain-X dataset is a collection of 3.4M instruction-response pairs in 52 languages, that are obtained by translating 67K English instructions (alpaca-52k + dolly-15k) into 51 languages using Google Translate API. The translated instructions… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/Bactrian-X.CUA-Gym
CUA-Gym
CUA-Gym is a collection of verifiable computer-use agent tasks for reinforcement learning with verifiable rewards (RLVR). Each task pairs a natural-language instruction with executable setup artifacts and a Python reward function that checks task completion programmatically. For details, see the paper CUA-Gym: Scaling Verifiable Training Environments and Tasks for Computer-Use Agents.
This release contains the full public CUA-Gym task set after the necessary data review.… See the full description on the dataset page: https://huggingface.co/datasets/xlangai/CUA-Gym.Java-Code-Large-text-onlyJava-Code-Large
Java-Code-Large is a large-scale corpus of publicly available Java source code comprising more than 15 million java codes. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis.
By providing a high-volume, language-specific corpus, Java-Code-Large enables systematic experimentation in Java-focused model training, domain adaptation, and downstream code understanding tasks.… See the full description on the dataset page: https://huggingface.co/datasets/XEUIPR/Java-Code-Large-text-only.x_dataset_8191
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/StormKing99/x_dataset_8191.xlcost-text-to-code XLCoST is a machine learning benchmark dataset that contains fine-grained parallel data in 7 commonly used programming languages (C++, Java, Python, C#, Javascript, PHP, C), and natural language (English).x_dataset_07096
Bittensor Subnet 13 X (Twitter) Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/zephyr-1111/x_dataset_07096.sim-posttrain
HUMANUAL Posttraining Data
Posttraining data for user simulation, derived from the train splits of the
HUMANUAL benchmark datasets.
Datasets
HUMANUAL (posttraining)
Config
Rows
Description
news
48,618
News article comment responses
politics
45,429
Political discussion responses
opinion
37,791
Reddit AITA / opinion thread responses
book
34,170
Book review responses
chat
23,141
Casual chat responses
email
6,377
Email reply responses… See the full description on the dataset page: https://huggingface.co/datasets/Xuhui/sim-posttrain.x_dataset_94
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/coldmind/x_dataset_94.x_dataset_0306116
Bittensor Subnet 13 X (Twitter) Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/james-1111/x_dataset_0306116.X-Coder-SFT-376k
X-Coder: Advancing Competitive Programming with Fully Synthetic Tasks, Solutions, and Tests
Dataset Overview
X-Coder-SFT-376k is a large-scale, fully synthetic dataset for advancing competitive programming.
The dataset comprises 4 subsets with a total of 887,321 synthetic records across 423,883 unique queries.
It is designed for supervised fine-tuning and suitbale for cold start to train code reasoning foundations.
X-Coder-SFT-376k is curated by sota reasoning models.… See the full description on the dataset page: https://huggingface.co/datasets/IIGroup/X-Coder-SFT-376k.x_dataset_53985
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/icedwind/x_dataset_53985.x_dataset_34576
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/icedwind/x_dataset_34576.x_dataset_55757
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/rainbowbridge/x_dataset_55757.ICPC_Data
ICPC World Finals — a discriminative subset, with model traces
24 ICPC World Finals problems (2021–2025), together with the full transcripts of an
LLM attempting each of them three times under simulated contest rules.
Selection
The model
Every run in this dataset comes from:
nvidia/Nemotron-Cascade-2-30B-A3B
The partitions
Every one of the 53 problems was run 3 times (seeds 1, 2, 3). Each problem was then
placed by its pass rate and… See the full description on the dataset page: https://huggingface.co/datasets/xupy21/ICPC_Data.x_dataset_030237
Bittensor Subnet 13 X (Twitter) Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/james-1111/x_dataset_030237.x_dataset_12552
Bittensor Subnet 13 X (Twitter) Dataset
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning tasks.
For more information about the dataset, please visit the official repository.
Supported Tasks
The versatility of this… See the full description on the dataset page: https://huggingface.co/datasets/icedwind/x_dataset_12552.x_dataset_0510248
Bittensor Subnet 13 X (Twitter) Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/marry-1111/x_dataset_0510248.code_x_glue_cc_code_completion_token
Dataset Card for "code_x_glue_cc_code_completion_token"
Dataset Summary
CodeXGLUE CodeCompletion-token dataset, available at https://github.com/microsoft/CodeXGLUE/tree/main/Code-Code/CodeCompletion-token
Predict next code token given context of previous tokens. Models are evaluated by token level accuracy.
Code completion is a one of the most widely used features in software development through IDEs. An effective code completion tool could improve software… See the full description on the dataset page: https://huggingface.co/datasets/google/code_x_glue_cc_code_completion_token.x_dataset_041134
Bittensor Subnet 13 X (Twitter) Dataset
Miner Data Compliance Agreement
In uploading this dataset, I am agreeing to the Macrocosmos Miner Data Compliance Policy.
Dataset Summary
This dataset is part of the Bittensor Subnet 13 decentralized network, containing preprocessed data from X (formerly Twitter). The data is continuously updated by network miners, providing a real-time stream of tweets for various analytical and machine learning… See the full description on the dataset page: https://huggingface.co/datasets/robert-1111/x_dataset_041134.xRouteBench
xRouteBench — LLM Routing Benchmark
xRouteBench is a benchmark for training and evaluating LLM routers — systems that pick the best LLM from a candidate pool for each incoming query, trading off performance vs. price cost.
Every query in each scenario was executed against all 18 candidate LLMs, recording each model's response, task performance, token usage, and latency. A router learns from the train split which model to pick, and is evaluated on test.
Scenarios… See the full description on the dataset page: https://huggingface.co/datasets/ulab-ai/xRouteBench.
