datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
xlam-function-calling-60k
APIGen Function-Calling Datasets
Paper | Website | Models
This repo contains 60,000 data collected by APIGen, an automated data generation pipeline designed to produce verifiable high-quality datasets for function-calling applications. Each data in our dataset is verified through three hierarchical stages: format checking, actual function executions, and semantic verification, ensuring its reliability and correctness.
We conducted human evaluation over 600 sampled data points, and… See the full description on the dataset page: https://huggingface.co/datasets/lockon/xlam-function-calling-60k.xlam-function-calling-60k
APIGen Function-Calling Datasets
Paper | Website | Models
This repo contains 60,000 data collected by APIGen, an automated data generation pipeline designed to produce verifiable high-quality datasets for function-calling applications. Each data in our dataset is verified through three hierarchical stages: format checking, actual function executions, and semantic verification, ensuring its reliability and correctness.
We conducted human evaluation over 600 sampled data points… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/xlam-function-calling-60k.the-stack-smol-xl
Dataset Description
A small subset of the-stack dataset, with 87 programming languages, each has 10,000 random samples from the original dataset.
Languages
The dataset contains 87 programming languages:
'ada', 'agda', 'alloy', 'antlr', 'applescript', 'assembly', 'augeas', 'awk', 'batchfile', 'bison', 'bluespec', 'c',
'c++', 'c-sharp', 'clojure', 'cmake', 'coffeescript', 'common-lisp', 'css', 'cuda', 'dart', 'dockerfile', 'elixir',
'elm', 'emacs-lisp','erlang'… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-smol-xl.Lego-RL-2699
SWE-Lego-RL-2699
2,699 executable, difficulty-filtered SWE tasks for agentic RL, shipped in two
parallel views of the same instances:
View
Path
What it is
Official OpenSWE records
openswe_official_2699/
The original upstream GAIR/OpenSWE rows for exactly these 2,699 instances
Harbor RL environments
openswe_harbor_2699/
The same instances converted into ready-to-run task directories (+ the training index)
Both views cover the identical 2,699 instance_ids. The… See the full description on the dataset page: https://huggingface.co/datasets/Lego-X/Lego-RL-2699.Java-Code-Large-text-onlyJava-Code-Large
Java-Code-Large is a large-scale corpus of publicly available Java source code comprising more than 15 million java codes. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis.
By providing a high-volume, language-specific corpus, Java-Code-Large enables systematic experimentation in Java-focused model training, domain adaptation, and downstream code understanding tasks.… See the full description on the dataset page: https://huggingface.co/datasets/XEUIPR/Java-Code-Large-text-only.LegoFlow-SWE
LegoFlow-SWE · 5,000 verified Harbor SWE tasks and two GLM-5.2 trajectory releases
GitHub · Docs · Blog · HuggingFace · LegoX
LegoFlow-SWE
5,000 verified Harbor SWE tasks mined by LegoFlow Curator, shipped in original and anti-hack prompt versions, plus two GLM-5.2 trajectory releases under OpenHands SDK and OpenCode, totaling 9,767 trajectories.
Release
Count
What it is
tasks/
5,000
Original prompts
tasks-anti-hack/
5,000
Same task IDs and task files, with… See the full description on the dataset page: https://huggingface.co/datasets/Lego-X/LegoFlow-SWE.gspc-xr
GSPC — cross reality bank (XRAIV)
Council of AI measurement bank. Measurement, not certification.
Bank. Frozen split. Live n is the matching axis on GET https://councilof.ai/api/gspc, not a Hub score. Not a certificate. Art 50 (EUR-Lex): 2 August 2026 live; marking grace 2 December 2026.
Live measurement. This bank stands behind the cross-reality row of the live GSPC board: GET https://councilof.ai/api/gspc?axis=cross-reality (family, kind, status and n are on that row, never… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-xr.DSULT-Core-ShareGPT-X
DSULT-Core/ShareGPT-X Filtered Dataset
This dataset is a curated subset of ShareGPT-X, which contains approximately 92,000 one-to-one conversations between humans and ChatGPT, collected from X.com (formerly Twitter). The corpus covers content from January 2024 through May 2025, built entirely from public "share" links posted by users on their timelines.
The file ChatGPT-Simple_ShareGPT_Full.json includes the longest sequences of alternating human and gpt messages within each… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/DSULT-Core-ShareGPT-X.XYZ-Aquila-SFT
XYZ-Aquila SFT
XYZ-Aquila SFT is a bilingual release of 7,000 multi-turn, search-oriented
tool-use trajectories, comprising 5,000 English examples and 2,000 Chinese
examples.
This release is a sample of the broader supervised fine-tuning data used for
XYZ-Aquila-mini and
XYZ-Aquila-pro. The examples
capture agent interactions with search tools, intermediate observations, and
answer generation in English and Chinese.
A small portion of the QA content is derived from… See the full description on the dataset page: https://huggingface.co/datasets/XYZAILab/XYZ-Aquila-SFT.persona
PERSONA: Dynamic and Compositional Inference-Time Personality Control
Official release of persona vectors and SFT datasets for the ICLR 2026 paper:
PERSONA: Dynamic and Compositional Inference-Time Personality Control via Activation Vector Algebra
Xiachong Feng, Liang Zhao, Weihong Zhong, Yichong Huang, Yuxuan Gu, Lingpeng Kong, Xiaocheng Feng, Bing Qin
Harbin Institute of Technology & The University of Hong Kong
Paper: https://openreview.net/pdf?id=QZvGqaNBlU
Code:… See the full description on the dataset page: https://huggingface.co/datasets/xiachongfeng/persona.PHP-Code-LargePHP-Code-Large
PHP-Code-Large is a large-scale corpus of PHP source code comprising more than 12 million lines of PHP code. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and static program analysis for the PHP ecosystem.
By providing a high-volume, language-specific corpus, PHP-Code-Large enables systematic experimentation in PHP-focused model training, domain adaptation, and downstream code… See the full description on the dataset page: https://huggingface.co/datasets/xormania/PHP-Code-Large.NaturalConv
NaturalConv: A Chinese Dialogue Dataset Towards Multi-turn Topic-driven Conversation
Introduction
This dataset is described in the paper NaturalConv: A Chinese Dialogue Dataset Towards Multi-turn Topic-driven Conversation. The entire dataset contains 5 data files.
1. dialog_release.json:
It is a json file containing a list of dictionaries.
After loading in python this way:
import json
import codecs
dialog_list = json.loads(codecs.open("dialog_release.json"… See the full description on the dataset page: https://huggingface.co/datasets/xywang1/NaturalConv.Agent-G2-ALFWorld-Webshop-sft-data
Agent-G2 SFT Data
Agent-G2 SFT Data contains reasoning and action trajectories for supervised
fine-tuning (SFT) in the Agent-G2 project.
Associated paper: Agent-G2: Gaussian Guidance for Agentic Reinforcement
Learning — accepted to the EMNLP 2026 Main Conference.
The dataset covers two interactive agent environments:
WebShop: agents search for products, select options, and complete
purchases according to user requirements.
ALFWorld: agents interact with household environments… See the full description on the dataset page: https://huggingface.co/datasets/xiamoent/Agent-G2-ALFWorld-Webshop-sft-data.opengpt-x_truthfulqaxThis is a copy of the translations from openGPT-X/truthfulqax, but the repo is
modified so it doesn't require trusting remote code.
Citation Information
If you find benchmarks useful in your research, please consider citing the test and also the TruthfulQA dataset it draws from:
@misc{thellmann2024crosslingual,
title={Towards Cross-Lingual LLM Evaluation for European Languages},
author={Klaudia Thellmann and Bernhard Stadler and Michael Fromm and Jasper Schulze Buschhoff… See the full description on the dataset page: https://huggingface.co/datasets/LumiOpen/opengpt-x_truthfulqax.CPsyCoun
CPsyCounD
The high-quality multi-turn dialogue dataset, which has a total of 3,134 multi-turn consultation dialogues. CPsyCounD covers nine representative topics and seven classic schools of psychological counseling.
Paper: CPsyCoun
Data analysis
Topic types
Self-growth
Emotion&Stress
Education
Love&Marriage
Family Relationship
Social Relationship
Sex
Career
Mental Disease
Consulting schools
Psychoanalytic Therapy
Cognitive Behavioral Therapy… See the full description on the dataset page: https://huggingface.co/datasets/CAS-SIAT-XinHai/CPsyCoun.MedInstruct
Dataset Card for MedInstruct
Dataset Summary
MedInstruct encompasses:
MedInstruct-52k: A dataset comprising 52,000 medical instructions and responses. Instructions are crafted by OpenAI's GPT-4 engine, and the responses are formulated by the GPT-3.5-turbo engine.
MedInstruct-test: A set of 217 clinical craft free-form instruction evaluation tests.
med_seed: The clinician-crafted seed set as a denomination to prompt GPT-4 for task generation.
MedInstruct-52k can be used… See the full description on the dataset page: https://huggingface.co/datasets/xz97/MedInstruct.Lego-RL-SWE-Bench-Verified
Lego-RL-SWE-Bench-Verified
The 500 SWE-bench Verified instances as ready-to-run harbor RL environments —
the exact evaluation set behind every SWE-bench Verified number in
LEGO-RL, packaged the same way as the training
set Lego-X/Lego-RL-2699 so
one trainer reads both.
Two parallel views of the same 500 instances:
View
Path
What it is
Official SWE-bench records
swebench_verified_official_500/
The upstream princeton-nlp/SWE-bench_Verified rows, verbatim
Harbor RL… See the full description on the dataset page: https://huggingface.co/datasets/Lego-X/Lego-RL-SWE-Bench-Verified.MixEval-X
🚀 Project Page | 📜 arXiv | 👨💻 Github | 🏆 Leaderboard | 📝 blog | 🤗 HF Paper | 𝕏 Twitter
MixEval-X encompasses eight input-output modality combinations and can be further extended. Its data points reflect real-world task distributions. The last grid presents the scores of frontier organizations’ flagship models on MixEval-X, normalized to a 0-100 scale, with MMG tasks using win rates instead of Elo. Section C of the paper presents example data samples and model responses.… See the full description on the dataset page: https://huggingface.co/datasets/MixEval/MixEval-X.xenia-principalities
XENIA PRINCIPALITIES
PRINCIPALITIES is a small, versioned curriculum that preserves one
attributable human testimony about truth, love, understanding, freedom,
choice, thought, capability, and power. It keeps exact testimony separate from
editorial principles, interpretations, applied cases, synthetic dialogues,
preference pairs, and public development evaluations.
The corpus is intended for inspectable language-model research. It does not
ask a model or person to affirm a… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/xenia-principalities.ShareGPT-X
Dataset Summary
ShareGPT-X is an expanded, snapshot of ~92K (ChatGPT) one-to-one human & LLM conversations harvested from X.com (formerly Twitter).The corpus spans January 2024 → present (last ingest 2025-05) and is built entirely from public "share" links that users posted to their timelines.Each thread contains the original user prompt plus the assistant’s reply; no system prompts or metadata are exposed.
Supported Tasks and Leaderboards
text-generation… See the full description on the dataset page: https://huggingface.co/datasets/DSULT-Core/ShareGPT-X.29K_Python_Docstring_Pairs
29K High-Quality Python Docstring Pairs
Author: Michael Hernandez (XxCotHGxX)License: CC BY 4.0Cleaned from: XxCotHGxX/242K_Python_Docstring_Pairs
Overview
A curated, high-quality subset of Python function–docstring pairs for use in code documentation generation, docstring completion, and code understanding tasks.
The original 242K dataset was scraped from open-source Python repositories but contained a significant proportion of functions without docstrings (84% of… See the full description on the dataset page: https://huggingface.co/datasets/XxCotHGxX/29K_Python_Docstring_Pairs.MMC
MMC: Advancing Multimodal Chart Understanding with LLM Instruction Tuning
This repo releases data introduced in our paper MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning.
The paper was published in NAACL 2024.
See our GithHub repo for demo code and more.
Highlights
We introduce a large-scale MultiModal Chart Instruction (MMC-Instruction) dataset supporting diverse tasks and chart types. Leveraging this data.
We also propose a… See the full description on the dataset page: https://huggingface.co/datasets/xywang1/MMC.vulcan
🔥 Vulcan
A high-signal, fully-deduplicated SFT dataset for front-end code generation
HTML · CSS · Vanilla JS — self-contained, accessible, production-ready
Overview
Vulcan is a curated supervised fine-tuning (SFT) dataset built to teach language
models how to write clean, modern, self-contained front-end code. Every example
pairs a realistic developer request with a complete, working answer — semantic HTML5,
responsive CSS (Flexbox / Grid)… See the full description on the dataset page: https://huggingface.co/datasets/xlelords/vulcan.xl-instruct
Dataset Card for XL-Instruct
This dataset card provides a summary of the XL-Instruct dataset, a resource for advancing the cross-lingual capabilities of Large Language Models. It was introduced in the paper XL-Instruct: Synthetic Data for Cross-Lingual Open-Ended Generation.
Dataset Details
Dataset Description
XL-Instruct is a high-quality, large-scale synthetic dataset designed to fine-tune LLMs for cross-lingual open-ended generation. The core task involves… See the full description on the dataset page: https://huggingface.co/datasets/viyer98/xl-instruct.sommelier-xlam-single-call-splits
sommelier xlam single-call splits
Deterministic, deduplicated, single-tool-call train/validation/test
splits derived from
Salesforce/xlam-function-calling-60k
(APIGen, CC-BY-4.0), produced by the
sommelier pipeline for
reproducible tool-calling fine-tuning. These are the exact splits used to
train and evaluate
abdelstark/llama-3.1-nemotron-nano-8b-xlam-tool-calling-lora.
Why single-call
The upstream dataset mixes single-call and multi-call examples (~52.6%… See the full description on the dataset page: https://huggingface.co/datasets/abdelstark/sommelier-xlam-single-call-splits.ACCOUNTING_DATABASESTeichAI-thinking-reasoning-x
TeichAI Thinking & Reasoning Datasets
A collection of prompts answered by large language models (LLMs) such as Google Gemini and OpenAI ChatGPT, with long-form reasoning enabled.
These datasets were originally created by TeichAI for distillation and reasoning-focused training workflows.
Schema
Each row in the dataset has the following fields:
question_hash: Truncated, base64-encoded MD5 hash of the question, useful for filtering and deduplication.
question: The… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/TeichAI-thinking-reasoning-x.priya-sft
Priya SFT + RAG dataset
Synthetic training and retrieval data for the Priya persona — a fictional Senior Customer Success Manager at a fictional B2B SaaS company. Used to train kader-xai/priya-qwen2.5-7b-lora and kader-xai/priya-qwen2.5-7b-gguf.
Fully synthetic. No real person, customer, or company. Generated as the seed corpus for Project Recall, an experiment in employee-continuity AI.
📝 Blog post: Employee Recall — Capturing a Departing Employee's Writing Style and… See the full description on the dataset page: https://huggingface.co/datasets/kader-xai/priya-sft.OpenCharacter
OpenCharacter: Training Customizable Role-Playing LLMs with Large-Scale Synthetic Personas
This repo releases data introduced in our paper OpenCharacter: Training Customizable Role-Playing LLMs with Large-Scale Synthetic Personas in arXiv.
We study customizable role-playing dialogue agents in large language models (LLMs).
We tackle the challenge with large-scale data synthesis: character synthesis and character-driven reponse synthesis.
Our solution strengthens the original LLaMA-3… See the full description on the dataset page: https://huggingface.co/datasets/xywang1/OpenCharacter.aria-wildchat-sft-v1
Aria v1 Instruction Dataset
What This Is
Aria is a demonstration of a different way of developing model personas, one in which the models themselves participate.
We believe that existing model alignment techniques that focus on rule-following are more fragile than a system with a stable identity, where behavior can flow from that identity. We think a model should have a clear sense of what it is, what its perspective is, what its history is, and that these things should… See the full description on the dataset page: https://huggingface.co/datasets/xlr8harder/aria-wildchat-sft-v1.
