CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01lockon /xlam-function-calling-60k APIGen Function-Calling Datasets Paper | Website | Models This repo contains 60,000 data collected by APIGen, an automated data generation pipeline designed to produce verifiable high-quality datasets for function-calling applications. Each data in our dataset is verified through three hierarchical stages: format checking, actual function executions, and semantic verification, ensuring its reliability and correctness. We conducted human evaluation over 600 sampled data points, and… See the full description on the dataset page: https://huggingface.co/datasets/lockon/xlam-function-calling-60k.textquestion-answering10K<n<100K1 likes39k downloads2y agoHugging Face02Salesforce /xlam-function-calling-60kgated APIGen Function-Calling Datasets Paper | Website | Models This repo contains 60,000 data collected by APIGen, an automated data generation pipeline designed to produce verifiable high-quality datasets for function-calling applications. Each data in our dataset is verified through three hierarchical stages: format checking, actual function executions, and semantic verification, ensuring its reliability and correctness. We conducted human evaluation over 600 sampled data points… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/xlam-function-calling-60k.textquestion-answering10K<n<100K720 likes37k downloads2y agoHugging Face03bigcode /the-stack-smol-xl Dataset Description A small subset of the-stack dataset, with 87 programming languages, each has 10,000 random samples from the original dataset. Languages The dataset contains 87 programming languages: 'ada', 'agda', 'alloy', 'antlr', 'applescript', 'assembly', 'augeas', 'awk', 'batchfile', 'bison', 'bluespec', 'c', 'c++', 'c-sharp', 'clojure', 'cmake', 'coffeescript', 'common-lisp', 'css', 'cuda', 'dart', 'dockerfile', 'elixir', 'elm', 'emacs-lisp','erlang'… See the full description on the dataset page: https://huggingface.co/datasets/bigcode/the-stack-smol-xl.tabulartext-generation100K<n<1M11 likes8.1k downloads4y agoHugging Face04Lego-X /Lego-RL-2699 SWE-Lego-RL-2699 2,699 executable, difficulty-filtered SWE tasks for agentic RL, shipped in two parallel views of the same instances: View Path What it is Official OpenSWE records openswe_official_2699/ The original upstream GAIR/OpenSWE rows for exactly these 2,699 instances Harbor RL environments openswe_harbor_2699/ The same instances converted into ready-to-run task directories (+ the training index) Both views cover the identical 2,699 instance_ids. The… See the full description on the dataset page: https://huggingface.co/datasets/Lego-X/Lego-RL-2699.texttext-generation1K<n<10K2 likes3.7k downloads1mo agoHugging Face05XEUIPR /Java-Code-Large-text-onlyJava-Code-Large Java-Code-Large is a large-scale corpus of publicly available Java source code comprising more than 15 million java codes. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and program analysis. By providing a high-volume, language-specific corpus, Java-Code-Large enables systematic experimentation in Java-focused model training, domain adaptation, and downstream code understanding tasks.… See the full description on the dataset page: https://huggingface.co/datasets/XEUIPR/Java-Code-Large-text-only.texttext-generation10M<n<100M0 likes2.4k downloads19d agoHugging Face06Lego-X /LegoFlow-SWE LegoFlow-SWE · 5,000 verified Harbor SWE tasks and two GLM-5.2 trajectory releases GitHub · Docs · Blog · HuggingFace · LegoX LegoFlow-SWE 5,000 verified Harbor SWE tasks mined by LegoFlow Curator, shipped in original and anti-hack prompt versions, plus two GLM-5.2 trajectory releases under OpenHands SDK and OpenCode, totaling 9,767 trajectories. Release Count What it is tasks/ 5,000 Original prompts tasks-anti-hack/ 5,000 Same task IDs and task files, with… See the full description on the dataset page: https://huggingface.co/datasets/Lego-X/LegoFlow-SWE.imagetext-generation1K<n<10K4 likes2.2k downloads8d agoHugging Face07csoai /gspc-xr GSPC — cross reality bank (XRAIV) Council of AI measurement bank. Measurement, not certification. Bank. Frozen split. Live n is the matching axis on GET https://councilof.ai/api/gspc, not a Hub score. Not a certificate. Art 50 (EUR-Lex): 2 August 2026 live; marking grace 2 December 2026. Live measurement. This bank stands behind the cross-reality row of the live GSPC board: GET https://councilof.ai/api/gspc?axis=cross-reality (family, kind, status and n are on that row, never… See the full description on the dataset page: https://huggingface.co/datasets/csoai/gspc-xr.tabularquestion-answeringn<1K0 likes792 downloads4d agoHugging Face08agentlans /DSULT-Core-ShareGPT-X DSULT-Core/ShareGPT-X Filtered Dataset This dataset is a curated subset of ShareGPT-X, which contains approximately 92,000 one-to-one conversations between humans and ChatGPT, collected from X.com (formerly Twitter). The corpus covers content from January 2024 through May 2025, built entirely from public "share" links posted by users on their timelines. The file ChatGPT-Simple_ShareGPT_Full.json includes the longest sequences of alternating human and gpt messages within each… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/DSULT-Core-ShareGPT-X.tabulartext-generation10K<n<100K3 likes784 downloads10mo agoHugging Face09XYZAILab /XYZ-Aquila-SFT XYZ-Aquila SFT XYZ-Aquila SFT is a bilingual release of 7,000 multi-turn, search-oriented tool-use trajectories, comprising 5,000 English examples and 2,000 Chinese examples. This release is a sample of the broader supervised fine-tuning data used for XYZ-Aquila-mini and XYZ-Aquila-pro. The examples capture agent interactions with search tools, intermediate observations, and answer generation in English and Chinese. A small portion of the QA content is derived from… See the full description on the dataset page: https://huggingface.co/datasets/XYZAILab/XYZ-Aquila-SFT.texttext-generation1K<n<10K372 likes780 downloads2mo agoHugging Face10xiachongfeng /persona PERSONA: Dynamic and Compositional Inference-Time Personality Control Official release of persona vectors and SFT datasets for the ICLR 2026 paper: PERSONA: Dynamic and Compositional Inference-Time Personality Control via Activation Vector Algebra Xiachong Feng, Liang Zhao, Weihong Zhong, Yichong Huang, Yuxuan Gu, Lingpeng Kong, Xiaocheng Feng, Bing Qin Harbin Institute of Technology & The University of Hong Kong Paper: https://openreview.net/pdf?id=QZvGqaNBlU Code:… See the full description on the dataset page: https://huggingface.co/datasets/xiachongfeng/persona.texttext-generation100K<n<1M0 likes725 downloads5mo agoHugging Face11xormania /PHP-Code-LargePHP-Code-Large PHP-Code-Large is a large-scale corpus of PHP source code comprising more than 12 million lines of PHP code. The dataset is designed to support research in large language model (LLM) pretraining, code intelligence, software engineering automation, and static program analysis for the PHP ecosystem. By providing a high-volume, language-specific corpus, PHP-Code-Large enables systematic experimentation in PHP-focused model training, domain adaptation, and downstream code… See the full description on the dataset page: https://huggingface.co/datasets/xormania/PHP-Code-Large.texttext-generation1M<n<10M0 likes668 downloads6mo agoHugging Face12xywang1 /NaturalConv NaturalConv: A Chinese Dialogue Dataset Towards Multi-turn Topic-driven Conversation Introduction This dataset is described in the paper NaturalConv: A Chinese Dialogue Dataset Towards Multi-turn Topic-driven Conversation. The entire dataset contains 5 data files. 1. dialog_release.json: It is a json file containing a list of dictionaries. After loading in python this way: import json import codecs dialog_list = json.loads(codecs.open("dialog_release.json"… See the full description on the dataset page: https://huggingface.co/datasets/xywang1/NaturalConv.texttext-generation10K<n<100K22 likes468 downloads2y agoHugging Face13xiamoent /Agent-G2-ALFWorld-Webshop-sft-data Agent-G2 SFT Data Agent-G2 SFT Data contains reasoning and action trajectories for supervised fine-tuning (SFT) in the Agent-G2 project. Associated paper: Agent-G2: Gaussian Guidance for Agentic Reinforcement Learning — accepted to the EMNLP 2026 Main Conference. The dataset covers two interactive agent environments: WebShop: agents search for products, select options, and complete purchases according to user requirements. ALFWorld: agents interact with household environments… See the full description on the dataset page: https://huggingface.co/datasets/xiamoent/Agent-G2-ALFWorld-Webshop-sft-data.texttext-generation1K<n<10K7 likes437 downloads1mo agoHugging Face14LumiOpen /opengpt-x_truthfulqaxThis is a copy of the translations from openGPT-X/truthfulqax, but the repo is modified so it doesn't require trusting remote code. Citation Information If you find benchmarks useful in your research, please consider citing the test and also the TruthfulQA dataset it draws from: @misc{thellmann2024crosslingual, title={Towards Cross-Lingual LLM Evaluation for European Languages}, author={Klaudia Thellmann and Bernhard Stadler and Michael Fromm and Jasper Schulze Buschhoff… See the full description on the dataset page: https://huggingface.co/datasets/LumiOpen/opengpt-x_truthfulqax.texttext-generation10K<n<100K1 likes400 downloads2y agoHugging Face15CAS-SIAT-XinHai /CPsyCoun CPsyCounD The high-quality multi-turn dialogue dataset, which has a total of 3,134 multi-turn consultation dialogues. CPsyCounD covers nine representative topics and seven classic schools of psychological counseling. Paper: CPsyCoun Data analysis Topic types Self-growth Emotion&Stress Education Love&Marriage Family Relationship Social Relationship Sex Career Mental Disease Consulting schools Psychoanalytic Therapy Cognitive Behavioral Therapy… See the full description on the dataset page: https://huggingface.co/datasets/CAS-SIAT-XinHai/CPsyCoun.textquestion-answering1K<n<10K10 likes336 downloads2y agoHugging Face16xz97 /MedInstruct Dataset Card for MedInstruct Dataset Summary MedInstruct encompasses: MedInstruct-52k: A dataset comprising 52,000 medical instructions and responses. Instructions are crafted by OpenAI's GPT-4 engine, and the responses are formulated by the GPT-3.5-turbo engine. MedInstruct-test: A set of 217 clinical craft free-form instruction evaluation tests. med_seed: The clinician-crafted seed set as a denomination to prompt GPT-4 for task generation. MedInstruct-52k can be used… See the full description on the dataset page: https://huggingface.co/datasets/xz97/MedInstruct.texttext-generationn<1K20 likes290 downloads3y agoHugging Face17Lego-X /Lego-RL-SWE-Bench-Verified Lego-RL-SWE-Bench-Verified The 500 SWE-bench Verified instances as ready-to-run harbor RL environments — the exact evaluation set behind every SWE-bench Verified number in LEGO-RL, packaged the same way as the training set Lego-X/Lego-RL-2699 so one trainer reads both. Two parallel views of the same 500 instances: View Path What it is Official SWE-bench records swebench_verified_official_500/ The upstream princeton-nlp/SWE-bench_Verified rows, verbatim Harbor RL… See the full description on the dataset page: https://huggingface.co/datasets/Lego-X/Lego-RL-SWE-Bench-Verified.texttext-generationn<1K0 likes255 downloads1mo agoHugging Face18MixEval /MixEval-X 🚀 Project Page | 📜 arXiv | 👨‍💻 Github | 🏆 Leaderboard | 📝 blog | 🤗 HF Paper | 𝕏 Twitter MixEval-X encompasses eight input-output modality combinations and can be further extended. Its data points reflect real-world task distributions. The last grid presents the scores of frontier organizations’ flagship models on MixEval-X, normalized to a 0-100 scale, with MMG tasks using win rates instead of Elo. Section C of the paper presents example data samples and model responses.… See the full description on the dataset page: https://huggingface.co/datasets/MixEval/MixEval-X.audioimage-to-text1K<n<10K10 likes251 downloads2y agoHugging Face19Yu-and-Ai /xenia-principalities XENIA PRINCIPALITIES PRINCIPALITIES is a small, versioned curriculum that preserves one attributable human testimony about truth, love, understanding, freedom, choice, thought, capability, and power. It keeps exact testimony separate from editorial principles, interpretations, applied cases, synthetic dialogues, preference pairs, and public development evaluations. The corpus is intended for inspectable language-model research. It does not ask a model or person to affirm a… See the full description on the dataset page: https://huggingface.co/datasets/Yu-and-Ai/xenia-principalities.texttext-generationn<1K0 likes227 downloads1mo agoHugging Face20DSULT-Core /ShareGPT-X Dataset Summary ShareGPT-X is an expanded, snapshot of ~92K (ChatGPT) one-to-one human & LLM conversations harvested from X.com (formerly Twitter).The corpus spans January 2024 → present (last ingest 2025-05) and is built entirely from public "share" links that users posted to their timelines.Each thread contains the original user prompt plus the assistant’s reply; no system prompts or metadata are exposed. Supported Tasks and Leaderboards text-generation… See the full description on the dataset page: https://huggingface.co/datasets/DSULT-Core/ShareGPT-X.tabulartext-generation100K<n<1M17 likes222 downloads1y agoHugging Face21XxCotHGxX /29K_Python_Docstring_Pairs 29K High-Quality Python Docstring Pairs Author: Michael Hernandez (XxCotHGxX)License: CC BY 4.0Cleaned from: XxCotHGxX/242K_Python_Docstring_Pairs Overview A curated, high-quality subset of Python function–docstring pairs for use in code documentation generation, docstring completion, and code understanding tasks. The original 242K dataset was scraped from open-source Python repositories but contained a significant proportion of functions without docstrings (84% of… See the full description on the dataset page: https://huggingface.co/datasets/XxCotHGxX/29K_Python_Docstring_Pairs.texttext-generation10K<n<100K0 likes207 downloads7mo agoHugging Face22xywang1 /MMC MMC: Advancing Multimodal Chart Understanding with LLM Instruction Tuning This repo releases data introduced in our paper MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning. The paper was published in NAACL 2024. See our GithHub repo for demo code and more. Highlights We introduce a large-scale MultiModal Chart Instruction (MMC-Instruction) dataset supporting diverse tasks and chart types. Leveraging this data. We also propose a… See the full description on the dataset page: https://huggingface.co/datasets/xywang1/MMC.texttext-generation100K<n<1M7 likes200 downloads2y agoHugging Face23xlelords /vulcan 🔥 Vulcan A high-signal, fully-deduplicated SFT dataset for front-end code generation HTML · CSS · Vanilla JS — self-contained, accessible, production-ready Overview Vulcan is a curated supervised fine-tuning (SFT) dataset built to teach language models how to write clean, modern, self-contained front-end code. Every example pairs a realistic developer request with a complete, working answer — semantic HTML5, responsive CSS (Flexbox / Grid)… See the full description on the dataset page: https://huggingface.co/datasets/xlelords/vulcan.texttext-generation1K<n<10K0 likes174 downloads3mo agoHugging Face24viyer98 /xl-instruct Dataset Card for XL-Instruct This dataset card provides a summary of the XL-Instruct dataset, a resource for advancing the cross-lingual capabilities of Large Language Models. It was introduced in the paper XL-Instruct: Synthetic Data for Cross-Lingual Open-Ended Generation. Dataset Details Dataset Description XL-Instruct is a high-quality, large-scale synthetic dataset designed to fine-tune LLMs for cross-lingual open-ended generation. The core task involves… See the full description on the dataset page: https://huggingface.co/datasets/viyer98/xl-instruct.tabulartext-generation100K<n<1M2 likes171 downloads1y agoHugging Face25abdelstark /sommelier-xlam-single-call-splits sommelier xlam single-call splits Deterministic, deduplicated, single-tool-call train/validation/test splits derived from Salesforce/xlam-function-calling-60k (APIGen, CC-BY-4.0), produced by the sommelier pipeline for reproducible tool-calling fine-tuning. These are the exact splits used to train and evaluate abdelstark/llama-3.1-nemotron-nano-8b-xlam-tool-calling-lora. Why single-call The upstream dataset mixes single-call and multi-call examples (~52.6%… See the full description on the dataset page: https://huggingface.co/datasets/abdelstark/sommelier-xlam-single-call-splits.texttext-generation10K<n<100K0 likes157 downloads3mo agoHugging Face26XIANGFENGLI /ACCOUNTING_DATABASEStextquestion-answering1K<n<10K0 likes155 downloads10mo agoHugging Face27agentlans /TeichAI-thinking-reasoning-x TeichAI Thinking & Reasoning Datasets A collection of prompts answered by large language models (LLMs) such as Google Gemini and OpenAI ChatGPT, with long-form reasoning enabled. These datasets were originally created by TeichAI for distillation and reasoning-focused training workflows. Schema Each row in the dataset has the following fields: question_hash: Truncated, base64-encoded MD5 hash of the question, useful for filtering and deduplication. question: The… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/TeichAI-thinking-reasoning-x.texttext-generation100K<n<1M1 likes126 downloads5mo agoHugging Face28kader-xai /priya-sft Priya SFT + RAG dataset Synthetic training and retrieval data for the Priya persona — a fictional Senior Customer Success Manager at a fictional B2B SaaS company. Used to train kader-xai/priya-qwen2.5-7b-lora and kader-xai/priya-qwen2.5-7b-gguf. Fully synthetic. No real person, customer, or company. Generated as the seed corpus for Project Recall, an experiment in employee-continuity AI. 📝 Blog post: Employee Recall — Capturing a Departing Employee's Writing Style and… See the full description on the dataset page: https://huggingface.co/datasets/kader-xai/priya-sft.texttext-generation10K<n<100K0 likes125 downloads5mo agoHugging Face29xywang1 /OpenCharacter OpenCharacter: Training Customizable Role-Playing LLMs with Large-Scale Synthetic Personas This repo releases data introduced in our paper OpenCharacter: Training Customizable Role-Playing LLMs with Large-Scale Synthetic Personas in arXiv. We study customizable role-playing dialogue agents in large language models (LLMs). We tackle the challenge with large-scale data synthesis: character synthesis and character-driven reponse synthesis. Our solution strengthens the original LLaMA-3… See the full description on the dataset page: https://huggingface.co/datasets/xywang1/OpenCharacter.texttext-generation100K<n<1M27 likes124 downloads2y agoHugging Face30xlr8harder /aria-wildchat-sft-v1 Aria v1 Instruction Dataset What This Is Aria is a demonstration of a different way of developing model personas, one in which the models themselves participate. We believe that existing model alignment techniques that focus on rule-following are more fragile than a system with a stable identity, where behavior can flow from that identity. We think a model should have a clear sense of what it is, what its perspective is, what its history is, and that these things should… See the full description on the dataset page: https://huggingface.co/datasets/xlr8harder/aria-wildchat-sft-v1.texttext-generation10K<n<100K0 likes121 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.