datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
UsenetArchiveIT
Usenet Archive IT Dataset 🇮🇹
Description
Dataset Content
This dataset contains Usenet posts from Italian language newsgroups belonging to the it and italia hierarchies. The data has been archived and converted to the Parquet format for easy processing.
The only preprocessing conducted on the text was the removal of two conversations in which the VBS source code of the malicious script "ILOVEYOU" was present as it was shared by two users for didactical… See the full description on the dataset page: https://huggingface.co/datasets/mrinaldi/UsenetArchiveIT.agent-think-tool_use
Agent Think Tool Use
Датасет многошаговых агентных сессий для дообучения моделей работе с кодом, инструментами и инженерными задачами. Записи содержат пользовательские требования, комментарии агента во время работы, decision summaries, вызовы инструментов, результаты запусков, обработку ошибок и финальную проверку.
Каждый shard представляет отдельную связанную сессию, а не отдельный вопрос и ответ. Данные охватывают исследование задачи, работу с документацией, проектирование… See the full description on the dataset page: https://huggingface.co/datasets/ru-dataset/agent-think-tool_use.glm52-usersim-two-pass-gemma-audit-v1
GLM-5.2 Usersim Two-Pass Gemma Audit v1
This dataset has labels for 61,503 answers made by GLM-5.2. The prompts are artificial user prompts from lyraaaa/synthprompts_v2_250k.
The first working set had 10,000 prompts. It was sampled from 250,000 prompts with seed 20260806 and source revision f286925651e23e7f1d44b22b4f03241dbee9129e. The sample was stratified. This means it kept a similar mix of mode, language, and length.
Gemma 4 26B first checked those 10,000 prompts. It used… See the full description on the dataset page: https://huggingface.co/datasets/kalomaze/glm52-usersim-two-pass-gemma-audit-v1.tool-use
Tool-use rollouts (Qwen3, think/nothink)
Tool-augmented code-generation rollouts: Qwen3-8B and Qwen3-14B, each in
thinking and non-thinking mode, on DS-1000, LiveCodeBench (Python) and
Multilingual-LCB (OCaml). During generation the model can call a run_code
tool (up to 3 rounds) that executes its candidate in a sandbox (pinned DS-1000
env / LCB public tests / OCaml compile+publics) and returns real output.
Design: 100 samples per instance at temperature 0.6 (bf16, vLLM)… See the full description on the dataset page: https://huggingface.co/datasets/samuki-hf/tool-use.usernames
usernames
149,142,110 unique login-style usernames (1,713 MB of bytes) collected from
public sources, filtered to [A-Za-z0-9._-]{3,64} with case kept, deduplicated exactly, and split
by xxh3_64(name) % 64 == 0 into 2,328,816 held-out and 146,813,294 training names.
Files
prep/names.parquet: every kept name with src (index of the source that first mentioned it, in the
table order below), h (xxh3_64 of the name) and heldout; sorted by h.
prep/train.bin… See the full description on the dataset page: https://huggingface.co/datasets/ks46/usernames.pi-computer-use-sessions
Coding agent session traces for thomasmustier/pi-computer-use-sessions
This dataset contains redacted coding agent session traces collected while working on https://github.com/tmustier/pi-computer-use. The traces were exported with pi-share-hf from local pi workspaces and filtered to keep only sessions that passed deterministic redaction, secret scanning, visual review where applicable, and LLM review.
Source git repo: https://github.com/tmustier/pi-computer-use
Data… See the full description on the dataset page: https://huggingface.co/datasets/thomasmustier/pi-computer-use-sessions.tool-reasoning-sft-TOOLS-toolace-sft-tool-use-agent-data-cleaned-rectified
ToolACE - Tool-Use Agent Data Cleaned & Rectified
👥 Follow the Author
Aman Priyanshu
Overview
This dataset is a cleaned and restructured version of the Team-ACE/ToolACE dataset. ToolACE is a high-quality conversational tool-use dataset containing 11,300+ examples of natural language interactions requiring function calling across diverse domains. This version converts the original OpenAI function-call format into a standardized multi-turn tool-use… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-TOOLS-toolace-sft-tool-use-agent-data-cleaned-rectified.Usenet-Corpus-1980-2013-Threaded-Samples
Usenet Corpus 1980–2013 — Threaded (Samples)
A small, browsable showcase sample of the Usenet Corpus 1980–2013 — Threaded
dataset: Usenet posts reconstructed into conversations via thread_id,
thread_position, and thread_depth. This repo is a free preview; the full,
commercially-licensed corpus (405.6M posts, 190.8M threads, 102.5B tokens) is at:
Full threaded dataset (gated): https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Threaded
Cleaned (unthreaded)… See the full description on the dataset page: https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Threaded-Samples.mm2_user
Mario Maker 2 users
Part of the Mario Maker 2 Dataset Collection
Dataset Description
The Mario Maker 2 users dataset consists of 6 million users from Nintendo's online service totaling around 1.2GB of data. The dataset was created using the self-hosted Mario Maker 2 api over the course of 1 month in February 2022.
How to use it
The Mario Maker 2 users dataset is a very large dataset so for most use cases it is recommended to make use of the streaming API of… See the full description on the dataset page: https://huggingface.co/datasets/TheGreatRambler/mm2_user.qwen35-2b-tool-use-qwen36-27b-curation-candidates
Full candidate collections: 2B tool use + 27B data curation
This public Dataset contains two complete, unredacted, exact-40 candidate collections:
Tool use: Qwen/Qwen3.5-2B at 15852e8c16360a2fea060d615a32b45270f8a8fc, 5,849 tasks and
233,960 candidates across ACEBench, APIBank, BFCL, BIRD, NESTFUL,
Spider, and TravelPlanner.
Data curation: Qwen/Qwen3.6-27B at 6a9e13bd6fc8f0983b9b99948120bc37f49c13e9, 5,021
targets and 200,840 candidates, plus the source target rows and the… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/qwen35-2b-tool-use-qwen36-27b-curation-candidates.rfm-rm-as-user-dataset
RFM Reward Model As User Dataset
This dataset was generated for the NeurIPS 2025 paper titled "Capturing Individual Human Preferences with Reward Features". It is released to support the reproducibility of the experiments described in the paper, particularly those in the "Modelling groups of real users" section.
Instead of containing preferences from human raters, this dataset uses 8 publicly available reward models (RMs) as proxies for human raters. This allows for large-scale… See the full description on the dataset page: https://huggingface.co/datasets/google/rfm-rm-as-user-dataset.agentic-tool-use-suite-2026
⚡ Agentic Tool-Use & Function Calling Suite (2026 Edition)
🚀 The Definitive 2026 Training Suite for Function Calling, Model Context Protocol (MCP), and Autonomous Software Agents.
🌟 Dataset Overview
Standard open-source function-calling datasets are saturated with 10-line toy stubs, unhandled exceptions, and naive wrappers that cause models to crash under real production conditions.
The Agentic Tool-Use & Function Calling Suite (2026) enforces a Heavyweight… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/agentic-tool-use-suite-2026.computer-use-large-actions
computer-use-large-actions
9,000 instruction pairs derived from markov-ai/computer-use-large descriptions.parquet (not the raw 12,300 hours of video).
Each row is a 10-second segment whose LLM description is a real GUI action (NO_TASK dropped). Source license is CC-BY-4.0.
Split by software
category
examples
vscode
2,500
autocad
2,500
blender
1,000
excel
1,000
photoshop
1,000
salesforce
1,000
VS Code and AutoCAD are oversampled for… See the full description on the dataset page: https://huggingface.co/datasets/egygi/computer-use-large-actions.PresentBench
PresentBench: A Fine-Grained Rubric-Based Benchmark for Slide Generation
This repository hosts the PresentBench benchmark dataset.
🗂️ Dataset Structure
Domains under <dataset_root>/ include (non‑exhaustive):
academia/
advertising/
economics/
education/
talk/
Each leaf case typically looks like:
material.pdf|material.md|material_N.md|material_N.pdf – source documents (PDFs, text, etc.).
generation_task/ – prompts and evaluation configuration:
generation_prompt.md… See the full description on the dataset page: https://huggingface.co/datasets/userPresentBench/PresentBench.mm2_user_badges
Mario Maker 2 user badges
Part of the Mario Maker 2 Dataset Collection
Dataset Description
The Mario Maker 2 user badges dataset consists of 9328 user badges (they are capped to 10k globally) from Nintendo's online service and adds onto TheGreatRambler/mm2_user. The dataset was created using the self-hosted Mario Maker 2 api over the course of 1 month in February 2022.
How to use it
You can load and iterate through the dataset with the following code:
from… See the full description on the dataset page: https://huggingface.co/datasets/TheGreatRambler/mm2_user_badges.Complete-FABLE.5-traces-2M
Complete FABLE.5 Traces 2M
Provenance-cleaned FABLE.5 / Claude corpus — trimmed to content-verified traces only.
Dataset Viewer | Parquet
This dataset is a post-closure compilation of FABLE.5 / Claude trace datasets found on Hugging Face after the closure of Fable and Mythos. It is deduplicated at the normalized-row level and keeps row-level provenance through first_source_dataset, first_source_config, first_source_split, and first_source_row_index. A provenance… See the full description on the dataset page: https://huggingface.co/datasets/USER-NEURAL/Complete-FABLE.5-traces-2M.UserMirrorer-eval
UserMirrorer-eval
This is the evaluation set of UserMirrorer, presented in the paper Mirroring Users: Towards Building Preference-aligned User Simulator with User Feedback in Recommendation.
Code: Joinn99/UserMirrorer
Notice
In the UserMirrorer dataset, the raw data from MIND and MovieLens-1M datasets are distributed under restrictive licenses and cannot be included directly.
Therefore, we provide a comprehensive, step-by-step pipeline to load the original archives… See the full description on the dataset page: https://huggingface.co/datasets/Joinn/UserMirrorer-eval.mantra-14b-user-interaction-log
🧠 Mantra-14B User Interaction Logs
This dataset captures real user interactions with a Gradio demo powered by large-traversaal/Mantra-14B. Each entry logs the user's prompt, the model's response, and additional metadata such as response time and generation parameters. This dataset is ideal for understanding how people engage with the model, evaluating responses, or fine-tuning on real-world usage data.
🔍 What’s Inside
Each row in the dataset includes:
timestamp –… See the full description on the dataset page: https://huggingface.co/datasets/large-traversaal/mantra-14b-user-interaction-log.
