CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01mrinaldi /UsenetArchiveIT Usenet Archive IT Dataset 🇮🇹 Description Dataset Content This dataset contains Usenet posts from Italian language newsgroups belonging to the it and italia hierarchies. The data has been archived and converted to the Parquet format for easy processing. The only preprocessing conducted on the text was the removal of two conversations in which the VBS source code of the malicious script "ILOVEYOU" was present as it was shared by two users for didactical… See the full description on the dataset page: https://huggingface.co/datasets/mrinaldi/UsenetArchiveIT.tabulartext-generation10M<n<100M11 likes1.3k downloads2y agoHugging Face02ru-dataset /agent-think-tool_use Agent Think Tool Use Датасет многошаговых агентных сессий для дообучения моделей работе с кодом, инструментами и инженерными задачами. Записи содержат пользовательские требования, комментарии агента во время работы, decision summaries, вызовы инструментов, результаты запусков, обработку ошибок и финальную проверку. Каждый shard представляет отдельную связанную сессию, а не отдельный вопрос и ответ. Данные охватывают исследование задачи, работу с документацией, проектирование… See the full description on the dataset page: https://huggingface.co/datasets/ru-dataset/agent-think-tool_use.tabulartext-generationn<1K2 likes789 downloads4d agoHugging Face03kalomaze /glm52-usersim-two-pass-gemma-audit-v1 GLM-5.2 Usersim Two-Pass Gemma Audit v1 This dataset has labels for 61,503 answers made by GLM-5.2. The prompts are artificial user prompts from lyraaaa/synthprompts_v2_250k. The first working set had 10,000 prompts. It was sampled from 250,000 prompts with seed 20260806 and source revision f286925651e23e7f1d44b22b4f03241dbee9129e. The sample was stratified. This means it kept a similar mix of mode, language, and length. Gemma 4 26B first checked those 10,000 prompts. It used… See the full description on the dataset page: https://huggingface.co/datasets/kalomaze/glm52-usersim-two-pass-gemma-audit-v1.tabulartext-generation100K<n<1M4 likes776 downloads1mo agoHugging Face04samuki-hf /tool-use Tool-use rollouts (Qwen3, think/nothink) Tool-augmented code-generation rollouts: Qwen3-8B and Qwen3-14B, each in thinking and non-thinking mode, on DS-1000, LiveCodeBench (Python) and Multilingual-LCB (OCaml). During generation the model can call a run_code tool (up to 3 rounds) that executes its candidate in a sandbox (pinned DS-1000 env / LCB public tests / OCaml compile+publics) and returns real output. Design: 100 samples per instance at temperature 0.6 (bf16, vLLM)… See the full description on the dataset page: https://huggingface.co/datasets/samuki-hf/tool-use.tabulartext-generation1M<n<10M1 likes411 downloads2mo agoHugging Face05ks46 /usernames usernames 149,142,110 unique login-style usernames (1,713 MB of bytes) collected from public sources, filtered to [A-Za-z0-9._-]{3,64} with case kept, deduplicated exactly, and split by xxh3_64(name) % 64 == 0 into 2,328,816 held-out and 146,813,294 training names. Files prep/names.parquet: every kept name with src (index of the source that first mentioned it, in the table order below), h (xxh3_64 of the name) and heldout; sorted by h. prep/train.bin… See the full description on the dataset page: https://huggingface.co/datasets/ks46/usernames.tabulartext-generation100M<n<1B0 likes229 downloads4d agoHugging Face06thomasmustier /pi-computer-use-sessions Coding agent session traces for thomasmustier/pi-computer-use-sessions This dataset contains redacted coding agent session traces collected while working on https://github.com/tmustier/pi-computer-use. The traces were exported with pi-share-hf from local pi workspaces and filtered to keep only sessions that passed deterministic redaction, secret scanning, visual review where applicable, and LLM review. Source git repo: https://github.com/tmustier/pi-computer-use Data… See the full description on the dataset page: https://huggingface.co/datasets/thomasmustier/pi-computer-use-sessions.tabulartext-generationn<1K0 likes146 downloads3mo agoHugging Face07AmanPriyanshu /tool-reasoning-sft-TOOLS-toolace-sft-tool-use-agent-data-cleaned-rectified ToolACE - Tool-Use Agent Data Cleaned & Rectified 👥 Follow the Author Aman Priyanshu Overview This dataset is a cleaned and restructured version of the Team-ACE/ToolACE dataset. ToolACE is a high-quality conversational tool-use dataset containing 11,300+ examples of natural language interactions requiring function calling across diverse domains. This version converts the original OpenAI function-call format into a standardized multi-turn tool-use… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-TOOLS-toolace-sft-tool-use-agent-data-cleaned-rectified.tabulartext-generation10K<n<100K0 likes132 downloads7mo agoHugging Face08OwnedByDanes /Usenet-Corpus-1980-2013-Threaded-Samples Usenet Corpus 1980–2013 — Threaded (Samples) A small, browsable showcase sample of the Usenet Corpus 1980–2013 — Threaded dataset: Usenet posts reconstructed into conversations via thread_id, thread_position, and thread_depth. This repo is a free preview; the full, commercially-licensed corpus (405.6M posts, 190.8M threads, 102.5B tokens) is at: Full threaded dataset (gated): https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Threaded Cleaned (unthreaded)… See the full description on the dataset page: https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Threaded-Samples.tabulartext-generation10K<n<100K0 likes93 downloads12d agoHugging Face09TheGreatRambler /mm2_user Mario Maker 2 users Part of the Mario Maker 2 Dataset Collection Dataset Description The Mario Maker 2 users dataset consists of 6 million users from Nintendo's online service totaling around 1.2GB of data. The dataset was created using the self-hosted Mario Maker 2 api over the course of 1 month in February 2022. How to use it The Mario Maker 2 users dataset is a very large dataset so for most use cases it is recommended to make use of the streaming API of… See the full description on the dataset page: https://huggingface.co/datasets/TheGreatRambler/mm2_user.tabularother1M<n<10M3 likes90 downloads4y agoHugging Face10asingh15 /qwen35-2b-tool-use-qwen36-27b-curation-candidates Full candidate collections: 2B tool use + 27B data curation This public Dataset contains two complete, unredacted, exact-40 candidate collections: Tool use: Qwen/Qwen3.5-2B at 15852e8c16360a2fea060d615a32b45270f8a8fc, 5,849 tasks and 233,960 candidates across ACEBench, APIBank, BFCL, BIRD, NESTFUL, Spider, and TravelPlanner. Data curation: Qwen/Qwen3.6-27B at 6a9e13bd6fc8f0983b9b99948120bc37f49c13e9, 5,021 targets and 200,840 candidates, plus the source target rows and the… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/qwen35-2b-tool-use-qwen36-27b-curation-candidates.tabulartext-generation100K<n<1M0 likes75 downloads1mo agoHugging Face11google /rfm-rm-as-user-dataset RFM Reward Model As User Dataset This dataset was generated for the NeurIPS 2025 paper titled "Capturing Individual Human Preferences with Reward Features". It is released to support the reproducibility of the experiments described in the paper, particularly those in the "Modelling groups of real users" section. Instead of containing preferences from human raters, this dataset uses 8 publicly available reward models (RMs) as proxies for human raters. This allows for large-scale… See the full description on the dataset page: https://huggingface.co/datasets/google/rfm-rm-as-user-dataset.tabulartext-generation10K<n<100K10 likes70 downloads11mo agoHugging Face12beatsprom /agentic-tool-use-suite-2026 ⚡ Agentic Tool-Use & Function Calling Suite (2026 Edition) 🚀 The Definitive 2026 Training Suite for Function Calling, Model Context Protocol (MCP), and Autonomous Software Agents. 🌟 Dataset Overview Standard open-source function-calling datasets are saturated with 10-line toy stubs, unhandled exceptions, and naive wrappers that cause models to crash under real production conditions. The Agentic Tool-Use & Function Calling Suite (2026) enforces a Heavyweight… See the full description on the dataset page: https://huggingface.co/datasets/beatsprom/agentic-tool-use-suite-2026.tabulartext-generation1K<n<10K0 likes65 downloads11d agoHugging Face13egygi /computer-use-large-actions computer-use-large-actions 9,000 instruction pairs derived from markov-ai/computer-use-large descriptions.parquet (not the raw 12,300 hours of video). Each row is a 10-second segment whose LLM description is a real GUI action (NO_TASK dropped). Source license is CC-BY-4.0. Split by software category examples vscode 2,500 autocad 2,500 blender 1,000 excel 1,000 photoshop 1,000 salesforce 1,000 VS Code and AutoCAD are oversampled for… See the full description on the dataset page: https://huggingface.co/datasets/egygi/computer-use-large-actions.tabulartext-generation1K<n<10K0 likes58 downloads25d agoHugging Face14userPresentBench /PresentBench PresentBench: A Fine-Grained Rubric-Based Benchmark for Slide Generation This repository hosts the PresentBench benchmark dataset. 🗂️ Dataset Structure Domains under <dataset_root>/ include (non‑exhaustive): academia/ advertising/ economics/ education/ talk/ Each leaf case typically looks like: material.pdf|material.md|material_N.md|material_N.pdf – source documents (PDFs, text, etc.). generation_task/ – prompts and evaluation configuration: generation_prompt.md… See the full description on the dataset page: https://huggingface.co/datasets/userPresentBench/PresentBench.documentany-to-anyn<1K0 likes56 downloads5mo agoHugging Face15TheGreatRambler /mm2_user_badges Mario Maker 2 user badges Part of the Mario Maker 2 Dataset Collection Dataset Description The Mario Maker 2 user badges dataset consists of 9328 user badges (they are capped to 10k globally) from Nintendo's online service and adds onto TheGreatRambler/mm2_user. The dataset was created using the self-hosted Mario Maker 2 api over the course of 1 month in February 2022. How to use it You can load and iterate through the dataset with the following code: from… See the full description on the dataset page: https://huggingface.co/datasets/TheGreatRambler/mm2_user_badges.tabularother1K<n<10K1 likes42 downloads4y agoHugging Face16USER-NEURAL /Complete-FABLE.5-traces-2M Complete FABLE.5 Traces 2M Provenance-cleaned FABLE.5 / Claude corpus — trimmed to content-verified traces only. Dataset Viewer | Parquet This dataset is a post-closure compilation of FABLE.5 / Claude trace datasets found on Hugging Face after the closure of Fable and Mythos. It is deduplicated at the normalized-row level and keeps row-level provenance through first_source_dataset, first_source_config, first_source_split, and first_source_row_index. A provenance… See the full description on the dataset page: https://huggingface.co/datasets/USER-NEURAL/Complete-FABLE.5-traces-2M.tabulartext-generation10K<n<100K0 likes40 downloads2mo agoHugging Face17Joinn /UserMirrorer-eval UserMirrorer-eval This is the evaluation set of UserMirrorer, presented in the paper Mirroring Users: Towards Building Preference-aligned User Simulator with User Feedback in Recommendation. Code: Joinn99/UserMirrorer Notice In the UserMirrorer dataset, the raw data from MIND and MovieLens-1M datasets are distributed under restrictive licenses and cannot be included directly. Therefore, we provide a comprehensive, step-by-step pipeline to load the original archives… See the full description on the dataset page: https://huggingface.co/datasets/Joinn/UserMirrorer-eval.tabulartext-generation1K<n<10K0 likes37 downloads5mo agoHugging Face18large-traversaal /mantra-14b-user-interaction-log 🧠 Mantra-14B User Interaction Logs This dataset captures real user interactions with a Gradio demo powered by large-traversaal/Mantra-14B. Each entry logs the user's prompt, the model's response, and additional metadata such as response time and generation parameters. This dataset is ideal for understanding how people engage with the model, evaluating responses, or fine-tuning on real-world usage data. 🔍 What’s Inside Each row in the dataset includes: timestamp –… See the full description on the dataset page: https://huggingface.co/datasets/large-traversaal/mantra-14b-user-interaction-log.tabulartext-generationn<1K0 likes12 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.