CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01nvidia /Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1 Dataset Description: We created an RL dataset for conversational tool-use by utilizing existing expert tool-use trajectories. We pose each assistant step of the trajectory as a separate behavior cloning problem where the policy model is incentivized to match the tool call choices of the expert model. Each trajectory includes the use of tools for authentication, data lookup, servicing (i.e. booking reservations, changing them, getting discounts, etc), and more across 838 different… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1.tabular10K<n<100K32 likes1.4k downloads7mo agoHugging Face02ehejin /user_study-preference-personalized_0423_base_filtered Filtered user study dataset Source repo: ehejin/user_study-preference-personalized_0423_base Each row is ONE item review (pre-rating, conversation, post-rating). Submission-level fields (prolific_pid, demographics, background) are duplicated across rows that share a submission. The 25-50 rows here are the FIRST review for each unique pool index, selected the same way the analysis plot uses — see scripts/plot_vote_shift_3way.py. Total rows: 50 tabularn<1K0 likes226 downloads5mo agoHugging Face03rafmacalaba /data-use-ner Data-use-ner (human holdout) GLiNER-format human-adjudicated holdout: 473 spans — annotator190 (190, origin=fcv_pads_east_africa) + jdc283 (283, origin=jdc_operational). Never trained on. Source: rafmacalaba/datause-displacement-reviewed holdout (gliner_reviewed token spans + readable_reviewed passages, v2.4 labels) with v3 probe head_score (outputs/gliner_datause_v3_probe_human473.jsonl). Columns text (full passage = " ".join(tokenized_text); span char offsets… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/data-use-ner.tabulartoken-classification10K<n<100K0 likes214 downloads13d agoHugging Face04thomasmustier /pi-computer-use-sessions Coding agent session traces for thomasmustier/pi-computer-use-sessions This dataset contains redacted coding agent session traces collected while working on https://github.com/tmustier/pi-computer-use. The traces were exported with pi-share-hf from local pi workspaces and filtered to keep only sessions that passed deterministic redaction, secret scanning, visual review where applicable, and LLM review. Source git repo: https://github.com/tmustier/pi-computer-use Data… See the full description on the dataset page: https://huggingface.co/datasets/thomasmustier/pi-computer-use-sessions.tabulartext-generationn<1K0 likes154 downloads3mo agoHugging Face05yufan /amazon2023-user-interactions Amazon Reviews 2023 — User Interactions (5-core, leave-one-out, sequential) User–item interaction data for five Amazon Reviews 2023 categories, processed into ready-to-use sequential / generative recommendation splits with the de-facto standard recipe (5-core filtering → chronological ordering → leave-one-out split). Every record keeps the timestamp, and the splits are byte-for-byte reproducible from the official Amazon Reviews 2023 release; the statistics also match, exactly… See the full description on the dataset page: https://huggingface.co/datasets/yufan/amazon2023-user-interactions.tabularother10M<n<100M0 likes138 downloads3mo agoHugging Face06tppllm /us-earthquake U.S. Earthquake Dataset This dataset contains earthquake events in the U.S. from January 1, 2020, to December 31, 2023. It inclucdes 3,009 sequences with 29,521 events across 3 magnitude types. The original data can be accessed via USGS Earthquake Search. The detailed data preprocessing steps used to create this dataset can be found in the TPP-LLM paper and TPP-Embedding paper. Update (2025-10-28): Added three timestamp fields (timestamp_event, timestamp_since_start… See the full description on the dataset page: https://huggingface.co/datasets/tppllm/us-earthquake.tabular1K<n<10K1 likes124 downloads10mo agoHugging Face07OwnedByDanes /Usenet-Corpus-1980-2013-Threaded-Samples Usenet Corpus 1980–2013 — Threaded (Samples) A small, browsable showcase sample of the Usenet Corpus 1980–2013 — Threaded dataset: Usenet posts reconstructed into conversations via thread_id, thread_position, and thread_depth. This repo is a free preview; the full, commercially-licensed corpus (405.6M posts, 190.8M threads, 102.5B tokens) is at: Full threaded dataset (gated): https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Threaded Cleaned (unthreaded)… See the full description on the dataset page: https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Threaded-Samples.tabulartext-generation10K<n<100K0 likes92 downloads12d agoHugging Face08asingh15 /qwen35-2b-tool-use-qwen36-27b-curation-candidates Full candidate collections: 2B tool use + 27B data curation This public Dataset contains two complete, unredacted, exact-40 candidate collections: Tool use: Qwen/Qwen3.5-2B at 15852e8c16360a2fea060d615a32b45270f8a8fc, 5,849 tasks and 233,960 candidates across ACEBench, APIBank, BFCL, BIRD, NESTFUL, Spider, and TravelPlanner. Data curation: Qwen/Qwen3.6-27B at 6a9e13bd6fc8f0983b9b99948120bc37f49c13e9, 5,021 targets and 200,840 candidates, plus the source target rows and the… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/qwen35-2b-tool-use-qwen36-27b-curation-candidates.tabulartext-generation100K<n<1M0 likes77 downloads1mo agoHugging Face09userPresentBench /PresentBench PresentBench: A Fine-Grained Rubric-Based Benchmark for Slide Generation This repository hosts the PresentBench benchmark dataset. 🗂️ Dataset Structure Domains under <dataset_root>/ include (non‑exhaustive): academia/ advertising/ economics/ education/ talk/ Each leaf case typically looks like: material.pdf|material.md|material_N.md|material_N.pdf – source documents (PDFs, text, etc.). generation_task/ – prompts and evaluation configuration: generation_prompt.md… See the full description on the dataset page: https://huggingface.co/datasets/userPresentBench/PresentBench.documentany-to-anyn<1K0 likes63 downloads5mo agoHugging Face10joeygambino /mobile-apps-user-sentiment-reviews Top Mobile Apps User Sentiment & Review Corpus (Google Play) Overview This dataset contains clean, structured public data exported directly from production runs of Apify actors. It serves as a benchmark and sample for lead qualification, market intelligence, research, and machine learning pipelines. Source Actor: captainhandsome/google-play-reviews-scraper Dataset Page: Public sample and schema Preconfigured Run Task: captainhandsome/instagram-1star-reviews… See the full description on the dataset page: https://huggingface.co/datasets/joeygambino/mobile-apps-user-sentiment-reviews.tabularothern<1K0 likes58 downloads5d agoHugging Face11egygi /computer-use-large-actions computer-use-large-actions 9,000 instruction pairs derived from markov-ai/computer-use-large descriptions.parquet (not the raw 12,300 hours of video). Each row is a 10-second segment whose LLM description is a real GUI action (NO_TASK dropped). Source license is CC-BY-4.0. Split by software category examples vscode 2,500 autocad 2,500 blender 1,000 excel 1,000 photoshop 1,000 salesforce 1,000 VS Code and AutoCAD are oversampled for… See the full description on the dataset page: https://huggingface.co/datasets/egygi/computer-use-large-actions.tabulartext-generation1K<n<10K0 likes55 downloads24d agoHugging Face12cfahlgren1 /us-ev-charging-locations United States Electric Vehicle Charging Locations tabular10K<n<100K3 likes48 downloads2y agoHugging Face13anon-user-423 /ACE ACE: Episodic Memory Dataset (StackOverflow Jan–Jun 2025) (v1.0.0) StackOverflow-derived events and monthly episodic rollups (Jan–Jun 2025). Dataset contents ACE contains two related components: events: canonical event records (~96K examples) derived from StackOverflow Q&A threads. episodes: grouped rollups of events for each month, ordered chronologically and packaged in fixed-size windows. Each event includes a question, an accepted answer (or top-scored substitute)… See the full description on the dataset page: https://huggingface.co/datasets/anon-user-423/ACE.tabulartext-retrieval10K<n<100K0 likes41 downloads8mo agoHugging Face14ehejin /user_study-preference-personalized_0505_base_filtered Filtered user study dataset Source repo: ehejin/user_study-preference-personalized_0505_base Each row is ONE item review (pre-rating, conversation, post-rating). Submission-level fields (prolific_pid, demographics, background) are duplicated across rows that share a submission. The 25-50 rows here are the FIRST review for each unique pool index, selected the same way the analysis plot uses — see scripts/plot_vote_shift_3way.py. Total rows: 50 tabularn<1K0 likes37 downloads5mo agoHugging Face15ehejin /user_study-preference-personalized_0505_base_personalized_filtered Filtered user study dataset Source repo: ehejin/user_study-preference-personalized_0505_base_personalized Each row is ONE item review (pre-rating, conversation, post-rating). Submission-level fields (prolific_pid, demographics, background) are duplicated across rows that share a submission. The 25-50 rows here are the FIRST review for each unique pool index, selected the same way the analysis plot uses — see scripts/plot_vote_shift_3way.py. Total rows: 50 tabularn<1K0 likes33 downloads5mo agoHugging Face16Arsh9210 /Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1 Dataset Description: We created an RL dataset for conversational tool-use by utilizing existing expert tool-use trajectories. We pose each assistant step of the trajectory as a separate behavior cloning problem where the policy model is incentivized to match the tool call choices of the expert model. Each trajectory includes the use of tools for authentication, data lookup, servicing (i.e. booking reservations, changing them, getting discounts, etc), and more across 838… See the full description on the dataset page: https://huggingface.co/datasets/Arsh9210/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1.tabular10K<n<100K0 likes31 downloads2mo agoHugging Face17tppllm /us-earthquake-description U.S. Earthquake Description Dataset This dataset contains earthquake events in the U.S. from January 1, 2020, to December 31, 2023. It inclucdes 3,009 sequences with 29,521 events across 3 magnitude types. The original data can be accessed via USGS Earthquake Search. The detailed data preprocessing steps used to create this dataset can be found in the TPP-LLM paper and TPP-Embedding paper. If you find this dataset useful, we kindly invite you to cite the following papers:… See the full description on the dataset page: https://huggingface.co/datasets/tppllm/us-earthquake-description.tabular1K<n<10K1 likes27 downloads10mo agoHugging Face18ehejin /user_study-preference-personalized_0423_5_2_filtered Filtered user study dataset Source repo: ehejin/user_study-preference-personalized_0423_5_2 Each row is ONE item review (pre-rating, conversation, post-rating). Submission-level fields (prolific_pid, demographics, background) are duplicated across rows that share a submission. The 25-50 rows here are the FIRST review for each unique pool index, selected the same way the analysis plot uses — see scripts/plot_vote_shift_3way.py. Total rows: 50 tabularn<1K0 likes27 downloads5mo agoHugging Face19ehejin /user_study-preference-personalized_0505_NP1_filtered Filtered user study dataset Source repo: ehejin/user_study-preference-personalized_0505_NP1 Each row is ONE item review (pre-rating, conversation, post-rating). Submission-level fields (prolific_pid, demographics, background) are duplicated across rows that share a submission. The 25-50 rows here are the FIRST review for each unique pool index, selected the same way the analysis plot uses — see scripts/plot_vote_shift_3way.py. Total rows: 50 tabularn<1K0 likes24 downloads5mo agoHugging Face20ehejin /user_study-preference-personalized_0505_NP3_filtered Filtered user study dataset Source repo: ehejin/user_study-preference-personalized_0505_NP3 Each row is ONE item review (pre-rating, conversation, post-rating). Submission-level fields (prolific_pid, demographics, background) are duplicated across rows that share a submission. The 25-50 rows here are the FIRST review for each unique pool index, selected the same way the analysis plot uses — see scripts/plot_vote_shift_3way.py. Total rows: 50 tabularn<1K0 likes20 downloads5mo agoHugging Face21ehejin /user_study-preference-281_all_filtered Combined user study dataset (0505) Merged from 5 filtered sub-studies. Each row carries a condition and source_repo field. Sub-studies: 0505 NP2 → ehejin/user_study-preference-personalized_BASE_filtered 0505 NP1 → ehejin/user_study-preference-personalized_0423_base_filtered 0505 NP3 → ehejin/user_study-preference-personalized_0423_base_personalized_filtered 0505 base → ehejin/user_study-preference-personalized_0505_base_filtered 0505 base personalized →… See the full description on the dataset page: https://huggingface.co/datasets/ehejin/user_study-preference-281_all_filtered.tabularn<1K0 likes18 downloads4mo agoHugging Face22unicon5553 /user-studytabularn<1K0 likes18 downloads1mo agoHugging Face23ehejin /user_study-preference-personalized_0505_NP2_filtered Filtered user study dataset Source repo: ehejin/user_study-preference-personalized_0505_NP2 Each row is ONE item review (pre-rating, conversation, post-rating). Submission-level fields (prolific_pid, demographics, background) are duplicated across rows that share a submission. The 25-50 rows here are the FIRST review for each unique pool index, selected the same way the analysis plot uses — see scripts/plot_vote_shift_3way.py. Total rows: 50 tabularn<1K0 likes17 downloads5mo agoHugging Face24ehejin /user_study-preference-personalized_0423_4_2_filtered Filtered user study dataset Source repo: ehejin/user_study-preference-personalized_0423_4_2 Each row is ONE item review (pre-rating, conversation, post-rating). Submission-level fields (prolific_pid, demographics, background) are duplicated across rows that share a submission. The 25-50 rows here are the FIRST review for each unique pool index, selected the same way the analysis plot uses — see scripts/plot_vote_shift_3way.py. Total rows: 50 tabularn<1K0 likes16 downloads5mo agoHugging Face25ehejin /user_study-preference-combined Combined user study dataset Item-level rows merged from 8 source repos: ehejin/user_study-preference-base_DETAILED ehejin/user_study-preference-base_DETAILED_checkpoint ehejin/user_study-preference-personalized_BASE ehejin/user_study-preference-personalized_0417_250 ehejin/user_study-preference-personalized_0423_6_2 ehejin/user_study-preference-personalized_0423_5_2 ehejin/user_study-preference-personalized_0423_4_2 ehejin/user_study-preference-personalized_0423_6_2_REAL Each row… See the full description on the dataset page: https://huggingface.co/datasets/ehejin/user_study-preference-combined.tabularn<1K0 likes16 downloads5mo agoHugging Face26anis-mselmi /random-user-profiles Random User Profiles 2,000 synthetic user profiles with random ids, ages, plans and scores. For testing only. Note: This is randomly generated synthetic data with no real-world meaning. Generated for testing and demonstration purposes. tabulartabular-classification1K<n<10K0 likes16 downloads2mo agoHugging Face27deathbyknowledge /Dolci-Instruct-SFT-Tool-Use-Codemode Dolci Instruct SFT Tool Use – Codemode Augmentation This dataset is a transformed version of allenai/Dolci-Instruct-SFT-Tool-Use. It preserves the original conversation content and metadata, but rewrites tool calls into executable JavaScript <codemode> blocks plus structured environment outputs. Only the final codemode-augmented dataset is published here; the original data remains available from the AllenAI dataset above. Source and Attribution Source dataset:… See the full description on the dataset page: https://huggingface.co/datasets/deathbyknowledge/Dolci-Instruct-SFT-Tool-Use-Codemode.tabular10K<n<100K0 likes15 downloads10mo agoHugging Face28ehejin /user_study-preference-personalized_0423_base_personalized_filtered Filtered user study dataset Source repo: ehejin/user_study-preference-personalized_0423_base_personalized Each row is ONE item review (pre-rating, conversation, post-rating). Submission-level fields (prolific_pid, demographics, background) are duplicated across rows that share a submission. The 25-50 rows here are the FIRST review for each unique pool index, selected the same way the analysis plot uses — see scripts/plot_vote_shift_3way.py. Total rows: 50 tabularn<1K0 likes13 downloads5mo agoHugging Face29large-traversaal /mantra-14b-user-interaction-log 🧠 Mantra-14B User Interaction Logs This dataset captures real user interactions with a Gradio demo powered by large-traversaal/Mantra-14B. Each entry logs the user's prompt, the model's response, and additional metadata such as response time and generation parameters. This dataset is ideal for understanding how people engage with the model, evaluating responses, or fine-tuning on real-world usage data. 🔍 What’s Inside Each row in the dataset includes: timestamp –… See the full description on the dataset page: https://huggingface.co/datasets/large-traversaal/mantra-14b-user-interaction-log.tabulartext-generationn<1K0 likes12 downloads1y agoHugging Face30ehejin /user_study-preference-personalized_BASE_filtered Filtered user study dataset Source repo: ehejin/user_study-preference-personalized_BASE Each row is ONE item review (pre-rating, conversation, post-rating). Submission-level fields (prolific_pid, demographics, background) are duplicated across rows that share a submission. The 25-50 rows here are the FIRST review for each unique pool index, selected the same way the analysis plot uses — see scripts/plot_vote_shift_3way.py. Total rows: 50 tabularn<1K0 likes12 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.