datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1
Dataset Description:
We created an RL dataset for conversational tool-use by utilizing existing expert tool-use trajectories. We pose each assistant step of the trajectory as a separate behavior cloning problem where the policy model is incentivized to match the tool call choices of the expert model. Each trajectory includes the use of tools for authentication, data lookup, servicing (i.e. booking reservations, changing them, getting discounts, etc), and more across 838 different… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1.user_study-preference-personalized_0423_base_filtered
Filtered user study dataset
Source repo: ehejin/user_study-preference-personalized_0423_base
Each row is ONE item review (pre-rating, conversation, post-rating). Submission-level
fields (prolific_pid, demographics, background) are duplicated across rows that share
a submission.
The 25-50 rows here are the FIRST review for each unique pool index, selected the same
way the analysis plot uses — see scripts/plot_vote_shift_3way.py.
Total rows: 50
data-use-ner
Data-use-ner (human holdout)
GLiNER-format human-adjudicated holdout: 473 spans — annotator190 (190, origin=fcv_pads_east_africa) + jdc283 (283, origin=jdc_operational). Never trained on.
Source: rafmacalaba/datause-displacement-reviewed holdout (gliner_reviewed token spans + readable_reviewed passages, v2.4 labels) with v3 probe head_score (outputs/gliner_datause_v3_probe_human473.jsonl).
Columns
text (full passage = " ".join(tokenized_text); span char offsets… See the full description on the dataset page: https://huggingface.co/datasets/rafmacalaba/data-use-ner.pi-computer-use-sessions
Coding agent session traces for thomasmustier/pi-computer-use-sessions
This dataset contains redacted coding agent session traces collected while working on https://github.com/tmustier/pi-computer-use. The traces were exported with pi-share-hf from local pi workspaces and filtered to keep only sessions that passed deterministic redaction, secret scanning, visual review where applicable, and LLM review.
Source git repo: https://github.com/tmustier/pi-computer-use
Data… See the full description on the dataset page: https://huggingface.co/datasets/thomasmustier/pi-computer-use-sessions.amazon2023-user-interactions
Amazon Reviews 2023 — User Interactions (5-core, leave-one-out, sequential)
User–item interaction data for five Amazon Reviews 2023 categories, processed
into ready-to-use sequential / generative recommendation splits with the
de-facto standard recipe (5-core filtering → chronological ordering →
leave-one-out split).
Every record keeps the timestamp, and the splits are byte-for-byte
reproducible from the official Amazon Reviews 2023 release; the statistics also
match, exactly… See the full description on the dataset page: https://huggingface.co/datasets/yufan/amazon2023-user-interactions.us-earthquake
U.S. Earthquake Dataset
This dataset contains earthquake events in the U.S. from January 1, 2020, to December 31, 2023. It inclucdes 3,009 sequences with 29,521 events across 3 magnitude types. The original data can be accessed via USGS Earthquake Search. The detailed data preprocessing steps used to create this dataset can be found in the TPP-LLM paper and TPP-Embedding paper.
Update (2025-10-28): Added three timestamp fields (timestamp_event, timestamp_since_start… See the full description on the dataset page: https://huggingface.co/datasets/tppllm/us-earthquake.Usenet-Corpus-1980-2013-Threaded-Samples
Usenet Corpus 1980–2013 — Threaded (Samples)
A small, browsable showcase sample of the Usenet Corpus 1980–2013 — Threaded
dataset: Usenet posts reconstructed into conversations via thread_id,
thread_position, and thread_depth. This repo is a free preview; the full,
commercially-licensed corpus (405.6M posts, 190.8M threads, 102.5B tokens) is at:
Full threaded dataset (gated): https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Threaded
Cleaned (unthreaded)… See the full description on the dataset page: https://huggingface.co/datasets/OwnedByDanes/Usenet-Corpus-1980-2013-Threaded-Samples.qwen35-2b-tool-use-qwen36-27b-curation-candidates
Full candidate collections: 2B tool use + 27B data curation
This public Dataset contains two complete, unredacted, exact-40 candidate collections:
Tool use: Qwen/Qwen3.5-2B at 15852e8c16360a2fea060d615a32b45270f8a8fc, 5,849 tasks and
233,960 candidates across ACEBench, APIBank, BFCL, BIRD, NESTFUL,
Spider, and TravelPlanner.
Data curation: Qwen/Qwen3.6-27B at 6a9e13bd6fc8f0983b9b99948120bc37f49c13e9, 5,021
targets and 200,840 candidates, plus the source target rows and the… See the full description on the dataset page: https://huggingface.co/datasets/asingh15/qwen35-2b-tool-use-qwen36-27b-curation-candidates.PresentBench
PresentBench: A Fine-Grained Rubric-Based Benchmark for Slide Generation
This repository hosts the PresentBench benchmark dataset.
🗂️ Dataset Structure
Domains under <dataset_root>/ include (non‑exhaustive):
academia/
advertising/
economics/
education/
talk/
Each leaf case typically looks like:
material.pdf|material.md|material_N.md|material_N.pdf – source documents (PDFs, text, etc.).
generation_task/ – prompts and evaluation configuration:
generation_prompt.md… See the full description on the dataset page: https://huggingface.co/datasets/userPresentBench/PresentBench.mobile-apps-user-sentiment-reviews
Top Mobile Apps User Sentiment & Review Corpus (Google Play)
Overview
This dataset contains clean, structured public data exported directly from production runs of Apify actors.
It serves as a benchmark and sample for lead qualification, market intelligence, research, and machine learning pipelines.
Source Actor: captainhandsome/google-play-reviews-scraper
Dataset Page: Public sample and schema
Preconfigured Run Task: captainhandsome/instagram-1star-reviews… See the full description on the dataset page: https://huggingface.co/datasets/joeygambino/mobile-apps-user-sentiment-reviews.computer-use-large-actions
computer-use-large-actions
9,000 instruction pairs derived from markov-ai/computer-use-large descriptions.parquet (not the raw 12,300 hours of video).
Each row is a 10-second segment whose LLM description is a real GUI action (NO_TASK dropped). Source license is CC-BY-4.0.
Split by software
category
examples
vscode
2,500
autocad
2,500
blender
1,000
excel
1,000
photoshop
1,000
salesforce
1,000
VS Code and AutoCAD are oversampled for… See the full description on the dataset page: https://huggingface.co/datasets/egygi/computer-use-large-actions.us-ev-charging-locations
United States Electric Vehicle Charging Locations
ACE
ACE: Episodic Memory Dataset (StackOverflow Jan–Jun 2025) (v1.0.0)
StackOverflow-derived events and monthly episodic rollups (Jan–Jun 2025).
Dataset contents
ACE contains two related components:
events: canonical event records (~96K examples) derived from StackOverflow Q&A threads.
episodes: grouped rollups of events for each month, ordered chronologically and packaged in fixed-size windows.
Each event includes a question, an accepted answer (or top-scored substitute)… See the full description on the dataset page: https://huggingface.co/datasets/anon-user-423/ACE.user_study-preference-personalized_0505_base_filtered
Filtered user study dataset
Source repo: ehejin/user_study-preference-personalized_0505_base
Each row is ONE item review (pre-rating, conversation, post-rating). Submission-level
fields (prolific_pid, demographics, background) are duplicated across rows that share
a submission.
The 25-50 rows here are the FIRST review for each unique pool index, selected the same
way the analysis plot uses — see scripts/plot_vote_shift_3way.py.
Total rows: 50
user_study-preference-personalized_0505_base_personalized_filtered
Filtered user study dataset
Source repo: ehejin/user_study-preference-personalized_0505_base_personalized
Each row is ONE item review (pre-rating, conversation, post-rating). Submission-level
fields (prolific_pid, demographics, background) are duplicated across rows that share
a submission.
The 25-50 rows here are the FIRST review for each unique pool index, selected the same
way the analysis plot uses — see scripts/plot_vote_shift_3way.py.
Total rows: 50
Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1
Dataset Description:
We created an RL dataset for conversational tool-use by utilizing existing expert tool-use trajectories. We pose each assistant step of the trajectory as a separate behavior cloning problem where the policy model is incentivized to match the tool call choices of the expert model. Each trajectory includes the use of tools for authentication, data lookup, servicing (i.e. booking reservations, changing them, getting discounts, etc), and more across 838… See the full description on the dataset page: https://huggingface.co/datasets/Arsh9210/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1.us-earthquake-description
U.S. Earthquake Description Dataset
This dataset contains earthquake events in the U.S. from January 1, 2020, to December 31, 2023. It inclucdes 3,009 sequences with 29,521 events across 3 magnitude types. The original data can be accessed via USGS Earthquake Search. The detailed data preprocessing steps used to create this dataset can be found in the TPP-LLM paper and TPP-Embedding paper.
If you find this dataset useful, we kindly invite you to cite the following papers:… See the full description on the dataset page: https://huggingface.co/datasets/tppllm/us-earthquake-description.user_study-preference-personalized_0423_5_2_filtered
Filtered user study dataset
Source repo: ehejin/user_study-preference-personalized_0423_5_2
Each row is ONE item review (pre-rating, conversation, post-rating). Submission-level
fields (prolific_pid, demographics, background) are duplicated across rows that share
a submission.
The 25-50 rows here are the FIRST review for each unique pool index, selected the same
way the analysis plot uses — see scripts/plot_vote_shift_3way.py.
Total rows: 50
user_study-preference-personalized_0505_NP1_filtered
Filtered user study dataset
Source repo: ehejin/user_study-preference-personalized_0505_NP1
Each row is ONE item review (pre-rating, conversation, post-rating). Submission-level
fields (prolific_pid, demographics, background) are duplicated across rows that share
a submission.
The 25-50 rows here are the FIRST review for each unique pool index, selected the same
way the analysis plot uses — see scripts/plot_vote_shift_3way.py.
Total rows: 50
user_study-preference-personalized_0505_NP3_filtered
Filtered user study dataset
Source repo: ehejin/user_study-preference-personalized_0505_NP3
Each row is ONE item review (pre-rating, conversation, post-rating). Submission-level
fields (prolific_pid, demographics, background) are duplicated across rows that share
a submission.
The 25-50 rows here are the FIRST review for each unique pool index, selected the same
way the analysis plot uses — see scripts/plot_vote_shift_3way.py.
Total rows: 50
user_study-preference-281_all_filtered
Combined user study dataset (0505)
Merged from 5 filtered sub-studies. Each row carries a condition
and source_repo field.
Sub-studies:
0505 NP2 → ehejin/user_study-preference-personalized_BASE_filtered
0505 NP1 → ehejin/user_study-preference-personalized_0423_base_filtered
0505 NP3 → ehejin/user_study-preference-personalized_0423_base_personalized_filtered
0505 base → ehejin/user_study-preference-personalized_0505_base_filtered
0505 base personalized →… See the full description on the dataset page: https://huggingface.co/datasets/ehejin/user_study-preference-281_all_filtered.user-studyuser_study-preference-personalized_0505_NP2_filtered
Filtered user study dataset
Source repo: ehejin/user_study-preference-personalized_0505_NP2
Each row is ONE item review (pre-rating, conversation, post-rating). Submission-level
fields (prolific_pid, demographics, background) are duplicated across rows that share
a submission.
The 25-50 rows here are the FIRST review for each unique pool index, selected the same
way the analysis plot uses — see scripts/plot_vote_shift_3way.py.
Total rows: 50
user_study-preference-personalized_0423_4_2_filtered
Filtered user study dataset
Source repo: ehejin/user_study-preference-personalized_0423_4_2
Each row is ONE item review (pre-rating, conversation, post-rating). Submission-level
fields (prolific_pid, demographics, background) are duplicated across rows that share
a submission.
The 25-50 rows here are the FIRST review for each unique pool index, selected the same
way the analysis plot uses — see scripts/plot_vote_shift_3way.py.
Total rows: 50
user_study-preference-combined
Combined user study dataset
Item-level rows merged from 8 source repos:
ehejin/user_study-preference-base_DETAILED
ehejin/user_study-preference-base_DETAILED_checkpoint
ehejin/user_study-preference-personalized_BASE
ehejin/user_study-preference-personalized_0417_250
ehejin/user_study-preference-personalized_0423_6_2
ehejin/user_study-preference-personalized_0423_5_2
ehejin/user_study-preference-personalized_0423_4_2
ehejin/user_study-preference-personalized_0423_6_2_REAL
Each row… See the full description on the dataset page: https://huggingface.co/datasets/ehejin/user_study-preference-combined.random-user-profiles
Random User Profiles
2,000 synthetic user profiles with random ids, ages, plans and scores. For testing only.
Note: This is randomly generated synthetic data with no real-world meaning. Generated for testing and demonstration purposes.
Dolci-Instruct-SFT-Tool-Use-Codemode
Dolci Instruct SFT Tool Use – Codemode Augmentation
This dataset is a transformed version of allenai/Dolci-Instruct-SFT-Tool-Use.
It preserves the original conversation content and metadata, but rewrites tool calls into executable JavaScript <codemode> blocks plus structured environment outputs.
Only the final codemode-augmented dataset is published here; the original data remains available from the AllenAI dataset above.
Source and Attribution
Source dataset:… See the full description on the dataset page: https://huggingface.co/datasets/deathbyknowledge/Dolci-Instruct-SFT-Tool-Use-Codemode.user_study-preference-personalized_0423_base_personalized_filtered
Filtered user study dataset
Source repo: ehejin/user_study-preference-personalized_0423_base_personalized
Each row is ONE item review (pre-rating, conversation, post-rating). Submission-level
fields (prolific_pid, demographics, background) are duplicated across rows that share
a submission.
The 25-50 rows here are the FIRST review for each unique pool index, selected the same
way the analysis plot uses — see scripts/plot_vote_shift_3way.py.
Total rows: 50
mantra-14b-user-interaction-log
🧠 Mantra-14B User Interaction Logs
This dataset captures real user interactions with a Gradio demo powered by large-traversaal/Mantra-14B. Each entry logs the user's prompt, the model's response, and additional metadata such as response time and generation parameters. This dataset is ideal for understanding how people engage with the model, evaluating responses, or fine-tuning on real-world usage data.
🔍 What’s Inside
Each row in the dataset includes:
timestamp –… See the full description on the dataset page: https://huggingface.co/datasets/large-traversaal/mantra-14b-user-interaction-log.user_study-preference-personalized_BASE_filtered
Filtered user study dataset
Source repo: ehejin/user_study-preference-personalized_BASE
Each row is ONE item review (pre-rating, conversation, post-rating). Submission-level
fields (prolific_pid, demographics, background) are duplicated across rows that share
a submission.
The 25-50 rows here are the FIRST review for each unique pool index, selected the same
way the analysis plot uses — see scripts/plot_vote_shift_3way.py.
Total rows: 50
