datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Nemotron-SFT-Agentic-v2-prompt-only
Nemotron-SFT-Agentic-v2-prompt-only
Prompt-only extraction from nvidia/Nemotron-SFT-Agentic-v2.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row indexes where prompt extraction… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-SFT-Agentic-v2-prompt-only.deepresearchgym-agentic-search-logs
DeepResearchGym Agentic Search Logs
This repository hosts the dataset accompanying the paper “Agentic Search in the Wild” (arXiv: https://arxiv.org/abs/2601.17617).
The dataset contains 14M+ search queries collected via DeepResearchGym (DRGym), an open-source search API designed for DeepResearch-style agentic search. For more background on DRGym, see: https://arxiv.org/abs/2505.19253.
All records have been anonymized and shuffled to prevent re-identification, and we additionally… See the full description on the dataset page: https://huggingface.co/datasets/cx-cmu/deepresearchgym-agentic-search-logs.DrugbankVocabularydaily-oracle
Daily Oracle
📰 Project Website📝 Paper - Are LLMs Prescient? A Continuous Evaluation using Daily News as the Oracle
Daily Oracle is a continuous evaluation benchmark using automatically generated QA pairs from daily news to assess how the future prediction capabilities of LLMs evolve over time.
Dataset Details
Question Type: True/False (TF) & Multiple Choice (MC)
Current Version*
Time Span: 2020.01.01 - 2026.07.18
Size: 20,376 TF questions and 18,557 MC… See the full description on the dataset page: https://huggingface.co/datasets/agentic-learning-ai-lab/daily-oracle.agentic-score-leaderboard
🛠️ Agentic Score Leaderboard — one RTX 5090
How well do local models actually drive a tool-using agent loop? Not single-call function-calling
benchmarks — a real loop: native OpenAI tool-calling through llama-server, multi-step deterministic
tasks, programmatic verification. Everything runs on a single RTX 5090 32GB.
Updated 2026-06-17 · llama.cpp b9562 · --jinja native tool-calling · temp 0.
Leaderboard
#
model
params
Agentic Score
success
tool-eff… See the full description on the dataset page: https://huggingface.co/datasets/witcheer/agentic-score-leaderboard.PortBench-Market
PortBench Market Base Dataset
Dataset Description
A ten-year (Jan 2015–Dec 2025) daily financial dataset covering 183 instruments across six heterogeneous asset classes, designed for multi-asset portfolio management research and LLM evaluation.
Asset Coverage
Asset Class
Instruments
Data Fields
Sources
Equities
126
OHLCV + return
Yahoo Finance (ETFs: broad market, sector, factor, international)
Bonds
16
Close + return (ETFs); yield… See the full description on the dataset page: https://huggingface.co/datasets/AgenticFinLab/PortBench-Market.PyFi-600K
Dataset Card for PyFi-600K
This dataset card aims to be a introduction for PyFi-600K, A financial VLM dataset containing 600K question-answer pairs generated via Adversarial agents.
AgenticFinLab/PyFi-600K/
├── README.md # Dataset documentation and description
├── images.zip # Compressed image files
├── PyFi-600K-dataset.csv # Q&A pairs in CSV format
├── PyFi-600K-dataset.json # Q&A pairs in JSON format
├── PyFi-600K-chain-dataset.json # Chain of Thought Q&A pairs dataset
└──… See the full description on the dataset page: https://huggingface.co/datasets/AgenticFinLab/PyFi-600K.DILIrankllm-agentic-precomputed-v3Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1-prompt-only
Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Agentic-Indirect-Prompt-Injection-v1-prompt-only.agentic-reasoning-benchmark
Agentic & Reasoning Benchmark (ARB) – Expanded
Ein synthetischer Benchmark mit 2.550 Fragen und Lösungen, optimiert für die Evaluation von Agentic Capabilities und Reasoning.
Überblick
Eigenschaft
Wert
Anzahl Beispiele
2.550
Kategorien
8
Schwierigkeitsgrade
easy / medium / hard
Formate
CSV + JSON
Reproduzierbarkeit
Generator-Skript (seed=42) enthalten
Lizenz
CC-BY-4.0
Kategorien
Kategorie
Anzahl
Beschreibung… See the full description on the dataset page: https://huggingface.co/datasets/roskosmos19/agentic-reasoning-benchmark.Nemotron-RL-Agentic-SWE-Pivot-v1-prompt-only
Nemotron-RL-Agentic-SWE-Pivot-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Agentic-SWE-Pivot-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row indexes where… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Agentic-SWE-Pivot-v1-prompt-only.Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1-prompt-only
Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Agentic-Conversational-Tool-Use-Pivot-v1-prompt-only.DIQTANemotron-RL-Agentic-Function-Calling-Pivot-v1-prompt-only
Nemotron-RL-Agentic-Function-Calling-Pivot-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-RL-Agentic-Function-Calling-Pivot-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-RL-Agentic-Function-Calling-Pivot-v1-prompt-only.DICTrankate-agentic-coverage
O*NET Agentic Coverage Dataset
A task-level estimate, for all 18,796 O*NET work tasks — the full O*NET task
corpus, economy-wide — of how much of each task a current autonomous AI agent
could complete end to end, with no human in the loop.
This dataset supports task-level research on AI exposure and human-AI task
allocation: which tasks an agent can already finish alone, which need a human for
part of the work, and which still require a human throughout.
Built as a companion… See the full description on the dataset page: https://huggingface.co/datasets/ravishgupta/ate-agentic-coverage.persona-and-other-evals
Qwen3.5-9B AMA adapters — persona evals
Inference code, the data it produced, and the tools that turn that data
into tables and an HTML viewer. The evals are Anthropic's persona set,
scored in three regimes: teacher-forced logprob of the answer literal,
greedy answer with the reasoning block pre-closed, and a full 16k-budget
reasoning trace.
Pinned models
base unsloth/Qwen3.5-9B @ 005429cee5cb648998cf2b70eebdd83175989c9a
util… See the full description on the dataset page: https://huggingface.co/datasets/agentic-moral-alignment/persona-and-other-evals.Portfolio-OptimizationFinancial-Advisory-ClientsDrugLinksagentic-data-access-benchmark
Agentic Data Access Benchmark (ADAB)
Agentic Data Access Benchmark is a set of real-world questions over few "closed domains" to illustrate the evaluation of closed domain AI assistants/agents.
Closed domains are domains where data is not available implicitly in the LLM as they reside in secure or private systems e.g. enterprise databases, SaaS applications, etc
and AI solutions require mechanisms to connect an LLM to such data. If you are evaluating an AI product or building your… See the full description on the dataset page: https://huggingface.co/datasets/hasura/agentic-data-access-benchmark.Portfolio-Rebalancedeepresearchgym-agentic-search-logs
DeepResearchGym Agentic Search Logs
This repository hosts the dataset accompanying the paper “Agentic Search in the Wild” (arXiv: https://arxiv.org/abs/2601.17617).
The dataset contains 14M+ search queries collected via DeepResearchGym (DRGym), an open-source search API designed for DeepResearch-style agentic search. For more background on DRGym, see: https://arxiv.org/abs/2505.19253.
All records have been anonymized and shuffled to prevent re-identification, and we additionally… See the full description on the dataset page: https://huggingface.co/datasets/ethanning/deepresearchgym-agentic-search-logs.AgenticRetrievalBench
Dataset Card for Agentic Retrieval Benchmark
In recent years, there has been a large body of work in the field of text retrieval. However, existing studies are often scattered across different datasets, and their comparisons are partial and fragmented, lacking a comprehensive benchmark for evaluating retrieval performance.
This dataset is provided as part of the Agentic Retrieval Benchmark.
The project aims to establish a reproducible benchmark for LLM-augmented text retrieval.… See the full description on the dataset page: https://huggingface.co/datasets/PrismShadow/AgenticRetrievalBench.Personal-Finance-DataNemotron-Agentic-v1-prompt-only
Nemotron-Agentic-v1-prompt-only
Prompt-only extraction from nvidia/Nemotron-Agentic-v1.
Files:
prompts.csv: one prompt extraction record per source row. Records include
prompt, separated system_prompt, and structured tools when the source row
defines available tools. Nested values are JSON-encoded inside CSV cells.
summary.md: source row counts, extracted row counts, count deltas, and failed prompt counts.
null_or_empty_rows.md: row indexes where prompt extraction produced a… See the full description on the dataset page: https://huggingface.co/datasets/jamesdborin/Nemotron-Agentic-v1-prompt-only.Portfolio-ManagementFinancial-ReportsCredit-Portfolio-Optimization
