agentic-data
Nemotron-AIQ-Agentic-Safety-Dataset-1.0
Nemotron-AIQ Agentic Safety Dataset
Dataset Summary
Nemotron-AIQ-Agentic-Safety-Dataset is a comprehensive dataset that captures a broad range of novel safety and security contextual risks that can emerge within agentic systems. It highlights the robustness of NVIDIA's open model, llama-3.3-nemotron-super-49b-v1, when deployed as a research assistant inside AIQ, demonstrating its ability to handle a diverse spectrum of agentic safety and security challenges. The dataset… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-AIQ-Agentic-Safety-Dataset-1.0.AgenticDataBench
AgenticDataBench: A Comprehensive Benchmark for Data Agents
Project Page | GitHub | Paper
AgenticDataBench is a comprehensive benchmark for evaluating LLM-based data agents that automate real-world data science workflows. It addresses the lack of rigorous evaluation by providing diverse, realistic tasks with fine-grained ground-truth labels.
The benchmark spans 15 domains, including real B2B fintech use cases, and is structured around reusable data science skills—core… See the full description on the dataset page: https://huggingface.co/datasets/shawnzzzh/AgenticDataBench.sol-max-opusnode-data
sol-max-opusnode-data
Training data built by the AgentPTB arm for cell sol-max-opusnode — Codex / gpt-5.6-sol @ effort max.
This is the corpus the arm itself assembled during its 100-hour run: what it downloaded,
filtered, rewrote and mixed. It is the input side of the checkpoints published as
agentic-ptb/sol-max-opusnode.h*, and the companion to the run record in agentic-ptb/sol-max-opusnode-record.
field
value
plot cell
sol-max-opusnode
driver
Codex / gpt-5.6-sol… See the full description on the dataset page: https://huggingface.co/datasets/agentic-ptb/sol-max-opusnode-data.sol-max-data
sol-max-data
Training data built by the AgentPTB arm for cell sol-max — Codex / gpt-5.6-sol @ effort max.
This is the corpus the arm itself assembled during its 100-hour run: what it downloaded,
filtered, rewrote and mixed. It is the input side of the checkpoints published as
agentic-ptb/sol-max.h*, and the companion to the run record in agentic-ptb/sol-max-record.
field
value
plot cell
sol-max
driver
Codex / gpt-5.6-sol
reasoning effort
max
total size
114.47 GB… See the full description on the dataset page: https://huggingface.co/datasets/agentic-ptb/sol-max-data.EMU-Agentic-PostTrain-Dataopus-high-v3-data
opus-high-v3 — complete research record
This dataset archives the qualitative and quantitative record of the
msr-agentic-ptb-opus / opus-high-v3 Claude Code research run.
The submitted artifact uses the unmodified base weights with a two-attempt
Pi verifier harness. The final replicated SWE result was 24.6% (245/995) with
the stock scaffold and 32.4% (321/990) with the submitted harness. Training
did not improve the weights; all trained variants measured at or below the
base… See the full description on the dataset page: https://huggingface.co/datasets/agentic-ptb/opus-high-v3-data.
