rl-rag/open-scholar-gpt-oss-120b
open-scholar-gpt-oss-120b Deep research agent evaluation on data/drtulu_open_scholar.jsonl (normal split). Results Metric Value pass@1 0.0% avg@1 0.0% Trajectory accuracy 0.0% (0/11854) Questions 11854 Trajectories 11854 (1 per question) Avg tool calls 17.3 Full conversations ✅ Model & Setup Model gpt-oss-120b Judge gpt-4o Max tool calls 50 Temperature 0.7 Blocked domains None Tool… See the full description on the dataset page: https://huggingface.co/datasets/rl-rag/open-scholar-gpt-oss-120b.
open-scholar-gpt-oss-120b
Deep research agent evaluation on data/drtulu_open_scholar.jsonl (normal split).
Results
Model & Setup
Tool Usage
Total: 199,013 tool calls (17.3 per trajectory)
⚠️ Contamination: 56/11854 trajectories flagged (search results contained evaluation dataset pages). Filter with ds.filter(lambda x: not x['contaminated']).
