rajatagarwal457/sarvam-top10-review
Sarvam Top-10 Task Review 10 tasks from the 118-task gworkspace_hard_v1 benchmark, selected for trace richness (actions per goal ≥ 1.5). 5 tasks where Sarvam-105B passed; 5 engaged failures. Each row is a Sarvam-105B rollout (n=1); the other three models' results live in the peer_results column. Selection Passes (5): all Sarvam-105B passes in the dataset, per the trace file. Failures (5): Sarvam ≥7 actions, failed, capped at 2 per family for diversity.… See the full description on the dataset page: https://huggingface.co/datasets/rajatagarwal457/sarvam-top10-review.
Sarvam Top-10 Task Review
10 tasks from the 118-task gworkspace_hard_v1 benchmark, selected for trace richness (actions per goal ≥ 1.5). 5 tasks where Sarvam-105B passed; 5 engaged failures. Each row is a Sarvam-105B rollout (n=1); the other three models' results live in the peer_results column.
Selection
- Passes (5): all Sarvam-105B passes in the dataset, per the trace file.
- Failures (5): Sarvam ≥7 actions, failed, capped at 2 per family for diversity.
Failure-mode taxonomy
Tags are not mutually exclusive. A single trajectory can carry multiple tags.
Columns
How to view
Open the Dataset Viewer. Filter failure_modes contains empty_doc_share to see every row exhibiting that pattern. Click a row to expand trajectory and see the model's reasoning step by step.
Notes on data quality
- The dataset summary file
gworkspace_hard_v1.jsonhad 5 stale Sarvam pass/exec rows that disagreed with the trace file. This dataset uses the trace file as source of truth. - Failure-mode tags were generated by heuristic rules and reviewed manually. Where the ceiling model (gpt-5.2) also fails a predicate, the tag
prompt_ambiguityis added — those are candidates for prompt rewrites rather than model evaluation.
