asingh15/qwen35-2b-tool-use-explorer
Qwen3.5-2B tool-use candidate explorer
This free public static Space presents the complete, unredacted Qwen3.5-2B tool-use candidate collection in a human-readable form. It covers 5,849 tasks and 233,960 exact-40 candidates from ACEBench, APIBank, BFCL, BIRD, NESTFUL, Spider, and TravelPlanner. AppWorld is not included.
The viewer supports source, split, tier, evaluator-outcome, task-ID, and prompt filters; per-task success and fractional-score charts; parsed tool-call, tool-result, and final-answer timelines; side-by-side candidate comparison; and full raw responses with model provenance, seeds, and response hashes.
success is the source evaluator's binary exact-success label. frac is an independent source-specific fractional score, so the two values can disagree. They describe evaluator results, not whether collection itself succeeded.
The app downloads only a compact byte-offset index. It retrieves a selected task through a bounded byte-range request against an immutable public Dataset revision, with a pinned-head Dataset Viewer request as fallback. Candidate content is escaped before HTML rendering. The static runtime has no owner credentials, service secret, analytics, or browser persistence.
The Space code and generated candidate collection are Apache-2.0. Source-level attribution is preserved in the linked Dataset repository.
