goose
Datasets
All datasets matching “goose”harbor-goose-openhands-benchmark
Same Model, Opposite Results: Goose vs OpenHands Turn Budget Study on Harbor Terminal-Bench-Pro
Trial-level results from a small controlled study comparing two agent harnesses —
Goose and OpenHands-SDK —
on a frozen 40-task Harbor Terminal-Bench-Pro slice.
All runs used minimax/minimax-m2.5 via OpenRouter with Daytona as the sandbox backend.
Key Findings
Reducing the turn budget from 100 to 60 pushed the two harnesses in opposite directions under the base setup:… See the full description on the dataset page: https://huggingface.co/datasets/namanvats/harbor-goose-openhands-benchmark.GooseReason-0.7M
Xent GooseReason 0.7M
This is a reproducibly shuffled and benchmark-decontaminated derivative of
nvidia/Nemotron-Research-GooseReason-0.7M.
Splits
Each source subset has a training split plus validation and test splits containing
500 rows each. Holdouts are stratified by num_choices with largest-remainder
allocation, then every split is deterministically shuffled with seed 42.
Source subset
Original
Raw exact
Added by normalization
Invalid/mask removed… See the full description on the dataset page: https://huggingface.co/datasets/xent-labs/GooseReason-0.7M.Nemotron-Research-GooseReason-0.7M
GooseReason-0.7M
Synthesized with Golden Goose: A Simple Trick to Synthesize Unlimited RLVR Tasks from Unverifiable Internet Text
GooseReason-0.7M is a large-scale RLVR dataset with over 0.7 million tasks across mathematics, programming, and general scientific domains, synthesized by the Golden Goose pipeline. It is used to train GooseReason-4B-Instruct, which achieves new state-of-the-art results among 4B-Instruct models across 15 diverse benchmarks, spanning mathematics… See the full description on the dataset page: https://huggingface.co/datasets/nvidia/Nemotron-Research-GooseReason-0.7M.RWKV-World-v3
RWKV-7 (Goose) World v3 Corpus
Paper | Code
This is an itemised and annotated list of the RWKV World v3 corpus
which is a multilingual dataset with about 3.1T tokens used to train the
"Goose" RWKV-7 World model series.
RWKV World v3 was crafted from public datasets spanning >100 world languages
(80% English, 10% multilang, and 10% code). Also available as a HF Collection of Datasets.
Subsampled subsets (previews) of the corpus are available as 100k JSONL dataset and 1M JSONL dataset… See the full description on the dataset page: https://huggingface.co/datasets/Goose-World/RWKV-World-v3.tianjin-pm25-dataclinical-trial-outcomes-2020plus
Clinical Trial Outcomes (2020+) with Normalized Endpoints
124,790 normalized endpoints across 14,170 clinical studies that started on or after
2020-01-01 and have posted results on ClinicalTrials.gov.
Snapshot: 2026-09-08. Source: ClinicalTrials.gov API v2 (U.S. National Library of Medicine).
Built with ctgov — the same normalizer, released as an
MIT-licensed package with zero dependencies. So this snapshot is not a dead artifact: you can
re-run it against the live registry, or… See the full description on the dataset page: https://huggingface.co/datasets/GooseWithStories/clinical-trial-outcomes-2020plus.
