datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
scientific-question-outcomes
Scientific Question Outcomes
980 astronomy research questions, frozen at five historical cutoffs, each
labelled with what the following five years of literature actually did with
it.
Systems that propose research questions are usually evaluated by asking a
person or a model how good the questions sound. This dataset supplies the
alternative: questions frozen using only pre-cutoff literature, and outcome
labels drawn from the literature published afterwards. It is, to our… See the full description on the dataset page: https://huggingface.co/datasets/huiluckylucky/scientific-question-outcomes.gfjd-outcomes-evidence
GFJD Outcomes Evidence
Registry status
Registry ID: edithatogo/gfjd-outcomes-evidence
Family: gfjd
Repository role: outcomes_evidence
Canonical dataset: edithatogo/gfjd-source-catalogue
Operational status: reserved_empty_private
Rights status: no_payload_published
Authoritative catalog: edithatogo/dataset-estate-registry
Origin and provenance
Origin repository: unresolved; no origin has been inferred.
Upstream source: Origin not yet evidenced.… See the full description on the dataset page: https://huggingface.co/datasets/edithatogo/gfjd-outcomes-evidence.echr-outcomesgpqa-rethink-outcomes
Stage-3 rethink outcomes
994 attempts in which a 2B model, having already answered a GPQA question wrongly, was asked
again with one of three preambles. The question is whether telling a small model what it got
wrong actually helps.
Three arms, and the control is the point:
arm
what it was given
control
a content-free nudge, "double-check your reasoning carefully at each step"
meta
an 80B model's specific critique of its own flagged reasoning, answer withheld… See the full description on the dataset page: https://huggingface.co/datasets/waqas137/gpqa-rethink-outcomes.
