datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ovos-stt-bench-ami-en-GB
OVOS stt bench — ami-en-GB
Per-clip transcripts predictions of the registered
OVOS Plugin Arena
stt fighters over
edinburghcstr/ami.
One dedicated repo per modality; one dataset split per language; one JSONL
file per fighter under predictions/<lang>/<competitor_id>.jsonl. Rows follow
the arena §3.2 contract (pinned dataset_revision, plugin_version,
latency_ms). Produced by the reproducible benchmark script in the arena repo;
the arena's assemble workflow turns these rows into… See the full description on the dataset page: https://huggingface.co/datasets/OpenVoiceOS/ovos-stt-bench-ami-en-GB.reason-math-correctness
Correctness as a causal weight-space direction — MATH
Artifacts for the experiment: is "being correct" a single, causal direction in a language model's
weights? We train 12,976 rank-1 LoRA adapters on correct vs. wrong solutions to MATH problems with
Qwen/Qwen3-4B-Base, extract a direction from them, and show that steering the model's weights
along it causally controls accuracy on held-out MATH-500 — ablation collapses accuracy 62%→5%,
amplification lifts it to 77%, and a… See the full description on the dataset page: https://huggingface.co/datasets/amildravid4292/reason-math-correctness.drugbank_drug_target_label_mapping_amino_acid_pairhigh-temp-refusal-probe-artifactsptfa
