datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ceshi0119
Dataset Card for "super_glue"
Dataset Summary
SuperGLUE (https://super.gluebenchmark.com/) is a new benchmark styled after
GLUE with a new set of more difficult language understanding tasks, improved
resources, and a new public leaderboard.
BoolQ (Boolean Questions, Clark et al., 2019a) is a QA task where each example consists of a short
passage and a yes/no question about the passage. The questions are provided anonymously and
unsolicited by users of the Google search… See the full description on the dataset page: https://huggingface.co/datasets/Xieyiyiyi/ceshi0119.ceshi0129111122223334444444444
loracle-cool-shit-reportsairtab-cesar-datasetSEM-Discce_ss
Dataset Card for "ce_ss"
More Information needed
verdict-oscillation-experiment
Verdict Oscillation Experiment
Track how Qwen3-8B's guilty/innocent verdict oscillates sentence-by-sentence during chain-of-thought reasoning about academic misconduct, and compare with an activation oracle's predictions from residual stream activations.
Setup
Model: Qwen3-8B (base, enable_thinking=False)
Oracle: Trained activation oracle (ceselder/cot-oracle-v15-stochastic), layers [9, 18, 27], stride=5
Questions: 15 academic misconduct scenarios (5 clearly guilty, 5… See the full description on the dataset page: https://huggingface.co/datasets/ceselder/verdict-oscillation-experiment.cot-oracle-ao-benchmarkSEM-Structurecesaret-pgartce_ss1
Dataset Card for "ce_ss1"
More Information needed
twitter-ws_yiyi-2025.09.28-1972287606802247937-Sl_9boVo_ceSNvcE-part1ceshi1SEM-Imagesheatmappngverdict-oscillation-v4EEE515_Segmentation
