datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
triviaqa_ftThe dataset contains a random 0.7/0.1/0.2 train/dev/test splits of triviaqa dataset from KILT https://github.com/facebookresearch/KILT for benchmarking embedding model fine-tuning.
gnosis-qwen3-8b-triviaqa-dapo-sft-2comp
Gnosis SFT — Qwen3-8B (TriviaQA + DAPO-Math)
This dataset contains SFT-style samples used to train Gnosis, a lightweight self-awareness head for LLM correctness detection.
It mixes two sources:
Math: open-r1/DAPO-Math-17k-Processed (2 model completions per question)
Trivia: mandarjoshi/trivia_qa (1 completion per question)
Each row includes the generated completion plus fields needed for correctness supervision (e.g., question/prompt, completion, and binary correctness labels).
gnosis-qwen3-4b-thinking-triviaqa-dapo-sft-2comp
Gnosis SFT — Qwen3-4B-Thinking-2507 (TriviaQA + DAPO-Math)
This dataset contains SFT-style samples used to train Gnosis, a lightweight self-awareness head for LLM correctness detection.
It mixes two sources:
Math: open-r1/DAPO-Math-17k-Processed (2 model completions per question)
Trivia: mandarjoshi/trivia_qa (1 completion per question)
Each row includes the generated completion plus fields needed for correctness supervision (e.g., question/prompt, completion, and binary… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/gnosis-qwen3-4b-thinking-triviaqa-dapo-sft-2comp.gnosis-qwen3-1_7b-triviaqa-dapo_train
Gnosis SFT — Qwen3-1.7B (TriviaQA + DAPO-Math)
Paper | Code
This dataset contains Qwen3-1.7B generated completions with binary correctness labels for training Gnosis, a lightweight self-awareness mechanism. Gnosis enables frozen LLMs to perform intrinsic self-verification by decoding signals from hidden states and attention patterns to predict the correctness of generated outputs.
It mixes two source benchmarks:
open-r1/DAPO-Math-17k-Processed (math reasoning)
mandarjoshi/trivia_qa… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/gnosis-qwen3-1_7b-triviaqa-dapo_train.gnosis-gpt-oss-20b-triviaqa-dapo-sft-2comp-lowthinking
Gnosis SFT — GPT-OSS-20B (TriviaQA + DAPO-Math)
This dataset contains SFT-style samples used to train Gnosis, a lightweight self-awareness head for LLM correctness detection.
Paper: Can LLMs Predict Their Own Failures? Self-Awareness via Internal Circuits
GitHub Repository: Amirhosein-gh98/Gnosis
Dataset Description
The dataset is designed to provide supervision for detecting model failures by inspecting internal states. It mixes two primary sources:
Math: Derived… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/gnosis-gpt-oss-20b-triviaqa-dapo-sft-2comp-lowthinking.qwen25-triviaqa-slopeqwen3-non-think-triviaqa-slopetriviaqa_all_OLMoE-1B-7B-0924-Instructmem_agent-model_based-memagent-1-5b-step1024-triviaqa-val-c27000-t2048-1000s-agnosticgnosis-qwen3-4b-instruct-triviaqa-dapo-sft-2comp
Gnosis SFT — Qwen3-4B-Instruct-2507 (TriviaQA + DAPO-Math)
This dataset contains Qwen3-4B-Instruct-2507 generated completions with binary correctness labels for training Gnosis, a lightweight self-awareness mechanism that enables frozen LLMs to perform intrinsic self-verification by decoding signals from hidden states and attention patterns.
Paper: Can LLMs Predict Their Own Failures? Self-Awareness via Internal Circuits
Repository: https://github.com/Amirhosein-gh98/Gnosis… See the full description on the dataset page: https://huggingface.co/datasets/AmirhoseinGH/gnosis-qwen3-4b-instruct-triviaqa-dapo-sft-2comp.qwq-triviaqa-slopetrivia-qa-kg-processedr1-triviaqa-slopeqwen3-think-triviaqa-slopebertscore-llama-3-3-70b-i-triviaqa-llama-memorization-val-c4096-t2048-1000s-agnostictriviaqa_all_Llama-3.1-8B-Instructtriviaqa_full_valid_w_paraphrasestriviaqa_all_Qwen2.5-7B-Instructtriviaqa_ask4confbertscore-llama-3-3-70b-i-triviaqa-llama-memorization-val-c2048-t1024-1000s-agnosticmem_agent-model_based-qwen3-1-5b-oldgrpo-2086-triviaqa-llama-memorization-val-c27000-t2048-1000striviaqa-llama-memorization-filtered-8583mem_agent-bertscore-qwen2-5-14b-i-triviaqa-llama-memorization-val-c2048-t512-1000s-agnosticmem_agent-model_based-memagent-1-5b-separate-step720-triviaqa-llama-memorization-val-c27000-t204TriviaQA_conflict_5_half_smalltriviaqa-filtered-filtered-unknown-12660samplesbertscore-llama-3-3-70b-i-triviaqa-llama-memorization-val-c2048-t2048-1000s-agnosticmem_agent-model_based-rl-memoryagent-14b-triviaqa-llama-memorization-val-c27000-t2048-1000s-agnomem_agent-model_based-qwen3-1-5b-oldgrpo-2086-triviaqa-val-c4096-t2048-1000s-agnosticexaone-deep-triviaqa-slope
