datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
quantum-like-attention-framework-1.3b-untuned-validation
Quantum Like Attention Framework (Q.L.A.F) 1.3b untuned
This repository contains the model checkpoints, downstream evaluation scores, and pretraining convergence logs for the Quantum Like Attention Framework (Q.L.A.F) 1.3B configuration.
Key Specifications & Architecture
Model Name: Q.L.A.F 1.3b untuned (Quantum Like Attention Framework - Hybrid Architecture)
Parameters: 1.3B parameters total configuration (327M active parameter student subset)
Layer Count: 12… See the full description on the dataset page: https://huggingface.co/datasets/IgnisCogitationis/quantum-like-attention-framework-1.3b-untuned-validation.DLR-Web
DLR-Web: Multidisciplinary Reasoning Dataset from Web Corpus [Project Page]
This repository releases the Design-Logic-Reasoning-Web (DLR-Web) dataset from the paper DESIGNER: Design-Logic-Guided Multidisciplinary Data Synthesis for LLM Reasoning (ICLR 2026).
Field definitions
original_document: web-sourced raw document text, further filtered from FineFineWeb; thanks to the FineFineWeb authors and maintainers for providing this resource
design_logic: Design Logic in… See the full description on the dataset page: https://huggingface.co/datasets/Attention1115/DLR-Web.MLVU-Attention-Review
MLVU Attention Review 下载与解压说明
本仓库提供一个未压缩 TAR,里面是完整静态 HTML 展示、所有页面所需图片、QA/GT、模型原始回答、attention 元数据与未归一化 grids.npz,并附离线查看和逐文件 SHA256 校验脚本。归档不含模型权重、完整源视频或推理环境。
容量要求
下载和解压期间需同时存放 TAR 与解压内容,请预留至少归档大小约 2.2 倍的可用空间。
文件系统必须支持大于 4 GB 的单文件(NTFS、exFAT、ext4、APFS 等;FAT32 不可用)。
Windows 建议在较短路径中操作,例如 D:\reviews\MLVU,避免长路径限制。
下载
安装最新版 Hugging Face CLI:
pip install -U huggingface_hub
hf download GraciaChen/MLVU-Attention-Review MLVU-Attention-Review.tar SHA256SUMS… See the full description on the dataset page: https://huggingface.co/datasets/GraciaChen/MLVU-Attention-Review.PARTIAL-STAR41K-DeepSeek-R1-Distill-Qwen-7B-Size-16-Blockwise-InterIntra-Attention-0_8192fpga_cost_model_kernel_data_attention_p2
FPGA HLS Kernel Cost-Model Data
Evolved Vitis HLS C++ kernels paired with their ground-truth Vitis HLS
csynth results. Each row is one generated program from an evolutionary FPGA
optimisation run, linked to its kernel source, evaluator report.json, and raw
synthesis report.
Each row carries a split label: train marks the original benchmarks used
to fit the analytical cost model's learned correction term, and holdout marks
benchmarks added afterwards that were not used for… See the full description on the dataset page: https://huggingface.co/datasets/adimnaku/fpga_cost_model_kernel_data_attention_p2.VideoMMEv2-Attention-Review
VideoMMEv2 Attention Review 下载与解压说明
本仓库提供一个未压缩 TAR,里面是完整静态 HTML 展示、所有页面所需图片、QA/GT、模型原始回答、attention 元数据与未归一化 grids.npz,并附离线查看和逐文件 SHA256 校验脚本。归档不含模型权重、完整源视频或推理环境。
容量要求
下载和解压期间需同时存放 TAR 与解压内容,请预留至少归档大小约 2.2 倍的可用空间。
文件系统必须支持大于 4 GB 的单文件(NTFS、exFAT、ext4、APFS 等;FAT32 不可用)。
Windows 建议在较短路径中操作,例如 D:\reviews\VideoMMEv2,避免长路径限制。
下载
安装最新版 Hugging Face CLI:
pip install -U huggingface_hub
hf download GraciaChen/VideoMMEv2-Attention-Review… See the full description on the dataset page: https://huggingface.co/datasets/GraciaChen/VideoMMEv2-Attention-Review.mailabs_fywrepro-attention-fw-traces
Agent traces
Agent sessions published from a Trackio Logbook.
attention-uq-800q-colab
Attention/UQ 800-question Colab bundle
A deterministic 200-question subset for each of MultiModalQA, WebQA, HotpotQA, and TAT-QA. See manifest.json for exact upstream sources, hashes, counts, and the explicitly constructed WebQA distractor setting.
cy-sieve-attention-benchmark
CY-Sieve Attention — GPU benchmark results (NVIDIA L4, 2026-06-22)
Benchmark artifacts for the CY-Sieve positional-attention kernel, a falsifiable
engineering experiment from the Mirror-Map-Sieve
project. The bias derives from the weight-5 Apéry-like sequence
$S_{20}(n)=\sum_k \binom{n}{k}^4\binom{n+k}{k}$ (a Calabi–Yau 3-fold period;
the geometry fixes the long-range decay slope $\log\lambda=3.762$ and curvature
$\beta=2$).
⚠️ Headline: this is a documented NEGATIVE… See the full description on the dataset page: https://huggingface.co/datasets/callensxavier/cy-sieve-attention-benchmark.2026.RA.QKV-Attention-Interface
2026.RA.QKV-Attention-Interface
Tables behind the Q/K/V attention-interface ladder: a four-rung preregistered study asking whether you can make a language-model negotiator more rational by changing what its attention reads rather than what its prompt says. Base model Qwen/Qwen3-8B (frozen, bf16, thinking off) playing one seat in a six-party, five-issue negotiation.
Headline: the interface level was the whole story. The same fair-and-efficient candidate deal — the Nash bargaining… See the full description on the dataset page: https://huggingface.co/datasets/siddharthmb/2026.RA.QKV-Attention-Interface.Vibe-Coding-Claude-Fable-5attention-heatmap-visualizer
Attention Heatmap Visualizer 3.1.0
A Python scripts to generate full attention-head heat-maps for transformer-based Language Models. They show "where the model is looking" or "what tokens/features are most relevant" when processing a specific input element.
By analyzing these heatmaps across all layers and heads you can gain insights into how the model processes information, identifies relationships between tokens, and prioritizes specific parts of the input during inference.… See the full description on the dataset page: https://huggingface.co/datasets/ronniross/attention-heatmap-visualizer.attention-not-self
Attention, Not Self — Knowledge Graph
JSON-LD knowledge graph encoding the concept layer of the Attention, Not Self research line — a personal essay collection and structured cross-tradition mapping between five Buddhist traditions (Theravāda, Sarvāstivāda, Yogācāra, Chan/Zen, Pure Land) and contemporary frameworks in computational phenomenology (predictive processing, active inference, Global Workspace Theory, Parallel Distributed Processing).
What this dataset is… See the full description on the dataset page: https://huggingface.co/datasets/shimo4228/attention-not-self.OpenR1-Distillation-DeepSeek-R1-Distill-Qwen-7B-Size-16-Blockwise_Reasoning-Attentionzzsn-activations-3_unet.up_blocks.1.attentions.1attention_testdesigner-design-logics
DESIGNER: Design Logic Library [Project Page]
This repository contains a library of Mermaid-format Design Logics used in the paper DESIGNER: Design-Logic-Guided Multidisciplinary Data Synthesis for LLM Reasoning (ICLR 2026).
Field definitions
mermaid: Design Logic in Mermaid format, abstracted from the source question, which is a human-authored high-difficulty question.
difficulty: difficulty label of the source question
type: type label of the source question… See the full description on the dataset page: https://huggingface.co/datasets/Attention1115/designer-design-logics.taf-attention-decay
TAF Attention-Decay Measurements
First public dataset of attention-decay exponent γ measurements
across transformer LLMs.
Companion to the Thermodynamic Attention Framework (TAF) papers by
Carles Marín (2026):
Paper I: 10.5281/zenodo.19826343 — Predicting How Transformers Attend
Paper II: 10.5281/zenodo.19960573 — A Six-Axis Decomposition with the Learned Imprint, Sink-Dominated Precision Boundaries, Bimodal Phase Structure, and Honest Revisions
Paper III:… See the full description on the dataset page: https://huggingface.co/datasets/karlexmarin/taf-attention-decay.social-robotics-attention
Social Robotics: Attention / Engagement (03a)
Does the bystander look at the camera-wearer? Per-bystander gaze + head-pose engagement around each task.
One layer of the Social-Affective Filter (SAF) — dehydrated social-signal metadata extracted from egocentric (first-person) video so robots can learn to read human reactions. No raw pixels and no audio. Each row is one source video, keyed by video_id; rehydrate against your own legally-obtained Ego4D copies (below).
Rows: 828… See the full description on the dataset page: https://huggingface.co/datasets/louisye/social-robotics-attention.repro-softmax-as-linear-attention-in-the-large-prompt-regime-a-measure-based-perspecti-traces
Agent traces
Agent sessions published from a Trackio Logbook.
repro-hybrid-linear-full-attention-code
Reproduction: A Provable Expressiveness Hierarchy in Hybrid Linear-Full Attention
Paper: arXiv:2602.01763 · OpenReview JPnA2BI6U5
Authors: Xiaowei Ye, Xiaoyu He, Chao Liao, Chen Wu, Pinyan Lu
Scope
This is a theory paper (communication-complexity lower bounds). There is no official training code or dataset. This reproduction:
Checks statement fidelity of Theorems 1.1–1.2 and Tables 1–2 against the PDF.
Verifies algebraic consistency of the (Hdp \le n^{2^{-4L-2}})… See the full description on the dataset page: https://huggingface.co/datasets/junwatu/repro-hybrid-linear-full-attention-code.attention_attrib_lmgpt5-attention-is-all-you-need-quiz
Dataset Card for "Attention is all you need" quiz
GPT5-Thinking generated synthetic dataset of 15 questions, choices and answers.
Prompt used for generation
https://arxiv.org/pdf/1706.03762 Create a quiz of 15 questions with answers in the form of a hf dataset.
Dataset Card Contact
IDK
general_knowledge_dataset
Synthetic MMLU CoT
This dataset contains 27,689 synthetic chain-of-thought examples
generated with Qwen/Qwen3-14B on cais/mmlu auxiliary_train
multiple-choice questions.
Columns
question
choices
answer
answer_letter
teacher_output
Provenance and License
The original questions, answer choices, and gold labels come from
cais/mmlu, split auxiliary_train. The
Hugging Face dataset card for cais/mmlu lists its license as
mit. Those source fields retain… See the full description on the dataset page: https://huggingface.co/datasets/cs-552-2026-AttentionSeekers/general_knowledge_dataset.TEMP-STAR-41K-Distillation-DeepSeek-R1-Distill-Qwen-7B-Size-16-Blockwise_Inter-Intra-Attentionattention_attrib_lm_universitysciqgeneral_sft_datasetSTAR-41K-Distillation-DeepSeek-R1-Distill-Qwen-7B-Size-16-Blockwise_Reasoning-Attention
