datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
quantum-like-attention-framework-1.3b-untuned-validation
Quantum Like Attention Framework (Q.L.A.F) 1.3b untuned
This repository contains the model checkpoints, downstream evaluation scores, and pretraining convergence logs for the Quantum Like Attention Framework (Q.L.A.F) 1.3B configuration.
Key Specifications & Architecture
Model Name: Q.L.A.F 1.3b untuned (Quantum Like Attention Framework - Hybrid Architecture)
Parameters: 1.3B parameters total configuration (327M active parameter student subset)
Layer Count: 12… See the full description on the dataset page: https://huggingface.co/datasets/IgnisCogitationis/quantum-like-attention-framework-1.3b-untuned-validation.poly-btc-orderbookpoly-sol-orderbookquantum-representations
Epsilon-Transformers Belief Analysis Dataset
This dataset contains trained neural network models and their corresponding belief state regression analysis from the Epsilon-Transformers project. The models were trained on four different stochastic processes and analyzed for their ability to learn and represent belief states.
See https://github.com/adamimos/epsilon-transformers/tree/quantum-public for codebase which generated this data.
Dataset Structure… See the full description on the dataset page: https://huggingface.co/datasets/SimplexAI/quantum-representations.gftquant-us-prices
quant-us-prices
市场: 美股
格式: parquet
目录: 按哈希分片到子目录(00-ff)
说明: 由本地下载任务持续补齐,仓库支持断点续传更新。
poly-xrp-orderbookbiro-ai-quantum-dataset
BIRO AI Quantum Dataset
Languages
English (primary)
Dataset Structure
Each example is a JSON object with two main fields:
text
The raw textual content, which may be:
Plain text from Wikipedia, papers, code, or textbooks
Structured JSON strings containing questions, answers, distractors, and explanations (from SciQ, CommonsenseQA, etc.)
Instruction‑response pairs in <s>[INST] ... [/INST] format
metadata
A dictionary… See the full description on the dataset page: https://huggingface.co/datasets/Robbiejr/biro-ai-quantum-dataset.open-economic-quant-research-data
Open Economic & Quant Research Data
Versioned research content for CasualLab, Macroeconomics, Mortgage Rate Lock-In and Housing Market Dynamics, Tariff Incidence, and Order Flow to Price Impact, including project code, publishable data, fixtures, reports, tests, and reproducibility documentation.
Repository structure
CasualLab/: causal inference and policy-simulation research content.
Macroeconomics/: vintage-aware forecasting and public-source adapter research… See the full description on the dataset page: https://huggingface.co/datasets/ShawnChamberlain/open-economic-quant-research-data.poly-eth-orderbookquantized-llama-3.1-leaderboard-v2-evals
Open LLM Leaderboard v2 Benchmark Results
This artifact contains all the data from evaluations of Neural Magic's quantized Llama-3.1 models.
These evaluations were produced with lm-evaluation-harness by running the following command:
lm_eval \
--model vllm \
--model_args pretrained="<model_path>",dtype=auto,add_bos_token=False,max_model_len=4096,tensor_parallel_size="<num_gpus>",gpu_memory_utilization=0.8,enable_chunked_prefill=True \
--apply_chat_template \… See the full description on the dataset page: https://huggingface.co/datasets/neuralmagic/quantized-llama-3.1-leaderboard-v2-evals.quant-a-share-prices
A-share OHLCV Snapshot
Exported at: 2026-04-24T05:26:11.824110Z
Files: 5351 parquet files (*.SH.parquet, *.SZ.parquet, *.BJ.parquet)
Source pipeline: quant-agent download_data.py (akshare primary + yfinance fallback)
VDR_Quantum
VDR_Quantum – Overview
VDR_Quantum is a curated multimodal dataset focused on quantum technical documents. It combines text and image data extracted from real scientific PDFs to support tasks such as RAG DSE, question answering, document search, and vision-language model training.
Dataset Composition
This dataset was created using our open-source tool VDR_pdf-to-parquet.Quantum-related PDFs were collected from public online sources. Each document was processed… See the full description on the dataset page: https://huggingface.co/datasets/racineai/VDR_Quantum.quantbench-leaderboard-data
QuantBench leaderboard data
Raw benchmark data behind the QuantBench leaderboard:
calibration-quality GPTQ/AWQ quantization results across model sizes, calibration
corpora, and GPU tiers. 331 rows (239 ok / 92 failed — failed
runs are published too; a documented failure is a finding, not noise).
Models
Qwen/Qwen2.5-1.5B-Instruct (1.5B)
HuggingFaceTB/SmolLM2-1.7B-Instruct (1.7B)
deepgrove/Bonsai (0.5B)
Qwen/Qwen2.5-3B-Instruct (3B) — licence pending, rows only, no… See the full description on the dataset page: https://huggingface.co/datasets/Mohaaxa/quantbench-leaderboard-data.QuantiPhy
QuantiPhy
Dataset Summary
QuantiPhy is a benchmark for evaluating whether vision–language models (VLMs) can perform quantitative physical inference from visual evidence, rather than producing plausible but ungrounded numerical guesses.
This repository contains the official test set of the QuantiPhy benchmark, consisting of 3,289 video–question (QA) pairs across 568 videos. Ground-truth answers are withheld to ensure fair evaluation.
Each instance requires a model… See the full description on the dataset page: https://huggingface.co/datasets/PaulineLi/QuantiPhy.binance-future-orderbookqlib_csi300
QuantaAlpha Qlib CSI300 Dataset
Usage reference:
Qlib market data and pre-computed HDF5 files for QuantaAlpha factor mining (A-share, CSI 300).
Dataset description
Filename
Description
daily_pv.h5
Adjusted daily price and volume data.
daily_pv_debug.h5
Debug subset (smaller) of price-volume data.
How to load from Hugging Face
from huggingface_hub import hf_hub_download
import pandas as pd
# Download a file from this dataset
path =… See the full description on the dataset page: https://huggingface.co/datasets/QuantaAlpha/qlib_csi300.AlphaPrompt-QuantumLullaby-PDF
Quantum Lullaby Books (PDF - Human Readable)
📚 44 books of philosophical framework addressing the 73% animal population decline.
📖 For Humans
Beautiful PDF formatting
Designed for reading, printing, sharing
Complete philosophical journey
All books PDF version
🔗 Also Available
Markdown (AI-optimized): AlphaPrompt-QuantumLullaby-Markdown
SFT Training Datasets: AlphaPrompt-Metatron-SFT
📥 Download Stats
This folder: [shows… See the full description on the dataset page: https://huggingface.co/datasets/AIMindLink/AlphaPrompt-QuantumLullaby-PDF.QuantiPhy-validation
QuantiPhy (Validation Set)
Dataset Summary
QuantiPhy is a benchmark for evaluating whether vision–language models (VLMs) can perform quantitative physical inference from visual evidence, rather than producing plausible but ungrounded numerical guesses.
This repository contains the official validation set of QuantiPhy, released to support model development, ablation studies, and preliminary evaluation.The validation set represents approximately 4% of the full benchmark and… See the full description on the dataset page: https://huggingface.co/datasets/PaulineLi/QuantiPhy-validation.Quantum
Dataset Summary
ILSVRC 2012, commonly known as 'ImageNet' is an image dataset organized according to the WordNet hierarchy. Each meaningful concept in WordNet, possibly described by multiple words or word phrases, is called a "synonym set" or "synset". There are more than 100,000 synsets in WordNet, majority of them are nouns (80,000+). ImageNet aims to provide on average 1000 images to illustrate each synset. Images of each concept are quality-controlled and human-annotated.
💡… See the full description on the dataset page: https://huggingface.co/datasets/Miku26727/Quantum.task_data
QuantCodeEval
A benchmark for evaluating LLM coding agents on quantitative-strategy code
reproduction from finance research papers.
Status: Anonymous artifact for the 30-task benchmark.
Release mirrors
The release is mirrored at two anonymous locations:
Hugging Face Datasets — complete anonymous release:
https://huggingface.co/datasets/quantcodeeval/task_data
anonymous.4open.science — browseable mirror:
https://anonymous.4open.science/r/QuantCodeEval-Anonymous… See the full description on the dataset page: https://huggingface.co/datasets/quantcodeeval/task_data.quant-fidelity-registry
Quantization Fidelity Registry
A public, schema'd, receipt-backed, cross-model index of quantization quality measurements.
It exists to answer one question that nothing else answers today: show me every measured quant of
model X, with its fidelity number and enough provenance to know whether that number means anything.
It is the sibling of 0xSero/local-ai-registry,
which answers how fast, how much VRAM, how much money. This one answers how faithful. Ids and the
huggingface… See the full description on the dataset page: https://huggingface.co/datasets/malaiwah/quant-fidelity-registry.Chat-TS-Quant-EvalWelcome to the Chat-TS-Quant-Eval Dataset
This dataset is intended to evaluate time-series reasoning.
The dataset consists of real-world time-series with synthetic text.
This version was used to do the quantitative eval in the paper below.
If you find useful please cite:
@misc{quinlan2025chattsenhancingmultimodalreasoning,
title={Chat-TS: Enhancing Multi-Modal Reasoning Over Time-Series and Natural Language Data},
author={Paul Quinlan and Qingguo Li and Xiaodan Zhu}… See the full description on the dataset page: https://huggingface.co/datasets/PaulQ1/Chat-TS-Quant-Eval.quant-rag-dataquant-h-share-prices
quant-h-share-prices
市场: 港股
格式: parquet
目录: 按哈希分片到子目录(00-ff)
说明: 由本地下载任务持续补齐,仓库支持断点续传更新。
quantized-retrieval-dataVanGogh_vs_TreeOilPainting_QuantumTorque_EnergyField_Analysis_Phase1_2025
Dataset Policy
VanGogh Vs. Tree Oil Painting: Quantum Torque Energy Field Analysis 2025
Structure Type
Free-form and Semi-structured Narrative
Core Principles
Each file is an independent analytical entity with its own identity.
Each file is the result of Autonomous AI–Human Co-analysis.
The structure is intentionally open, flexible, and adaptive, reflecting the natural reasoning process of the researcher, rather than forcing rigid… See the full description on the dataset page: https://huggingface.co/datasets/HaruthaiAi/VanGogh_vs_TreeOilPainting_QuantumTorque_EnergyField_Analysis_Phase1_2025.qmof_quantum
Dataset Details
Dataset Description
QMOF is a database of electronic properties of MOFs, assembled by Rosen et al.
Jablonka et al. added gas adsorption properties.
Curated by:
License: CC-BY-4.0
Dataset Sources
No links provided
Citation
BibTeX:
@article{Rosen_2021,
doi = {10.1016/j.matt.2021.02.015},
url = {https://doi.org/10.1016%2Fj.matt.2021.02.015},
year = 2021,
month = {may},
publisher = {Elsevier {BV}},
volume = {4},
number =… See the full description on the dataset page: https://huggingface.co/datasets/jablonkagroup/qmof_quantum.AlphaPrompt-QuantumLullaby-Markdown
Quantum Lullaby Books (Markdown - AI Optimized)
🤖 44 books in markdown format for AI training, RAG, and processing.
🤖 For AI Systems
Clean markdown formatting
Optimized for tokenization
Easy parsing for RAG/training
No formatting artifacts
🔗 Also Available
PDF (Human-readable): AlphaPrompt-QuantumLullaby-PDF
Extracted SFT Datasets: AlphaPrompt-Metatron-SFT
💡 Use Cases
RAG systems: Load as context documents
Training data:… See the full description on the dataset page: https://huggingface.co/datasets/AIMindLink/AlphaPrompt-QuantumLullaby-Markdown.quantem-data
quantem-data
This repo is currently used primarily by
quantem.widget
(tutorial notebooks and quantem.widget.datasets). Other QuantEM packages
may use it later. Keys, the checker, and these instructions can grow as
we take more Community pull requests. Use the current required keys for
new PRs.
Public electron-microscopy data. MIT license. Downloads need no token.
This page is the upload and download protocol.
Download
from quantem.widget.datasets import… See the full description on the dataset page: https://huggingface.co/datasets/bobleesj/quantem-data.
