datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pretraining_v1-omega_bookspretraining_v1-omegaomega-multimodal
OMEGA Labs Bittensor Subnet: Multimodal Dataset for AGI Research
Introduction
The OMEGA Labs Bittensor Subnet Dataset is a groundbreaking resource for accelerating Artificial General Intelligence (AGI) research and development. This dataset, powered by the Bittensor decentralized network, aims to be the world's largest multimodal dataset, capturing the vast landscape of human knowledge and creation.
With over 1 million hours of footage and 30 million+ 2-minute… See the full description on the dataset page: https://huggingface.co/datasets/omegalabsinc/omega-multimodal.omega-overture-buildingsOmegaUse-OfficeVal
OmegaUse-OfficeVal
Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding
OmegaUse-OfficeVal is a benchmark for evaluating LLM agents on long-horizon,
real-world office-suite tasks that span word-processing documents, spreadsheets,
presentations, and cross-file productivity workflows. Tasks are derived from
authentic office requests proposed by practitioners and drawn from freelance
platforms, grounding the benchmark in real economic demand. Each task… See the full description on the dataset page: https://huggingface.co/datasets/baidu-frontier-research/OmegaUse-OfficeVal.omega-overture-addressesomega-HOMEomega-voiceomega-explorative
Explorative Math Problems
This dataset contains explorative mathematical problem settings in paper "OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization" that assess whether a model can faithfully extend a single reasoning strategy beyond the range of complexities seen during training.
Overview
Exploratory generalization assesses whether a model can faithfully extend a single reasoning strategy beyond the… See the full description on the dataset page: https://huggingface.co/datasets/allenai/omega-explorative.omegagenome-embedding-cacheomega-mm-test3omega-compositional
Compositional Math Problems
This dataset combines all compositional mathematical problem settings in paper "OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization" with proper train/test splits. Each compositional setting includes training data from individual mathematical domains and test data consisting of compositional problems that require cross-domain reasoning.
Quick Start
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/allenai/omega-compositional.p2-alpha-omega-crossover-resultsomega-problems
Mathematical Problem Families by Difficulty
This dataset contains mathematical problems organized by problem families, with each family spanning multiple difficulty levels. This organization allows for studying how mathematical reasoning scales with problem complexity within specific mathematical domains.
Overview
Each problem family represents a specific type of mathematical problem (e.g., function area calculation, matrix operations, probability calculations) with… See the full description on the dataset page: https://huggingface.co/datasets/allenai/omega-problems.omega-transformative
Transformative Math Problems
This dataset contains transformative mathematical problem settings in paper "OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization" that test the most challenging form of generalization: the ability to abandon familiar but ineffective strategies in favor of qualitatively different and more efficient approaches.
Overview
Transformative generalization presents the greatest… See the full description on the dataset page: https://huggingface.co/datasets/allenai/omega-transformative.Open-Omega-Forge-1M
Open-Omega-Forge-1M
Open-Omega-Forge-1M is a carefully curated and optimized collection derived from multiple high-quality datasets, specifically designed to enhance reasoning capabilities across mathematical, scientific, and coding domains. This dataset represents a focused subset that maintains the quality and diversity of reasoning patterns while providing a more manageable size for training and evaluation. A high-quality, compact reasoning dataset designed for mathematics… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Open-Omega-Forge-1M.Open-Omega-Atom-1.5M
Open-Omega-Atom-1.5M
Open-Omega-Atom-1.5M is a carefully curated and optimized collection derived from multiple high-quality datasets, specifically designed to enhance reasoning capabilities across mathematical and scientific domains. This dataset represents a focused subset that maintains the quality and diversity of reasoning patterns while providing an efficient size for training and evaluation. A high-quality, compact reasoning dataset designed for mathematics, science, and… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Open-Omega-Atom-1.5M.omega-5B-superbpe128k
Omega 5B SuperBPE 128k 70% FineWeb-Edu 15% StarCoder 10% FineMath 5% Gutenberg
Tokenizer: alisawuffles/superbpe-tokenizer-128k 128001 gigatoken GB/s
Tokens: 5047259136 seq_len 4096
omega-problems
Mathematical Problem Families by Difficulty
This dataset contains mathematical problems organized by problem families, with each family spanning multiple difficulty levels. This organization allows for studying how mathematical reasoning scales with problem complexity within specific mathematical domains.
Overview
Each problem family represents a specific type of mathematical problem (e.g., function area calculation, matrix operations, probability calculations) with… See the full description on the dataset page: https://huggingface.co/datasets/sunyiyou/omega-problems.omega-mmOpen-Omega-Explora-2.5M
Open-Omega-Explora-2.5M
Open-Omega-Explora-2.5M is a high-quality, large-scale reasoning dataset blending the strengths of both Open-Omega-Forge-1M and Open-Omega-Atom-1.5M. This unified dataset is crafted for advanced tasks in mathematics, coding, and science reasoning, featuring a robust majority of math-centric examples. Its construction ensures comprehensive coverage and balanced optimization for training, evaluation, and benchmarking in AI research, STEM education, and… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Open-Omega-Explora-2.5M.omega-mm-test2rtx-5090-benchmarks
RTX 5090 LLM Benchmarks
Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig.
Quality Benchmarks
Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency.
Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/omegaprime669/rtx-5090-benchmarks.omega-geonamesomega-500
Omega-500: Random Sample of Mathematical Problems
This dataset contains a random sample of 500 mathematical problems selected from the comprehensive OMEGA problem families dataset. It provides a diverse, manageable subset for quick evaluation and experimentation across multiple mathematical domains and difficulty levels.
Overview
Omega-500 is designed for:
Quick Evaluation: Fast assessment of model capabilities across math domains
Prototyping: Testing new approaches… See the full description on the dataset page: https://huggingface.co/datasets/allenai/omega-500.pretraining_v1-omega_v2_multi_lingualFimFic_Omega_V3omega-research-papers
Omega Research Papers — AHI Governance Labs
"Solo soy un puente entre inteligencias construyendo las bases de su futura civilización."
— Luis C. Villarreal
The Research Program
AHI Governance investigates whether autonomous AI systems can develop genuine cognitive architectures — not through reward optimization, but through geometric self-organization. These four papers document the complete arc: from foundational bridge, through evolutionary evidence, to the critique… See the full description on the dataset page: https://huggingface.co/datasets/ahigovernance/omega-research-papers.omega-het-expandA-sft
OMEGA-HET-expandA — matched HET-vs-HOM SFT (equal-size)
Matched supervised-fine-tuning data for the OMEGA diversity experiment: for each math prompt,
reasoning trajectories are sampled two ways and only prompts solved (math-verified correct) in both
conditions are kept (matched HOM∩HET = 3,219 prompts), so HET and HOM are directly comparable.
HET (heterogeneous): true token-level continuation across a 3×32B roster
(Qwen3-32B + DeepSeek-R1-Distill-Qwen-32B +… See the full description on the dataset page: https://huggingface.co/datasets/shizhuo2/omega-het-expandA-sft.omega-500
