datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
pretraining_v1-omega_booksomega-multimodal
OMEGA Labs Bittensor Subnet: Multimodal Dataset for AGI Research
Introduction
The OMEGA Labs Bittensor Subnet Dataset is a groundbreaking resource for accelerating Artificial General Intelligence (AGI) research and development. This dataset, powered by the Bittensor decentralized network, aims to be the world's largest multimodal dataset, capturing the vast landscape of human knowledge and creation.
With over 1 million hours of footage and 30 million+ 2-minute… See the full description on the dataset page: https://huggingface.co/datasets/omegalabsinc/omega-multimodal.OmegaUse-OfficeVal
OmegaUse-OfficeVal
Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding
OmegaUse-OfficeVal is a benchmark for evaluating LLM agents on long-horizon,
real-world office-suite tasks that span word-processing documents, spreadsheets,
presentations, and cross-file productivity workflows. Tasks are derived from
authentic office requests proposed by practitioners and drawn from freelance
platforms, grounding the benchmark in real economic demand. Each task… See the full description on the dataset page: https://huggingface.co/datasets/baidu-frontier-research/OmegaUse-OfficeVal.omega-overture-addressesomega-explorative
Explorative Math Problems
This dataset contains explorative mathematical problem settings in paper "OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization" that assess whether a model can faithfully extend a single reasoning strategy beyond the range of complexities seen during training.
Overview
Exploratory generalization assesses whether a model can faithfully extend a single reasoning strategy beyond the… See the full description on the dataset page: https://huggingface.co/datasets/allenai/omega-explorative.omegagenome-embedding-cacheomega-mm-test3omega-compositional
Compositional Math Problems
This dataset combines all compositional mathematical problem settings in paper "OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization" with proper train/test splits. Each compositional setting includes training data from individual mathematical domains and test data consisting of compositional problems that require cross-domain reasoning.
Quick Start
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/allenai/omega-compositional.omega-problems
Mathematical Problem Families by Difficulty
This dataset contains mathematical problems organized by problem families, with each family spanning multiple difficulty levels. This organization allows for studying how mathematical reasoning scales with problem complexity within specific mathematical domains.
Overview
Each problem family represents a specific type of mathematical problem (e.g., function area calculation, matrix operations, probability calculations) with… See the full description on the dataset page: https://huggingface.co/datasets/allenai/omega-problems.omega-transformative
Transformative Math Problems
This dataset contains transformative mathematical problem settings in paper "OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization" that test the most challenging form of generalization: the ability to abandon familiar but ineffective strategies in favor of qualitatively different and more efficient approaches.
Overview
Transformative generalization presents the greatest… See the full description on the dataset page: https://huggingface.co/datasets/allenai/omega-transformative.Open-Omega-Forge-1M
Open-Omega-Forge-1M
Open-Omega-Forge-1M is a carefully curated and optimized collection derived from multiple high-quality datasets, specifically designed to enhance reasoning capabilities across mathematical, scientific, and coding domains. This dataset represents a focused subset that maintains the quality and diversity of reasoning patterns while providing a more manageable size for training and evaluation. A high-quality, compact reasoning dataset designed for mathematics… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Open-Omega-Forge-1M.Open-Omega-Atom-1.5M
Open-Omega-Atom-1.5M
Open-Omega-Atom-1.5M is a carefully curated and optimized collection derived from multiple high-quality datasets, specifically designed to enhance reasoning capabilities across mathematical and scientific domains. This dataset represents a focused subset that maintains the quality and diversity of reasoning patterns while providing an efficient size for training and evaluation. A high-quality, compact reasoning dataset designed for mathematics, science, and… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Open-Omega-Atom-1.5M.omega-problems
Mathematical Problem Families by Difficulty
This dataset contains mathematical problems organized by problem families, with each family spanning multiple difficulty levels. This organization allows for studying how mathematical reasoning scales with problem complexity within specific mathematical domains.
Overview
Each problem family represents a specific type of mathematical problem (e.g., function area calculation, matrix operations, probability calculations) with… See the full description on the dataset page: https://huggingface.co/datasets/sunyiyou/omega-problems.omega-mmOpen-Omega-Explora-2.5M
Open-Omega-Explora-2.5M
Open-Omega-Explora-2.5M is a high-quality, large-scale reasoning dataset blending the strengths of both Open-Omega-Forge-1M and Open-Omega-Atom-1.5M. This unified dataset is crafted for advanced tasks in mathematics, coding, and science reasoning, featuring a robust majority of math-centric examples. Its construction ensures comprehensive coverage and balanced optimization for training, evaluation, and benchmarking in AI research, STEM education, and… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Open-Omega-Explora-2.5M.rtx-5090-benchmarks
RTX 5090 LLM Benchmarks
Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig.
Quality Benchmarks
Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency.
Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/omegaprime669/rtx-5090-benchmarks.omega-geonamesomega-500
Omega-500: Random Sample of Mathematical Problems
This dataset contains a random sample of 500 mathematical problems selected from the comprehensive OMEGA problem families dataset. It provides a diverse, manageable subset for quick evaluation and experimentation across multiple mathematical domains and difficulty levels.
Overview
Omega-500 is designed for:
Quick Evaluation: Fast assessment of model capabilities across math domains
Prototyping: Testing new approaches… See the full description on the dataset page: https://huggingface.co/datasets/allenai/omega-500.pretraining_v1-omega_v2_multi_lingualFimFic_Omega_V3omega-het-expandA-sft
OMEGA-HET-expandA — matched HET-vs-HOM SFT (equal-size)
Matched supervised-fine-tuning data for the OMEGA diversity experiment: for each math prompt,
reasoning trajectories are sampled two ways and only prompts solved (math-verified correct) in both
conditions are kept (matched HOM∩HET = 3,219 prompts), so HET and HOM are directly comparable.
HET (heterogeneous): true token-level continuation across a 3×32B roster
(Qwen3-32B + DeepSeek-R1-Distill-Qwen-32B +… See the full description on the dataset page: https://huggingface.co/datasets/shizhuo2/omega-het-expandA-sft.omega-500omega-globalml-buildingsomega-het-sft-rl
OMEGA-HET-SFT-RL
Reasoning-trajectory corpora for studying whether heterogeneous (HET) multi-model
SFT data improves post-RL out-of-distribution generalization on OMEGA math vs
homogeneous (HOM) single-model data, under matched controls.
Conditions
HOM: trajectories generated by a single model (Qwen3-4B).
HET: trajectories composed via true token-level continuation across a roster of
7 reasoning models (each model resumes the previous model's own assistant turn… See the full description on the dataset page: https://huggingface.co/datasets/shizhuo2/omega-het-sft-rl.pretraining_v2-omega_v2_multi_lingualomega-explorative
Explorative Math Problems
This dataset contains explorative mathematical problem settings in paper "OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization" that assess whether a model can faithfully extend a single reasoning strategy beyond the range of complexities seen during training.
Overview
Exploratory generalization assesses whether a model can faithfully extend a single reasoning strategy beyond the… See the full description on the dataset page: https://huggingface.co/datasets/shuqike/omega-explorative.omega-combinedKisan_Call_Centre_Transcriptsomega-explorative-combined
Combined Explorative Math Problems
This dataset contains a unified version of all explorative mathematical problem settings from the paper "OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization".
Unlike the subset-based version, this dataset combines all problems from different mathematical domains into single unified splits, making it easier to train on all explorative problems together.
Overview… See the full description on the dataset page: https://huggingface.co/datasets/sunyiyou/omega-explorative-combined.omega-mm
