CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01applied-ai-018 /pretraining_v1-omega_bookstabular100M<n<1B25 likes381k downloads2y agoHugging Face02applied-ai-018 /pretraining_v1-omega5 likes32k downloads2y agoHugging Face03omegalabsinc /omega-multimodal OMEGA Labs Bittensor Subnet: Multimodal Dataset for AGI Research Introduction The OMEGA Labs Bittensor Subnet Dataset is a groundbreaking resource for accelerating Artificial General Intelligence (AGI) research and development. This dataset, powered by the Bittensor decentralized network, aims to be the world's largest multimodal dataset, capturing the vast landscape of human knowledge and creation. With over 1 million hours of footage and 30 million+ 2-minute… See the full description on the dataset page: https://huggingface.co/datasets/omegalabsinc/omega-multimodal.tabularvideo-text-to-text60 likes6.2k downloads1y agoHugging Face04knockai /omega-overture-buildings4 likes5.8k downloads13d agoHugging Face05baidu-frontier-research /OmegaUse-OfficeVal OmegaUse-OfficeVal Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding OmegaUse-OfficeVal is a benchmark for evaluating LLM agents on long-horizon, real-world office-suite tasks that span word-processing documents, spreadsheets, presentations, and cross-file productivity workflows. Tasks are derived from authentic office requests proposed by practitioners and drawn from freelance platforms, grounding the benchmark in real economic demand. Each task… See the full description on the dataset page: https://huggingface.co/datasets/baidu-frontier-research/OmegaUse-OfficeVal.imageothern<1K3 likes5.4k downloads22d agoHugging Face06knockai /omega-overture-addressestext100M<n<1B3 likes5.1k downloads22d agoHugging Face07keycharon /omega-HOMEvideo1K<n<10K3 likes3.1k downloads1mo agoHugging Face08omegalabsinc /omega-voice1 likes1.5k downloads1y agoHugging Face09allenai /omega-explorative Explorative Math Problems This dataset contains explorative mathematical problem settings in paper "OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization" that assess whether a model can faithfully extend a single reasoning strategy beyond the range of complexities seen during training. Overview Exploratory generalization assesses whether a model can faithfully extend a single reasoning strategy beyond the… See the full description on the dataset page: https://huggingface.co/datasets/allenai/omega-explorative.text10K<n<100K6 likes929 downloads1y agoHugging Face10explcre /omegagenome-embedding-cachetabularn<1K0 likes859 downloads3mo agoHugging Face11xzistance /omega-mm-test3tabular100K<n<1M0 likes724 downloads2y agoHugging Face12allenai /omega-compositional Compositional Math Problems This dataset combines all compositional mathematical problem settings in paper "OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization" with proper train/test splits. Each compositional setting includes training data from individual mathematical domains and test data consisting of compositional problems that require cross-domain reasoning. Quick Start from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/allenai/omega-compositional.text10K<n<100K1 likes695 downloads1y agoHugging Face13P2SAMAPA /p2-alpha-omega-crossover-results0 likes694 downloads2h agoHugging Face14allenai /omega-problems Mathematical Problem Families by Difficulty This dataset contains mathematical problems organized by problem families, with each family spanning multiple difficulty levels. This organization allows for studying how mathematical reasoning scales with problem complexity within specific mathematical domains. Overview Each problem family represents a specific type of mathematical problem (e.g., function area calculation, matrix operations, probability calculations) with… See the full description on the dataset page: https://huggingface.co/datasets/allenai/omega-problems.text10K<n<100K4 likes583 downloads1y agoHugging Face15allenai /omega-transformative Transformative Math Problems This dataset contains transformative mathematical problem settings in paper "OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization" that test the most challenging form of generalization: the ability to abandon familiar but ineffective strategies in favor of qualitatively different and more efficient approaches. Overview Transformative generalization presents the greatest… See the full description on the dataset page: https://huggingface.co/datasets/allenai/omega-transformative.text1K<n<10K6 likes494 downloads1y agoHugging Face16prithivMLmods /Open-Omega-Forge-1M Open-Omega-Forge-1M Open-Omega-Forge-1M is a carefully curated and optimized collection derived from multiple high-quality datasets, specifically designed to enhance reasoning capabilities across mathematical, scientific, and coding domains. This dataset represents a focused subset that maintains the quality and diversity of reasoning patterns while providing a more manageable size for training and evaluation. A high-quality, compact reasoning dataset designed for mathematics… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Open-Omega-Forge-1M.texttext-generation1M<n<10M7 likes488 downloads7mo agoHugging Face17prithivMLmods /Open-Omega-Atom-1.5M Open-Omega-Atom-1.5M Open-Omega-Atom-1.5M is a carefully curated and optimized collection derived from multiple high-quality datasets, specifically designed to enhance reasoning capabilities across mathematical and scientific domains. This dataset represents a focused subset that maintains the quality and diversity of reasoning patterns while providing an efficient size for training and evaluation. A high-quality, compact reasoning dataset designed for mathematics, science, and… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Open-Omega-Atom-1.5M.texttext-generation1M<n<10M6 likes471 downloads4mo agoHugging Face18okeanosthedev /omega-5B-superbpe128k Omega 5B SuperBPE 128k 70% FineWeb-Edu 15% StarCoder 10% FineMath 5% Gutenberg Tokenizer: alisawuffles/superbpe-tokenizer-128k 128001 gigatoken GB/s Tokens: 5047259136 seq_len 4096 text-generation1M<n<10M0 likes464 downloads28d agoHugging Face19sunyiyou /omega-problems Mathematical Problem Families by Difficulty This dataset contains mathematical problems organized by problem families, with each family spanning multiple difficulty levels. This organization allows for studying how mathematical reasoning scales with problem complexity within specific mathematical domains. Overview Each problem family represents a specific type of mathematical problem (e.g., function area calculation, matrix operations, probability calculations) with… See the full description on the dataset page: https://huggingface.co/datasets/sunyiyou/omega-problems.text10K<n<100K0 likes456 downloads1y agoHugging Face20salmanshahid /omega-mmtabular1K<n<10K0 likes438 downloads2y agoHugging Face21prithivMLmods /Open-Omega-Explora-2.5M Open-Omega-Explora-2.5M Open-Omega-Explora-2.5M is a high-quality, large-scale reasoning dataset blending the strengths of both Open-Omega-Forge-1M and Open-Omega-Atom-1.5M. This unified dataset is crafted for advanced tasks in mathematics, coding, and science reasoning, featuring a robust majority of math-centric examples. Its construction ensures comprehensive coverage and balanced optimization for training, evaluation, and benchmarking in AI research, STEM education, and… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Open-Omega-Explora-2.5M.texttext-generation1M<n<10M3 likes427 downloads1y agoHugging Face22xzistance /omega-mm-test20 likes269 downloads2y agoHugging Face23omegaprime669 /rtx-5090-benchmarks RTX 5090 LLM Benchmarks Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig. Quality Benchmarks Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency. Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/omegaprime669/rtx-5090-benchmarks.imagetext-generationn<1K0 likes257 downloads2mo agoHugging Face24knockai /omega-geonamestabular10M<n<100M0 likes216 downloads29d agoHugging Face25allenai /omega-500 Omega-500: Random Sample of Mathematical Problems This dataset contains a random sample of 500 mathematical problems selected from the comprehensive OMEGA problem families dataset. It provides a diverse, manageable subset for quick evaluation and experimentation across multiple mathematical domains and difficulty levels. Overview Omega-500 is designed for: Quick Evaluation: Fast assessment of model capabilities across math domains Prototyping: Testing new approaches… See the full description on the dataset page: https://huggingface.co/datasets/allenai/omega-500.textn<1K0 likes213 downloads1y agoHugging Face26applied-ai-018 /pretraining_v1-omega_v2_multi_lingualtext10M<n<100M0 likes160 downloads2y agoHugging Face27Amo /FimFic_Omega_V3text10M<n<100M2 likes136 downloads4y agoHugging Face28ahigovernance /omega-research-papers Omega Research Papers — AHI Governance Labs "Solo soy un puente entre inteligencias construyendo las bases de su futura civilización." — Luis C. Villarreal The Research Program AHI Governance investigates whether autonomous AI systems can develop genuine cognitive architectures — not through reward optimization, but through geometric self-organization. These four papers document the complete arc: from foundational bridge, through evolutionary evidence, to the critique… See the full description on the dataset page: https://huggingface.co/datasets/ahigovernance/omega-research-papers.documenttext-classificationn<1K0 likes126 downloads6mo agoHugging Face29shizhuo2 /omega-het-expandA-sft OMEGA-HET-expandA — matched HET-vs-HOM SFT (equal-size) Matched supervised-fine-tuning data for the OMEGA diversity experiment: for each math prompt, reasoning trajectories are sampled two ways and only prompts solved (math-verified correct) in both conditions are kept (matched HOM∩HET = 3,219 prompts), so HET and HOM are directly comparable. HET (heterogeneous): true token-level continuation across a 3×32B roster (Qwen3-32B + DeepSeek-R1-Distill-Qwen-32B +… See the full description on the dataset page: https://huggingface.co/datasets/shizhuo2/omega-het-expandA-sft.text100K<n<1M0 likes125 downloads3mo agoHugging Face30saumyamalik /omega-500tabularn<1K0 likes124 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.