CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01applied-ai-018 /pretraining_v1-omega_bookstabular100M<n<1B25 likes381k downloads2y agoHugging Face02omegalabsinc /omega-multimodal OMEGA Labs Bittensor Subnet: Multimodal Dataset for AGI Research Introduction The OMEGA Labs Bittensor Subnet Dataset is a groundbreaking resource for accelerating Artificial General Intelligence (AGI) research and development. This dataset, powered by the Bittensor decentralized network, aims to be the world's largest multimodal dataset, capturing the vast landscape of human knowledge and creation. With over 1 million hours of footage and 30 million+ 2-minute… See the full description on the dataset page: https://huggingface.co/datasets/omegalabsinc/omega-multimodal.tabularvideo-text-to-text60 likes6.2k downloads1y agoHugging Face03baidu-frontier-research /OmegaUse-OfficeVal OmegaUse-OfficeVal Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding OmegaUse-OfficeVal is a benchmark for evaluating LLM agents on long-horizon, real-world office-suite tasks that span word-processing documents, spreadsheets, presentations, and cross-file productivity workflows. Tasks are derived from authentic office requests proposed by practitioners and drawn from freelance platforms, grounding the benchmark in real economic demand. Each task… See the full description on the dataset page: https://huggingface.co/datasets/baidu-frontier-research/OmegaUse-OfficeVal.imageothern<1K3 likes5.4k downloads22d agoHugging Face04knockai /omega-overture-addressestext100M<n<1B3 likes5.1k downloads22d agoHugging Face05allenai /omega-explorative Explorative Math Problems This dataset contains explorative mathematical problem settings in paper "OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization" that assess whether a model can faithfully extend a single reasoning strategy beyond the range of complexities seen during training. Overview Exploratory generalization assesses whether a model can faithfully extend a single reasoning strategy beyond the… See the full description on the dataset page: https://huggingface.co/datasets/allenai/omega-explorative.text10K<n<100K6 likes929 downloads1y agoHugging Face06explcre /omegagenome-embedding-cachetabularn<1K0 likes859 downloads3mo agoHugging Face07xzistance /omega-mm-test3tabular100K<n<1M0 likes724 downloads2y agoHugging Face08allenai /omega-compositional Compositional Math Problems This dataset combines all compositional mathematical problem settings in paper "OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization" with proper train/test splits. Each compositional setting includes training data from individual mathematical domains and test data consisting of compositional problems that require cross-domain reasoning. Quick Start from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/allenai/omega-compositional.text10K<n<100K1 likes695 downloads1y agoHugging Face09allenai /omega-problems Mathematical Problem Families by Difficulty This dataset contains mathematical problems organized by problem families, with each family spanning multiple difficulty levels. This organization allows for studying how mathematical reasoning scales with problem complexity within specific mathematical domains. Overview Each problem family represents a specific type of mathematical problem (e.g., function area calculation, matrix operations, probability calculations) with… See the full description on the dataset page: https://huggingface.co/datasets/allenai/omega-problems.text10K<n<100K4 likes583 downloads1y agoHugging Face10allenai /omega-transformative Transformative Math Problems This dataset contains transformative mathematical problem settings in paper "OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization" that test the most challenging form of generalization: the ability to abandon familiar but ineffective strategies in favor of qualitatively different and more efficient approaches. Overview Transformative generalization presents the greatest… See the full description on the dataset page: https://huggingface.co/datasets/allenai/omega-transformative.text1K<n<10K6 likes494 downloads1y agoHugging Face11prithivMLmods /Open-Omega-Forge-1M Open-Omega-Forge-1M Open-Omega-Forge-1M is a carefully curated and optimized collection derived from multiple high-quality datasets, specifically designed to enhance reasoning capabilities across mathematical, scientific, and coding domains. This dataset represents a focused subset that maintains the quality and diversity of reasoning patterns while providing a more manageable size for training and evaluation. A high-quality, compact reasoning dataset designed for mathematics… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Open-Omega-Forge-1M.texttext-generation1M<n<10M7 likes488 downloads7mo agoHugging Face12prithivMLmods /Open-Omega-Atom-1.5M Open-Omega-Atom-1.5M Open-Omega-Atom-1.5M is a carefully curated and optimized collection derived from multiple high-quality datasets, specifically designed to enhance reasoning capabilities across mathematical and scientific domains. This dataset represents a focused subset that maintains the quality and diversity of reasoning patterns while providing an efficient size for training and evaluation. A high-quality, compact reasoning dataset designed for mathematics, science, and… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Open-Omega-Atom-1.5M.texttext-generation1M<n<10M6 likes471 downloads4mo agoHugging Face13sunyiyou /omega-problems Mathematical Problem Families by Difficulty This dataset contains mathematical problems organized by problem families, with each family spanning multiple difficulty levels. This organization allows for studying how mathematical reasoning scales with problem complexity within specific mathematical domains. Overview Each problem family represents a specific type of mathematical problem (e.g., function area calculation, matrix operations, probability calculations) with… See the full description on the dataset page: https://huggingface.co/datasets/sunyiyou/omega-problems.text10K<n<100K0 likes456 downloads1y agoHugging Face14salmanshahid /omega-mmtabular1K<n<10K0 likes438 downloads2y agoHugging Face15prithivMLmods /Open-Omega-Explora-2.5M Open-Omega-Explora-2.5M Open-Omega-Explora-2.5M is a high-quality, large-scale reasoning dataset blending the strengths of both Open-Omega-Forge-1M and Open-Omega-Atom-1.5M. This unified dataset is crafted for advanced tasks in mathematics, coding, and science reasoning, featuring a robust majority of math-centric examples. Its construction ensures comprehensive coverage and balanced optimization for training, evaluation, and benchmarking in AI research, STEM education, and… See the full description on the dataset page: https://huggingface.co/datasets/prithivMLmods/Open-Omega-Explora-2.5M.texttext-generation1M<n<10M3 likes427 downloads1y agoHugging Face16omegaprime669 /rtx-5090-benchmarks RTX 5090 LLM Benchmarks Speed and quality benchmarks for quantized LLMs on NVIDIA RTX 5090 32GB, measured with llm-bench-rig. Quality Benchmarks Generative evaluation through llama-server chat completions. Replicates standard benchmark methodology using custom evaluators — no lm-evaluation-harness dependency. Results are split by reasoning mode: comparing a thinking-on (reasoning) model's quality against a thinking-off model is apples-to-oranges, so the two groups… See the full description on the dataset page: https://huggingface.co/datasets/omegaprime669/rtx-5090-benchmarks.imagetext-generationn<1K0 likes257 downloads2mo agoHugging Face17knockai /omega-geonamestabular10M<n<100M0 likes216 downloads29d agoHugging Face18allenai /omega-500 Omega-500: Random Sample of Mathematical Problems This dataset contains a random sample of 500 mathematical problems selected from the comprehensive OMEGA problem families dataset. It provides a diverse, manageable subset for quick evaluation and experimentation across multiple mathematical domains and difficulty levels. Overview Omega-500 is designed for: Quick Evaluation: Fast assessment of model capabilities across math domains Prototyping: Testing new approaches… See the full description on the dataset page: https://huggingface.co/datasets/allenai/omega-500.textn<1K0 likes213 downloads1y agoHugging Face19applied-ai-018 /pretraining_v1-omega_v2_multi_lingualtext10M<n<100M0 likes160 downloads2y agoHugging Face20Amo /FimFic_Omega_V3text10M<n<100M2 likes136 downloads4y agoHugging Face21shizhuo2 /omega-het-expandA-sft OMEGA-HET-expandA — matched HET-vs-HOM SFT (equal-size) Matched supervised-fine-tuning data for the OMEGA diversity experiment: for each math prompt, reasoning trajectories are sampled two ways and only prompts solved (math-verified correct) in both conditions are kept (matched HOM∩HET = 3,219 prompts), so HET and HOM are directly comparable. HET (heterogeneous): true token-level continuation across a 3×32B roster (Qwen3-32B + DeepSeek-R1-Distill-Qwen-32B +… See the full description on the dataset page: https://huggingface.co/datasets/shizhuo2/omega-het-expandA-sft.text100K<n<1M0 likes125 downloads3mo agoHugging Face22saumyamalik /omega-500tabularn<1K0 likes124 downloads1y agoHugging Face23knockai /omega-globalml-buildingstabular10M<n<100M0 likes116 downloads29d agoHugging Face24shizhuo2 /omega-het-sft-rl OMEGA-HET-SFT-RL Reasoning-trajectory corpora for studying whether heterogeneous (HET) multi-model SFT data improves post-RL out-of-distribution generalization on OMEGA math vs homogeneous (HOM) single-model data, under matched controls. Conditions HOM: trajectories generated by a single model (Qwen3-4B). HET: trajectories composed via true token-level continuation across a roster of 7 reasoning models (each model resumes the previous model's own assistant turn… See the full description on the dataset page: https://huggingface.co/datasets/shizhuo2/omega-het-sft-rl.texttext-generation100K<n<1M0 likes89 downloads4mo agoHugging Face25applied-ai-018 /pretraining_v2-omega_v2_multi_lingualtext10M<n<100M0 likes83 downloads2y agoHugging Face26shuqike /omega-explorative Explorative Math Problems This dataset contains explorative mathematical problem settings in paper "OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization" that assess whether a model can faithfully extend a single reasoning strategy beyond the range of complexities seen during training. Overview Exploratory generalization assesses whether a model can faithfully extend a single reasoning strategy beyond the… See the full description on the dataset page: https://huggingface.co/datasets/shuqike/omega-explorative.text10K<n<100K1 likes82 downloads9mo agoHugging Face27hamishivi /omega-combinedtext10K<n<100K0 likes70 downloads1y agoHugging Face28Omegaindebt /Kisan_Call_Centre_Transcriptstabular1M<n<10M0 likes61 downloads1y agoHugging Face29sunyiyou /omega-explorative-combined Combined Explorative Math Problems This dataset contains a unified version of all explorative mathematical problem settings from the paper "OMEGA: Can LLMs Reason Outside the Box in Math? Evaluating Exploratory, Compositional, and Transformative Generalization". Unlike the subset-based version, this dataset combines all problems from different mathematical domains into single unified splits, making it easier to train on all explorative problems together. Overview… See the full description on the dataset page: https://huggingface.co/datasets/sunyiyou/omega-explorative-combined.text10K<n<100K0 likes54 downloads1y agoHugging Face30xzistance /omega-mmtabular1K<n<10K0 likes50 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.