CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Infatoshi /kernelbench-mega-traces KernelBench-Mega agent traces Coding agents writing full GPU megakernels across Blackwell / H100 / B200, scored as speedup over reference; contamination-audited (23 verified cells). Each .jsonl file is one agent run in Claude-Code session format, viewable with the agent trace viewer. Filename = run id; manifest.csv maps each run to model / harness / problem / GPU / score. 23 agent traces · live leaderboard: https://kernelbench.com/mega Secrets redacted. Full reasoning for… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-mega-traces.tabularn<1K18 likes6.7k downloads7h agoHugging Face02thethanksforthegod /quran-asr-mega-corpustabular10K<n<100K1 likes1.2k downloads14d agoHugging Face03GindaChen /megatron-prof-data-v5text10K<n<100K0 likes885 downloads1y agoHugging Face04hltcoe /megawika-report-generation Dataset Card for MegaWika for Report Generation Dataset Summary MegaWika is a multi- and crosslingual text dataset containing 30 million Wikipedia passages with their scraped and cleaned web citations. The passages span 50 Wikipedias in 50 languages, and the articles in which the passages were originally embedded are included for convenience. Where a Wikipedia passage is in a non-English language, an automated English translation is provided. This dataset provides the… See the full description on the dataset page: https://huggingface.co/datasets/hltcoe/megawika-report-generation.textsummarization100K<n<1M6 likes861 downloads3y agoHugging Face05megagonlabs /cypherbench CypherBench CypherBench is a benchmark designed to evaluate text-to-Cypher translation for large language models (LLMs). It includes: 11 large-scale Neo4j property graphs transformed from Wikidata with 7.8 million entities. Over 10,000 (question, Cypher) pairs for training/evaluating text-to-Cypher translation. Paper: https://arxiv.org/pdf/2412.18702 Repository & Demo: https://github.com/megagonlabs/cypherbench Contact: yanlin@megagon.ai Sample Task { "qid":… See the full description on the dataset page: https://huggingface.co/datasets/megagonlabs/cypherbench.text10K<n<100K11 likes712 downloads1y agoHugging Face06metaeval /mega-acceptability-v2text10K<n<100K0 likes508 downloads4y agoHugging Face07ItalianNarratives /megamatt-translated-ITtext100K<n<1M0 likes498 downloads15d agoHugging Face08GindaChen /megatron-prof-data-v6text100K<n<1M0 likes463 downloads1y agoHugging Face09UniversityOfMontanaSAL /Rustins_Super_Mega_Awesome_VEDU_Model Rustin's Super Mega Awesome VEDU Model A reproducible, heavily-documented pipeline that maps Ventenata dubia ("VEDU", an invasive winter-annual grass) across Montana from satellite + environmental data. Science reference: docs/VEDU_48_predictors_detailed.md Data decisions & gotchas: docs/CONTRADICTIONS.md Parity with the Earth Engine build: docs/GEE_PARITY.md Continue-the-build guide: docs/HANDOFF.md Label inventory: docs/DATA_SOURCES.md What it produces 57… See the full description on the dataset page: https://huggingface.co/datasets/UniversityOfMontanaSAL/Rustins_Super_Mega_Awesome_VEDU_Model.imagen<1K0 likes443 downloads9d agoHugging Face10GindaChen /megatron-prof-data-v14text100K<n<1M0 likes294 downloads1y agoHugging Face11GindaChen /megatron-prof-data-v7text10K<n<100K0 likes223 downloads1y agoHugging Face12Mr-Philo /dolma3_dolmino_megatron_tokenize Dolma 3 / Dolmino Megatron-LM indexed dataset This repository contains immutable Megatron-LM indexed datasets (.bin and .idx) produced from pinned Dolma 3 and Dolmino releases. It intentionally contains no training checkpoints, experiment outputs, logs, or dataset caches. The indexed payloads were derived from these pinned public datasets: allenai/dolma3_mix-150B-1025@afa92bfb22366821c5e6cd427cdd036b34b713ef… See the full description on the dataset page: https://huggingface.co/datasets/Mr-Philo/dolma3_dolmino_megatron_tokenize.tabulartext-generationn<1K0 likes191 downloads23d agoHugging Face13GindaChen /megatron-prof-data-v3text10K<n<100K0 likes186 downloads1y agoHugging Face14GindaChen /megatron-prof-data-v4text1K<n<10K0 likes176 downloads1y agoHugging Face15nyu-dice-lab /lm-eval-results-Eurdem-megatron_2.1_MoE_2x7B-private Dataset Card for Evaluation run of Eurdem/megatron_2.1_MoE_2x7B Dataset automatically created during the evaluation run of model Eurdem/megatron_2.1_MoE_2x7B The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Eurdem-megatron_2.1_MoE_2x7B-private.tabular100K<n<1M0 likes131 downloads2y agoHugging Face16GindaChen /megatron-prof-data-v12text100K<n<1M0 likes108 downloads1y agoHugging Face17tasal9 /ZamAI-Pashto-Mega-Dataset ZamAI Pashto Mega Dataset Languages: psLicense: apache-2.0Task categories: text-generation, summarization, question-answeringSize categories: 1M<n<10M Summary This dataset is part of the ZamAI Pashto data collection. It is intended for text-generation, summarization, question-answering tasks in Pashto. How to use from datasets import load_dataset dataset = load_dataset("tasal9/ZamAI-Pashto-Mega-Dataset") print(dataset) Configs… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/ZamAI-Pashto-Mega-Dataset.texttext-generation1M<n<10M0 likes63 downloads2mo agoHugging Face18AksaraLLM /aksara-mega-sft Aksara Mega SFT — 64K+ Dataset Grade S AksaraLLM Community mempersembahkan 64797 pasangan instruksi Indonesia kualitas tinggi. Highlight Dataset Pemahaman 10 Bahasa Daerah (Jawa, Sunda, Minang, dll) via NusaX Ribuan QA Suku, Agama, dan Budaya Nusantara SQuAD ID & Dolly 15K Indonesian Orca Math Word Problems Indonesian Guanaco & xP3x High Quality Conversations 38 Provinsi Lengkap & Etika Lokal texttext-generation10K<n<100K0 likes61 downloads5mo agoHugging Face19open-llm-leaderboard /Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8-detailsgated Dataset Card for Evaluation run of Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8 Dataset automatically created during the evaluation run of model Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8-details.tabular10K<n<100K0 likes58 downloads2y agoHugging Face20open-llm-leaderboard /Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8.9-detailsgated Dataset Card for Evaluation run of Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8.9 Dataset automatically created during the evaluation run of model Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8.9 The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8.9-details.tabular10K<n<100K0 likes58 downloads2y agoHugging Face21pandakingpunc /megazeka-tr-spellfix-pairs Megazeka · Turkish Spelling Correction Pairs 40,313 (misspelled → correct) Turkish sentence pairs with a full record of every corruption that was applied, generated from the CC0 Common Voice Turkish Sentence Collector and from project-authored everyday first/second-person sentences. The distinguishing feature is the operations field: each pair carries the exact sequence of noise transformations that produced it, with before/after text at each step. That makes it possible to… See the full description on the dataset page: https://huggingface.co/datasets/pandakingpunc/megazeka-tr-spellfix-pairs.text10K<n<100K0 likes51 downloads14h agoHugging Face22andreaskoepf /megacode3-min100text1M<n<10M1 likes44 downloads3y agoHugging Face23open-llm-leaderboard /prithivMLmods__Megatron-Corpus-14B-Exp-detailsgated Dataset Card for Evaluation run of prithivMLmods/Megatron-Corpus-14B-Exp Dataset automatically created during the evaluation run of model prithivMLmods/Megatron-Corpus-14B-Exp The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prithivMLmods__Megatron-Corpus-14B-Exp-details.tabular10K<n<100K0 likes44 downloads2y agoHugging Face24tyzhu /megamath-web-pro-max-splittedtabular10M<n<100M0 likes39 downloads3d agoHugging Face25rombodawg /MegaCodeTraining VERSION 3 IS RELEASED DOWNLOAD HERE: https://huggingface.co/datasets/rombodawg/LosslessMegaCodeTrainingV3_2.2m_Evol This is a uncensored mega combined dataset using both razent/wizardlm-code-evol-32k and nickrosh/Evol-Instruct-Code-80k-v1 In this version many lines of instructions were removed in part of a uncensoring process. The Rombo's format.rar file is so you can use the training data in oobagooba text generation webui. Simply unzip it, and use it as a json file. All links bellow… See the full description on the dataset page: https://huggingface.co/datasets/rombodawg/MegaCodeTraining.text100K<n<1M14 likes38 downloads3y agoHugging Face26megagonlabs /recap RECAP RECAP: REwriting Conversations for Intent Understanding in Agentic Planning 📄 paper link Kushan Mitra, Dan Zhang, Hannah Kim, Estevam Hruschka RECAP is a benchmark designed to evaluate and advance agentic planning given a user-agent conversation. RECAP focuses on intent rewriting as an integral part towards understanding user goals and task fulfillment. The dataset comprises user-agent conversations across varied conversation lengths, topics and intent-related challenges.… See the full description on the dataset page: https://huggingface.co/datasets/megagonlabs/recap.textothern<1K0 likes35 downloads10mo agoHugging Face27Ker102 /n8n-mega-workflows 🚀 n8n Mega Workflows - The Largest n8n Workflow Dataset The world's largest open-source n8n workflow dataset for training AI workflow generators 🌟 Highlights 131,648 high-quality n8n workflows with valid connection skeletons 26 semantic categories for balanced coverage Instruction-tuning format ready for fine-tuning LLMs 15+ million lines of workflow JSON Perfect for: RAG pipelines, fine-tuning, workflow generation models 📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/Ker102/n8n-mega-workflows.texttext-generation100K<n<1M0 likes34 downloads9mo agoHugging Face28Khurram123 /urdu-poetry-mega-corpus 📜 Urdu Poetry Mega Corpus This dataset is a comprehensive collection of classical and modern Urdu poetry, meticulously curated for training Large Language Models (LLMs) like Qwen, Llama, and Mistral to generate authentic Urdu Ghazals and Nazms. 🌟 Dataset Overview The Urdu Poetry Mega Corpus contains over 42,639 high-quality Urdu couplets (ash'aar). It is designed to capture the structural nuances, rhythmic patterns (Beher), and stylistic essence of renowned poets such… See the full description on the dataset page: https://huggingface.co/datasets/Khurram123/urdu-poetry-mega-corpus.text10K<n<100K0 likes34 downloads7mo agoHugging Face29megagonlabs /magesql-spider-derived MageSQL — Spider-derived data and model These files are derived from / adapted from the Spider dataset (Yu et al., 2018), which is distributed under CC BY-SA 4.0. Modifications by Megagon Labs, Inc.: merged the Spider train splits (train_spider_and_others.json), extracted database schema text (db_id2schema_text.json), mapped questions to gold SQL (question2sql.json), generated question embeddings (question_embeddings.pt), and trained the database-routing classifier… See the full description on the dataset page: https://huggingface.co/datasets/megagonlabs/magesql-spider-derived.text1K<n<10K0 likes34 downloads4mo agoHugging Face30Alex01837178373 /mega-cot-ru-dataset-ShareGPT Mega CoT Russian Dataset (ShareGPT Format) Этот датасет объединяет три русскоязычных источника данных с цепочками рассуждений (Chain of Thought / CoT), приведенными к единому формату ShareGPT с явным выделением мыслей модели в тегах <think>...</think>. Общая статистика Всего записей: 2,874 диалогов. Формат: ShareGPT (conversations массив с ролями system, human, gpt). Разметка рассуждений: Все рассуждения обернуты в теги <think> ... </think> в начале сообщений от… See the full description on the dataset page: https://huggingface.co/datasets/Alex01837178373/mega-cot-ru-dataset-ShareGPT.text1K<n<10K0 likes29 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.