datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
kernelbench-mega-traces
KernelBench-Mega agent traces
Coding agents writing full GPU megakernels across Blackwell / H100 / B200, scored as speedup over reference; contamination-audited (23 verified cells).
Each .jsonl file is one agent run in Claude-Code session format, viewable with the agent trace viewer. Filename = run id; manifest.csv maps each run to model / harness / problem / GPU / score.
23 agent traces · live leaderboard: https://kernelbench.com/mega
Secrets redacted. Full reasoning for… See the full description on the dataset page: https://huggingface.co/datasets/Infatoshi/kernelbench-mega-traces.quran-asr-mega-corpusmegatron-prof-data-v5megawika-report-generation
Dataset Card for MegaWika for Report Generation
Dataset Summary
MegaWika is a multi- and crosslingual text dataset containing 30 million Wikipedia passages with their scraped and cleaned web citations. The passages span
50 Wikipedias in 50 languages, and the articles in which the passages were originally embedded are included for convenience. Where a Wikipedia passage is in a
non-English language, an automated English translation is provided.
This dataset provides the… See the full description on the dataset page: https://huggingface.co/datasets/hltcoe/megawika-report-generation.cypherbench
CypherBench
CypherBench is a benchmark designed to evaluate text-to-Cypher translation for large language models (LLMs). It includes:
11 large-scale Neo4j property graphs transformed from Wikidata with 7.8 million entities.
Over 10,000 (question, Cypher) pairs for training/evaluating text-to-Cypher translation.
Paper: https://arxiv.org/pdf/2412.18702
Repository & Demo: https://github.com/megagonlabs/cypherbench
Contact: yanlin@megagon.ai
Sample Task
{
"qid":… See the full description on the dataset page: https://huggingface.co/datasets/megagonlabs/cypherbench.mega-acceptability-v2megamatt-translated-ITmegatron-prof-data-v6Rustins_Super_Mega_Awesome_VEDU_Model
Rustin's Super Mega Awesome VEDU Model
A reproducible, heavily-documented pipeline that maps Ventenata dubia ("VEDU", an invasive
winter-annual grass) across Montana from satellite + environmental data.
Science reference: docs/VEDU_48_predictors_detailed.md
Data decisions & gotchas: docs/CONTRADICTIONS.md
Parity with the Earth Engine build: docs/GEE_PARITY.md
Continue-the-build guide: docs/HANDOFF.md
Label inventory: docs/DATA_SOURCES.md
What it produces
57… See the full description on the dataset page: https://huggingface.co/datasets/UniversityOfMontanaSAL/Rustins_Super_Mega_Awesome_VEDU_Model.megatron-prof-data-v14megatron-prof-data-v7dolma3_dolmino_megatron_tokenize
Dolma 3 / Dolmino Megatron-LM indexed dataset
This repository contains immutable Megatron-LM indexed datasets (.bin and
.idx) produced from pinned Dolma 3 and Dolmino releases. It intentionally
contains no training checkpoints, experiment outputs, logs, or dataset caches.
The indexed payloads were derived from these pinned public datasets:
allenai/dolma3_mix-150B-1025@afa92bfb22366821c5e6cd427cdd036b34b713ef… See the full description on the dataset page: https://huggingface.co/datasets/Mr-Philo/dolma3_dolmino_megatron_tokenize.megatron-prof-data-v3megatron-prof-data-v4lm-eval-results-Eurdem-megatron_2.1_MoE_2x7B-private
Dataset Card for Evaluation run of Eurdem/megatron_2.1_MoE_2x7B
Dataset automatically created during the evaluation run of model Eurdem/megatron_2.1_MoE_2x7B
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-Eurdem-megatron_2.1_MoE_2x7B-private.megatron-prof-data-v12ZamAI-Pashto-Mega-Dataset
ZamAI Pashto Mega Dataset
Languages: psLicense: apache-2.0Task categories: text-generation, summarization, question-answeringSize categories: 1M<n<10M
Summary
This dataset is part of the ZamAI Pashto data collection. It is intended for text-generation, summarization, question-answering tasks in Pashto.
How to use
from datasets import load_dataset
dataset = load_dataset("tasal9/ZamAI-Pashto-Mega-Dataset")
print(dataset)
Configs… See the full description on the dataset page: https://huggingface.co/datasets/tasal9/ZamAI-Pashto-Mega-Dataset.aksara-mega-sft
Aksara Mega SFT — 64K+ Dataset Grade S
AksaraLLM Community mempersembahkan 64797 pasangan instruksi Indonesia kualitas tinggi.
Highlight Dataset
Pemahaman 10 Bahasa Daerah (Jawa, Sunda, Minang, dll) via NusaX
Ribuan QA Suku, Agama, dan Budaya Nusantara
SQuAD ID & Dolly 15K Indonesian
Orca Math Word Problems Indonesian
Guanaco & xP3x High Quality Conversations
38 Provinsi Lengkap & Etika Lokal
Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8-details
Dataset Card for Evaluation run of Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8
Dataset automatically created during the evaluation run of model Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8-details.Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8.9-details
Dataset Card for Evaluation run of Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8.9
Dataset automatically created during the evaluation run of model Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8.9
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8.9-details.megazeka-tr-spellfix-pairs
Megazeka · Turkish Spelling Correction Pairs
40,313 (misspelled → correct) Turkish sentence pairs with a full record of every corruption that
was applied, generated from the CC0 Common Voice Turkish Sentence Collector and from project-authored
everyday first/second-person sentences.
The distinguishing feature is the operations field: each pair carries the exact sequence of noise
transformations that produced it, with before/after text at each step. That makes it possible to… See the full description on the dataset page: https://huggingface.co/datasets/pandakingpunc/megazeka-tr-spellfix-pairs.megacode3-min100prithivMLmods__Megatron-Corpus-14B-Exp-details
Dataset Card for Evaluation run of prithivMLmods/Megatron-Corpus-14B-Exp
Dataset automatically created during the evaluation run of model prithivMLmods/Megatron-Corpus-14B-Exp
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/prithivMLmods__Megatron-Corpus-14B-Exp-details.megamath-web-pro-max-splittedMegaCodeTraining
VERSION 3 IS RELEASED DOWNLOAD HERE:
https://huggingface.co/datasets/rombodawg/LosslessMegaCodeTrainingV3_2.2m_Evol
This is a uncensored mega combined dataset using both razent/wizardlm-code-evol-32k and nickrosh/Evol-Instruct-Code-80k-v1
In this version many lines of instructions were removed in part of a uncensoring process.
The Rombo's format.rar file is so you can use the training data in oobagooba text generation webui. Simply unzip it, and use it as a json file.
All links bellow… See the full description on the dataset page: https://huggingface.co/datasets/rombodawg/MegaCodeTraining.recap
RECAP
RECAP: REwriting Conversations for Intent Understanding in Agentic Planning 📄 paper link
Kushan Mitra, Dan Zhang, Hannah Kim, Estevam Hruschka
RECAP is a benchmark designed to evaluate and advance agentic planning given a user-agent conversation. RECAP focuses on intent rewriting as an integral part towards understanding user goals and task fulfillment. The dataset comprises user-agent conversations across varied conversation lengths, topics and intent-related challenges.… See the full description on the dataset page: https://huggingface.co/datasets/megagonlabs/recap.n8n-mega-workflows
🚀 n8n Mega Workflows - The Largest n8n Workflow Dataset
The world's largest open-source n8n workflow dataset for training AI workflow generators
🌟 Highlights
131,648 high-quality n8n workflows with valid connection skeletons
26 semantic categories for balanced coverage
Instruction-tuning format ready for fine-tuning LLMs
15+ million lines of workflow JSON
Perfect for: RAG pipelines, fine-tuning, workflow generation models
📊 Dataset Statistics… See the full description on the dataset page: https://huggingface.co/datasets/Ker102/n8n-mega-workflows.urdu-poetry-mega-corpus
📜 Urdu Poetry Mega Corpus
This dataset is a comprehensive collection of classical and modern Urdu poetry, meticulously curated for training Large Language Models (LLMs) like Qwen, Llama, and Mistral to generate authentic Urdu Ghazals and Nazms.
🌟 Dataset Overview
The Urdu Poetry Mega Corpus contains over 42,639 high-quality Urdu couplets (ash'aar). It is designed to capture the structural nuances, rhythmic patterns (Beher), and stylistic essence of renowned poets such… See the full description on the dataset page: https://huggingface.co/datasets/Khurram123/urdu-poetry-mega-corpus.magesql-spider-derived
MageSQL — Spider-derived data and model
These files are derived from / adapted from the
Spider dataset (Yu et al., 2018), which is
distributed under
CC BY-SA 4.0.
Modifications by Megagon Labs, Inc.: merged the Spider train splits
(train_spider_and_others.json), extracted database schema text
(db_id2schema_text.json), mapped questions to gold SQL (question2sql.json),
generated question embeddings (question_embeddings.pt), and trained the
database-routing classifier… See the full description on the dataset page: https://huggingface.co/datasets/megagonlabs/magesql-spider-derived.mega-cot-ru-dataset-ShareGPT
Mega CoT Russian Dataset (ShareGPT Format)
Этот датасет объединяет три русскоязычных источника данных с цепочками рассуждений (Chain of Thought / CoT), приведенными к единому формату ShareGPT с явным выделением мыслей модели в тегах <think>...</think>.
Общая статистика
Всего записей: 2,874 диалогов.
Формат: ShareGPT (conversations массив с ролями system, human, gpt).
Разметка рассуждений: Все рассуждения обернуты в теги <think> ... </think> в начале сообщений от… See the full description on the dataset page: https://huggingface.co/datasets/Alex01837178373/mega-cot-ru-dataset-ShareGPT.
