manta
Datasets
All datasets matching “manta”mantaMANTA-1M
Abstract
We introduce MANTA, an automated pipeline that generates high-quality large-scale instruction fine-tuning datasets from massive web corpora while preserving their diversity and scalability. By extracting structured syllabi from web documents and leveraging high-performance LLMs, our approach enables highly effective query-response generation with minimal human intervention. Extensive experiments on 8B-scale LLMs demonstrate that fine-tuning on the MANTA-1M dataset… See the full description on the dataset page: https://huggingface.co/datasets/LGAI-EXAONE/MANTA-1M.so101_testThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": "so101",
"total_episodes": 10,
"total_frames": 5960,
"total_tasks": 1,
"total_videos": 10,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 30,
"splits": {
"train": "0:10"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/ManTang034/so101_test.mantamanta-1m-seqlen-513-1024
MANTA-1M: Calibration Subset (Seq Len 513-1024)
Overview
This dataset is a length-specific subset of the LGAI-EXAONE/MANTA-1M dataset, curated specifically for Post-Training Quantization (PTQ) Calibration.
Following the insights from the MaCa (Matryoshka Calibration) paper, this dataset provides length-specific calibration samples to ensure that the quantization process accounts for the variable weight importance across different input scales. By focusing on the… See the full description on the dataset page: https://huggingface.co/datasets/haesol-shin/manta-1m-seqlen-513-1024.MANTA_1M_ENG_KO_70_30
Abstract
We introduce MANTA, an automated pipeline that generates high-quality large-scale instruction fine-tuning datasets from massive web corpora while preserving their diversity and scalability. By extracting structured syllabi from web documents and leveraging high-performance LLMs, our approach enables highly effective query-response generation with minimal human intervention. Extensive experiments on 8B-scale LLMs demonstrate that fine-tuning on the MANTA-1M dataset… See the full description on the dataset page: https://huggingface.co/datasets/JustcallmeJo/MANTA_1M_ENG_KO_70_30.
