datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mt_pubmedmtp-selfdata-qwen3-8b-finewikimtp-selfdata-qwen3-8b-finewiki-40kmtp-selfdata-llama3.2-3b-finewikiredhat-docs_dataset
🖥️ Red Hat Technical Documentation Dataset
📌 Overview
This dataset contains 55,741 structured technical documentation entries sourced from Red Hat, covering:✅ System Administration Guides – User management, permissions, kernel tuning✅ Networking & Security – Firewall rules, SELinux, VPN setup✅ Virtualization & Containers – KVM, Podman, OpenShift, Kubernetes✅ Enterprise Software Documentation – RHEL, Ansible, Satellite, OpenStack
📊 Dataset Details
This… See the full description on the dataset page: https://huggingface.co/datasets/mtpti5iD/redhat-docs_dataset.mtp-selfdata-qwen3.6-35b-a3b-finewikirecap_dish_20260611This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so101_follower",
"total_episodes": 71,
"total_frames": 39618,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:71"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mt-prox/recap_dish_20260611.recap_intervention0625This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "bi_so101_follower",
"total_episodes": 42,
"total_frames": 64742,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 500,
"fps": 30,
"splits": {
"train": "0:42"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/mt-prox/recap_intervention0625.MT-prefnexus-sft-mix-v6-phase2-mtprce
NEXUS SFT Mix v6 Phase 2 — MTP-RCE Option B
149,799 samples for NEXUS phase 2 continuation training. Includes MTP-RCE mode annotations on 16% of samples.
Composition
105,000 rehearsal v5 : stratified subset of mix v5 (50 sources, all preserved for anti-forgetting). Preserves identity (4,730 nexus_identity_v2) + style + bilingue 60% FR / 40% EN.
20,799 HF reasoning : math/reasoning filtered ≤4000 chars (NEXUS seq_len 4096 safe).
12K… See the full description on the dataset page: https://huggingface.co/datasets/jescy525/nexus-sft-mix-v6-phase2-mtprce.tool-calling-conversations-mtp1p1c0
Tool Calling Conversations
An Arena-style dataset of anonymized, multi-turn conversations focused on real-world
tool use. It is intended for research, evaluation, and training of models that decide
when and how to call tools.
The conversations include:
Tool selection and no-tool decisions
Structured tool arguments
Sequential and parallel tool calls
Tool results and error recovery
Multi-step agent workflows
Final responses after tool execution
Data is organized into… See the full description on the dataset page: https://huggingface.co/datasets/dakr-pandas/tool-calling-conversations-mtp1p1c0.datasets_for_magnetic_MTP_NatSR2024_verification
Cite this dataset Kotykhov, A. S., Gubaev, K., Hodapp, M., Tantardini, C., Shapeev, A. V., and Novikov, I. S. datasets for magnetic MTP NatSR2024 verification. ColabFit, 2024. https://doi.org/10.60732/acd42be9
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_wu6xd9i8cf7i_0
Visit the ColabFit Exchange to search additional datasets by… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/datasets_for_magnetic_MTP_NatSR2024_verification.datasets_for_magnetic_MTP_NatSR2024_training
Cite this dataset Kotykhov, A. S., Gubaev, K., Hodapp, M., Tantardini, C., Shapeev, A. V., and Novikov, I. S. datasets for magnetic MTP NatSR2024 training. ColabFit, 2024. https://doi.org/10.60732/9d635e27
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_mf8sn11cn6wa_0
Visit the ColabFit Exchange to search additional datasets by author… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/datasets_for_magnetic_MTP_NatSR2024_training.token_classification_datasetZn_MTP_CMS2023
Cite this dataset Mei, H., Cheng, L., Chen, L., Wang, F., Li, J., and Kong, L. Zn MTP CMS2023. ColabFit, 2024. https://doi.org/10.60732/54902e18
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_58y020ce6b6j_0
Visit the ColabFit Exchange to search additional datasets by author, description, element content and more.… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/Zn_MTP_CMS2023.MTPu_2023
Cite this dataset Zongo, K., Sun, H., Ouellet-Plamondon, C., and Beland, L. K. MTPu 2023. ColabFit, 2024. https://doi.org/10.60732/41115bd2
This dataset has been curated and formatted for the ColabFit Exchange
This dataset is also available on the ColabFit Exchange:
https://materials.colabfit.org/id/DS_326i4urabisb_0
Visit the ColabFit Exchange to search additional datasets by author, description, element content and more.… See the full description on the dataset page: https://huggingface.co/datasets/colabfit/MTPu_2023.MT-pref-humanmt-pref-latin-to-englishtest-gemma4-mtp
Generated Responses Dataset (Gemma 4 + MTP)
This dataset contains generated responses for prompts from davanstrien/haiku_dpo,
produced with Google Gemma 4 and vLLM's Multi-Token Prediction support.
Generation Details
Source Dataset: davanstrien/haiku_dpo
Source Split: train
Input Column: question (plain text prompts)
Model: google/gemma-4-26B-A4B-it
Rows Processed: 5
Batches: 3 (chunk size: 2)
Generation Date: 2026-05-06T13:33:39.595844
Script: gemma4-mtp.py… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/test-gemma4-mtp.test-gemma4-mtp-uv
Generated Responses Dataset (Gemma 4 + MTP)
This dataset contains generated responses for prompts from davanstrien/haiku_dpo,
produced with Google Gemma 4 and vLLM's Multi-Token Prediction support.
Generation Details
Source Dataset: davanstrien/haiku_dpo
Source Split: train
Input Column: question (plain text prompts)
Model: google/gemma-4-26B-A4B-it
Rows Processed: 5
Batches: 3 (chunk size: 2)
Generation Date: 2026-05-06T13:42:18.500799
Script: gemma4-mtp.py… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/test-gemma4-mtp-uv.test-gemma4-mtp-ON
Generated Responses Dataset (Gemma 4 + MTP)
This dataset contains generated responses for prompts from davanstrien/haiku_dpo,
produced with Google Gemma 4 and vLLM's Multi-Token Prediction support.
Generation Details
Source Dataset: davanstrien/haiku_dpo
Source Split: train
Input Column: question (plain text prompts)
Model: google/gemma-4-26B-A4B-it
Rows Processed: 30
Batches: 3 (chunk size: 10)
Generation Date: 2026-05-06T13:52:41.595582
Script: gemma4-mtp.py… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/test-gemma4-mtp-ON.test-gemma4-mtp-OFF
Generated Responses Dataset (Gemma 4 + MTP)
This dataset contains generated responses for prompts from davanstrien/haiku_dpo,
produced with Google Gemma 4 and vLLM's Multi-Token Prediction support.
Generation Details
Source Dataset: davanstrien/haiku_dpo
Source Split: train
Input Column: question (plain text prompts)
Model: google/gemma-4-26B-A4B-it
Rows Processed: 30
Batches: 3 (chunk size: 10)
Generation Date: 2026-05-06T13:52:42.477251
Script: gemma4-mtp.py… See the full description on the dataset page: https://huggingface.co/datasets/davanstrien/test-gemma4-mtp-OFF.
