datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
SAINetset_v8.0
SAINetset - Wildfire Smoke Detection Dataset
Dataset of real-world images captured by SAI (Sistema de Alerta de Incendios / Fire Alert System) surveillance nodes for wildfire smoke detection in Cordoba, Argentina.
Current version: v8.0 (January 2026)
About SAI
The SAI (Fire Alert System) is an open-source early wildfire detection platform developed by AlterMundi, a civil association in Argentina. The system uses distributed camera nodes with YOLO-based AI (powered by… See the full description on the dataset page: https://huggingface.co/datasets/SAINetset/SAINetset_v8.0.genomes-v4-genome_set-animals-intervals-v8_256_128vision-opd-vqa14k-fullimage-curriculum-v8
Vision-OPD VQA14K Full-Image Curriculum v8
Private single-image visual-question-answering dataset.
Split
Rows
Train
14,000
Diagnostic validation
609
The repository contains 14,609 content-addressed media files (4,580,273,467 bytes). Paths in both Parquet files are relative to the repository root and follow media/<sha256-prefix>/<filename>.
from pathlib import Path
import pyarrow.parquet as pq
from huggingface_hub import snapshot_download
root =… See the full description on the dataset page: https://huggingface.co/datasets/yyy051007/vision-opd-vqa14k-fullimage-curriculum-v8.sequences_only_correct_V8babylon-native-v8-noise-op-wiseai2thor-perspective-qa-800-qa-v8-splitsngld-grape-leaf-vlm-w-img-without-diff-ref-v8promptriever-ours-v8-vanilla-add_qkaggle-native-v8-noise-op-wisenova-v8-dataset
NOVA v8 Pre-Training Dataset
This is the official pre-training dataset used to train the NOVA v8 architecture (a 710M parameter hybrid Sparse Neural Network).
Dataset Structure
Dataset Mix (10 Billion Tokens)
This is NOT a generic web crawl. This dataset is an ultra-dense "university education" designed to make the 710M model punch far above its weight class in reasoning, logic, code, and structural awareness.
Dataset
Weight
Description… See the full description on the dataset page: https://huggingface.co/datasets/ReXeeD/nova-v8-dataset.salabs-virtual-spatial-digitaltwin-v8
🌐 SALabs 10,000,000-Node 3D Virtual Spatial & Digital Twin Avatar Kinematics Dataset (v8.0)
[!IMPORTANT]
💳 Click Here to Purchase Enterprise Commercial License ($2,000 USD) & Instant 8.0GB Master DownloadInstant download of the complete 8.0GB master archive containing 10,000,000 verified 3D spatial nodes, 18-DoF avatar kinematics, B-spline 4D motion tensors, Laplace-Beltrami spectral resonance, and commercial license certificate.
🌟 Executive Summary
The… See the full description on the dataset page: https://huggingface.co/datasets/suitai/salabs-virtual-spatial-digitaltwin-v8.lm-eval-results-bobofrut-ladybird-base-7B-v8-private
Dataset Card for Evaluation run of bobofrut/ladybird-base-7B-v8
Dataset automatically created during the evaluation run of model bobofrut/ladybird-base-7B-v8
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-bobofrut-ladybird-base-7B-v8-private.details_zhengr__MixTAO-7Bx2-MoE-v8.1
Dataset Card for Evaluation run of zhengr/MixTAO-7Bx2-MoE-v8.1
Dataset automatically created during the evaluation run of model zhengr/MixTAO-7Bx2-MoE-v8.1.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_zhengr__MixTAO-7Bx2-MoE-v8.1.regulated_layout_dataset_v8_20260709swe-zero-grounded-v8danish-tool-dialogues-v8
danish-tool-dialogues-v1
Danish multi-turn tool-use conversations with reasoning, translated from the
Glaive subset of
Nanbeige/ToolMind
(Apache-2.0) by scripts/translate_toolmind_da.py.
Complements danish-tool-calls-v1, which is single-turn and synthetic. Here
the conversations run several turns, tool results are fed back, and the
assistant reasons before calling.
split
rows
train
34,168
eval_seen_tools
698
eval_unseen_tools
768
eval_seen_sym
752… See the full description on the dataset page: https://huggingface.co/datasets/jensjepsen/danish-tool-dialogues-v8.promptriever-ours-v8-vanilla-instructiongcd-native-v8-noise-op-wiselm-eval-results-zhengr-MixTAO-7Bx2-MoE-v8.1-private
Dataset Card for Evaluation run of zhengr/MixTAO-7Bx2-MoE-v8.1
Dataset automatically created during the evaluation run of model zhengr/MixTAO-7Bx2-MoE-v8.1
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-zhengr-MixTAO-7Bx2-MoE-v8.1-private.LR_Robotics_real_diagonal_v8india-crop-yield-prediction
🇮🇳 India Crop Yield Prediction Dataset (2000–2026)
A comprehensive dataset for building crop yield prediction models for Indian agriculture, covering 27 years (2000–2026), 31 states/UTs, and 62 crop types.
Dataset Summary
Metric
Value
Total Records
21,750
Year Range
2000 – 2026
States/UTs
31
Crop Types
62
Train Split
17,400 (80%)
Test Split
4,350 (20%)
Missing Values
0
Features
Column
Type
Description
Year… See the full description on the dataset page: https://huggingface.co/datasets/v8shanth/india-crop-yield-prediction.aletheias-phoenix-v8-1-kimi-k3-distillation
Phoenix 8.1 Kimi K3 distillation annotations
This is the reproducibility artifact for SAIN Groningen's Phoenix Wright 8.1
submission to Aletheia's Quest. Phoenix 8.1 attained 0.9661 mean private
per-dataset AUROC. The matching rank-16 Qwen3.5-9B adapter is
Jazhyc/aletheias-phoenix-v8-1-kimi-k3-liars-full-r16-ep2.
The train split has 13,149 rows:
source
rows
Aletheia's Quest public competition training rows
6,573
Liars' Bench harm-pressure choice
800
Liars' Bench… See the full description on the dataset page: https://huggingface.co/datasets/Jazhyc/aletheias-phoenix-v8-1-kimi-k3-distillation.Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8-details
Dataset Card for Evaluation run of Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8
Dataset automatically created during the evaluation run of model Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 3 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8-details.Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8.9-details
Dataset Card for Evaluation run of Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8.9
Dataset automatically created during the evaluation run of model Lunzima/NQLSG-Qwen2.5-14B-MegaFusion-v8.9
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Lunzima__NQLSG-Qwen2.5-14B-MegaFusion-v8.9-details.zeroin-v8-r4-dataset
Zeroin Methodology v8-r4 QA
Training corpus used to fine-tune
KG-ZEROIN/gpt-oss-20b-zeroin-v8-r4.
The corpus captures question/answer pairs derived from the Zeroin
fund-evaluation methodology Korean domain document. It is organized
as a Harmony-ready chat-messages dataset for supervised full
fine-tuning of
openai/gpt-oss-20b.
Released under CC BY-NC 4.0 — free for non-commercial research,
evaluation, and educational use. See LICENSE and NOTICE.
Contents
File… See the full description on the dataset page: https://huggingface.co/datasets/KG-ZEROIN/zeroin-v8-r4-dataset.zhengr__MixTAO-7Bx2-MoE-v8.1-details
Dataset Card for Evaluation run of zhengr/MixTAO-7Bx2-MoE-v8.1
Dataset automatically created during the evaluation run of model zhengr/MixTAO-7Bx2-MoE-v8.1
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/zhengr__MixTAO-7Bx2-MoE-v8.1-details.babylon-native-v8-vocab-noisedatlas_region2_harmful_v8baab-next-v8-rawcontext-conditioned-molecule-transfer-v8-skin-reaction-mixed-continuous-intern
Skin_Reaction context-conditioned molecule transfer V8
V8 preserves the V7 direct panels and appends training-only, post-aggregate continuous assay-evidence transfer pairs. Query values remain hidden from prompts.
Train rows: 46,008
Validation rows: 6,073
Test rows: 18,697
