datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GeoMeld
🌍 GeoMeld Multi-Modal Earth Observation Dataset (WebDataset)
GeoMeld is a large-scale multi-modal remote sensing dataset introduced in our CVPRW 2026 paper on semantically grounded foundation modeling.
GeoMeld contains approximately 2.5 million spatially aligned samples spanning heterogeneous sensing modalities and spatial resolutions, paired with semantically grounded captions generated through an agentic pipeline.
The dataset is designed to support multimodal representation… See the full description on the dataset page: https://huggingface.co/datasets/vimageiitb/GeoMeld.ViMD_ReigonGroup
Multi-Dialect Vietnamese: Task, Dataset, Baseline Models and Challenges (Main EMNLP 2024)
Introduction
This document presents the accompanying dataset for the paper titled "Multi-Dialect Vietnamese: Task, Dataset, Baseline Models, and Challenges". The dataset, referred to as the Vietnamese Multi-Dialect (ViMD) dataset, is a comprehensive resource designed to capture the linguistic diversity represented by 63 provincial dialects spoken across Vietnam. The paper is… See the full description on the dataset page: https://huggingface.co/datasets/notlee203/ViMD_ReigonGroup.ViMUL-Bench
ViMUL-Bench: A Culturally-diverse Multilingual Multimodal Video Benchmark
Overview
The evaluation toolkit to be used is lmms-eval. This toolkit facilitates the evaluation of models across multiple tasks and languages.
Key Features
🌍 14 Languages: English, Chinese, Spanish, French, German, Hindi, Arabic, Russian, Bengali, Urdu, Sinhala, Tamil, Swedish, Japanese🎭 15 Categories: Including 8 culturally diverse categories (lifestyles, festivals, foods… See the full description on the dataset page: https://huggingface.co/datasets/MBZUAI/ViMUL-Bench.stillwarm-bench-results
stillwarm-bench results dataset
Every measured row behind the stillwarm project: does saving/restoring a
llama-server conversation's KV cache to disk actually work, when does it beat
recomputing, and what silently breaks it?
Hardware/build (frozen): MacBook Pro, Apple M3 Max, 36 GB unified memory,
macOS 26.5.1; llama.cpp release b9871 (ef2d770…), Release build, Metal;
models pinned by SHA-256 (Llama-3.1-8B-Instruct Q4_K_M, Qwen2.5-7B-Instruct
Q4_K_M, Gemma-3-4B-it Q4_K_M — the… See the full description on the dataset page: https://huggingface.co/datasets/vimalnakrani/stillwarm-bench-results.vimedaqa-nli-g1vi-mmarcostillwarm-kv-cache-artifact
A downloadable KV-cache save file — with the honest math
One llama-server slot save: the first 8,192 Llama-tokens of Frankenstein
(public domain), prefilled by Qwen2.5-7B-Instruct Q4_K_M (Apache-2.0 model —
chosen over Llama specifically for artifact licensing) and saved with a
stillwarm sidecar.
This file is USELESS unless your setup matches the sidecar exactly:
field
value
llama.cpp build
b9871 (ef2d770117db45b05aa7ecd1b0acca36370c5470) — advisory: ±5 weeks measured… See the full description on the dataset page: https://huggingface.co/datasets/vimalnakrani/stillwarm-kv-cache-artifact.trace-forge-kimi-k3-dry-v0
trace-forge-kimi-k3-dry-v0
Reasoning traces from kimi-k3 (Moonshot native API) over a
16-prompt self-authored bank, 2 samples per prompt,
generated on 2026-07-25. All numbers in this card are measured.
Author and maintainer: Vimal Nakrani (vimalnakrani), sole author and maintainer.
Configs
The raw config has all 32 records: prompt, final answer in
content, the model's reasoning in its own reasoning field,
finish_reason, verification status, token usage, seed, and… See the full description on the dataset page: https://huggingface.co/datasets/vimalnakrani/trace-forge-kimi-k3-dry-v0.sometimesanotion__Qwen2.5-14B-Vimarckoso-v3-model_stock-details
Dataset Card for Evaluation run of sometimesanotion/Qwen2.5-14B-Vimarckoso-v3-model_stock
Dataset automatically created during the evaluation run of model sometimesanotion/Qwen2.5-14B-Vimarckoso-v3-model_stock
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwen2.5-14B-Vimarckoso-v3-model_stock-details.context_instruct_vimhopfcVMTEB-ViMedNLI_defaultvi_msa_sar
VN-SarMSA-vi
Vietnamese aspect-level sentiment (7-point, −3…+3) and sarcasm (binary) annotations
for hotel/restaurant reviews.
17,236 (sentence, aspect) pairs; splits 13,758 / 1,739 / 1,739 (train/validation/test)
Splits are leakage-free by construction: reviews sharing any normalized sentence are
grouped before splitting, so no sentence or review crosses splits
source column marks synthetic augmentation ('augmented'); evaluate sentiment on
source != 'augmented' for a… See the full description on the dataset page: https://huggingface.co/datasets/phamluan/vi_msa_sar.details_sometimesanotion__Qwen2.5-14B-Vimarckoso-v3
Dataset Card for Evaluation run of sometimesanotion/Qwen2.5-14B-Vimarckoso-v3
Dataset automatically created during the evaluation run of model sometimesanotion/Qwen2.5-14B-Vimarckoso-v3.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_sometimesanotion__Qwen2.5-14B-Vimarckoso-v3.sometimesanotion__Qwen2.5-14B-Vimarckoso-details
Dataset Card for Evaluation run of sometimesanotion/Qwen2.5-14B-Vimarckoso
Dataset automatically created during the evaluation run of model sometimesanotion/Qwen2.5-14B-Vimarckoso
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwen2.5-14B-Vimarckoso-details.sometimesanotion__Qwen2.5-14B-Vimarckoso-v3-Prose01-details
Dataset Card for Evaluation run of sometimesanotion/Qwen2.5-14B-Vimarckoso-v3-Prose01
Dataset automatically created during the evaluation run of model sometimesanotion/Qwen2.5-14B-Vimarckoso-v3-Prose01
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwen2.5-14B-Vimarckoso-v3-Prose01-details.sometimesanotion__Qwen2.5-14B-Vimarckoso-v2-details
Dataset Card for Evaluation run of sometimesanotion/Qwen2.5-14B-Vimarckoso-v2
Dataset automatically created during the evaluation run of model sometimesanotion/Qwen2.5-14B-Vimarckoso-v2
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwen2.5-14B-Vimarckoso-v2-details.sometimesanotion__Qwen2.5-14B-Vimarckoso-v3-details
Dataset Card for Evaluation run of sometimesanotion/Qwen2.5-14B-Vimarckoso-v3
Dataset automatically created during the evaluation run of model sometimesanotion/Qwen2.5-14B-Vimarckoso-v3
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwen2.5-14B-Vimarckoso-v3-details.gold_evidence_instruct_vimhopfcDatasetsPADvimedaqa-rft-poolsometimesanotion__Qwen2.5-14B-Vimarckoso-v3-IF-Variant-details
Dataset Card for Evaluation run of sometimesanotion/Qwen2.5-14B-Vimarckoso-v3-IF-Variant
Dataset automatically created during the evaluation run of model sometimesanotion/Qwen2.5-14B-Vimarckoso-v3-IF-Variant
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train"… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/sometimesanotion__Qwen2.5-14B-Vimarckoso-v3-IF-Variant-details.vimedaqa-status-runsvimedaqa-bon-candidates
