datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
UGround-V1-8k
UGround-WebHybrid-8K
This dataset is a curated 8K-sample subset from the original UGround-V1-Data (Web-Hybrid), as mentioned in our paper. It serves as part of the training corpus for GUI grounding tasks, focusing on diverse web interface screenshots across resolutions and aspect ratios.
Paper and Code
Paper: ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding
Code: https://github.com/zonghanHZH/ZonUI-3B
Dataset Details
Source:… See the full description on the dataset page: https://huggingface.co/datasets/zonghanHZH/UGround-V1-8k.UltraData-SFT-2605-no-think-8k-32k
UltraData-SFT-2605 · no_think · 8k–32k
A length-filtered subset of the no_think split of
openbmb/UltraData-SFT-2605,
containing conversations whose token length falls in the 8k–32k range.
This is the medium-length tier intended for standard long-context SFT.
Two companion tiers were produced from the same source:
Dataset
Length range
Records
this repo — fxmeng/UltraData-SFT-2605-no-think-8k-32k
8k–32k tokens
623,421
fxmeng/UltraData-SFT-2605-no-think-32k-200k… See the full description on the dataset page: https://huggingface.co/datasets/fxmeng/UltraData-SFT-2605-no-think-8k-32k.AMEX-8k
AMEX-8K
This dataset is a curated 8K-sample subset from the original AMEX dataset, as mentioned in our paper. It serves as part of the training corpus for GUI grounding tasks, specifically capturing mobile app interfaces across diverse platforms and screen densities.
Dataset Details
Source: Sampled from AMEX
Domain: Mobile GUI screenshots
Diversity: Includes a variety of app types and device form factors
Use case: GUI grounding pretraining, especially for mobile… See the full description on the dataset page: https://huggingface.co/datasets/zonghanHZH/AMEX-8k.ShowUI-web-8kzh-meme-sft-8k
zh-meme-sft-8k
📖 简介 | Introduction
zh-meme-sft-8k 是一个高质量的中文互联网梗文化指令微调数据集。该数据集基于抖音、小红书、B站等平台的真实评论互动构建,经过多轮清洗、增强和格式化处理,专门用于训练能够理解和使用网络热梗、具备幽默感的对话模型。
🎯 这个数据集是 Meme-Qwen-7B-Instruct 模型的训练数据,如果你想看微调后的效果,可以直接体验模型!
这个数据集的特点是:
🎯 真实来源:基于真实社交平台的用户互动,保留原本网络表达
🔄 对话结构:包含帖子-评论、评论-回复的完整对话链
🧹 精细清洗:经过多轮规则清洗和LLM增强,去除噪声的同时保留热梗
💬 ChatML格式:标准化为ChatML格式,开箱即用
📊 数据统计 | Data Statistics
数据集
样本数量
占比
训练集
7,377
85%
验证集
868
10%
测试集
435
5%
总计
8… See the full description on the dataset page: https://huggingface.co/datasets/GaryYang123/zh-meme-sft-8k.ShowUI-web-8k
ShowUI-web-8K
This dataset is a curated 8K-sample subset from the original ShowUI-web dataset, as mentioned in our paper. It contributes to the training of GUI grounding models, with a focus on realistic web user interfaces collected from diverse websites.
Dataset Details
Source: Sampled from ShowUI-web
Domain: Web GUI screenshots
Diversity: Covers a wide variety of website layouts and components
Use case: GUI grounding pretraining for web environments… See the full description on the dataset page: https://huggingface.co/datasets/zonghanHZH/ShowUI-web-8k.JiRack-Magpie-Pro-MT-300K_8k-DatasetUGround-V1-8kalimeeting-eval-8k
AliMeeting Eval — 8 kHz CH0 clips
Unknown-(N) eval clips from AliMeeting Eval (M2MeT / OpenSLR 119), far-field channel 0, resampled to 8 kHz.
Manifest
Clips
manifests/eval_all.jsonl
1280 (full local eval)
manifests/eval_n100.jsonl
100 stratified subset
manifests/eval_n200.jsonl
200 stratified subset
manifests/eval_*mix.jsonl
by (N=1\ldots4)
Headset s{k}.wav stems (near, TextGrid-gated) are present for the n200 subset (eval_n200_headset.jsonl). Other clips… See the full description on the dataset page: https://huggingface.co/datasets/playwithmino/alimeeting-eval-8k.MedThoughts-8K
MedThoughts-8K
English|中文
This dataset is distilled from the full-scale DeepSeek-R1 (671B) in the medical domain. For more detailed information, please refer to our GitHub project MedR1.
1. Original Dataset
The data in this dataset is sourced from the US/train partition of MedQA (5 options).
2. Dataset Format
The keys in the dataset are explained as follows:
"question_id": The unique identifier for the question,
"question": The question itself,
"options": The… See the full description on the dataset page: https://huggingface.co/datasets/hw-hwei/MedThoughts-8K.benchmark_8k
Benchmark 8K Dataset
A curated dataset of 1,000 high-quality prompts designed for benchmarking Large Language Model (LLM) performance across various metrics including latency, throughput, and response quality. This dataset features longer, more complex prompts ideal for testing models' capabilities with extended context and detailed analysis tasks.
Dataset Overview
Size: 100 prompts
Format: JSONL (JSON Lines)
Average Token Length: Variable (extended context; computed… See the full description on the dataset page: https://huggingface.co/datasets/raffel36/benchmark_8k.tigerbot-gsm-8k-enTigerbot 基于gsm8k数据集加工而来
GSM8K(Grade School Math 8K)是一个包含 8.5K 高质量语言多样化小学数学单词问题的数据集。创建数据集是为了支持对需要多步推理的基本数学问题的问答任务。
原始来源:https://huggingface.co/datasets/gsm8k
Usage
import datasets
ds_sft = datasets.load_dataset('TigerResearch/tigerbot-gsm-8k-en')
arabic-rag-chat-8k-eval
arabic-rag-chat-8k-eval
Per-row evaluation artifacts for the 8,192-token Arabic multi-turn RAG models:
the test split, every model's raw replies, every judge verdict, and the rendered
report for each. Thirteen judged models, all scored on the same 1,651 prompts
by the same judge at temperature 0.0, so the comparison below is like-for-like
and can be recomputed offline without a GPU or a judge server.
This is the measurement half of
oddadmix/100M-8192-Nawah-dsv4;
the training… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-chat-8k-eval.deepseek-v4-distill-8k
🐳 DeepSeek-V4-Distill-8100x
Dataset Summary
DeepSeek-V4-Distill-8100x is a supervised fine-tuning dataset for reasoning-oriented distillation. The question prompts come from Jackrong/GLM-5.1-Reasoning-1M-Cleaned, and the answers were generated by the teacher model DeepSeek-V4-Flash.
After the cleaning process, the released train split contains 7,716 high-quality JSONL examples.
[!NOTE]
The answer pool was cleaned to remove real-time questions… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/deepseek-v4-distill-8k.grimjim__Magnolia-v2-Gemma2-8k-9B-details
Dataset Card for Evaluation run of grimjim/Magnolia-v2-Gemma2-8k-9B
Dataset automatically created during the evaluation run of model grimjim/Magnolia-v2-Gemma2-8k-9B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/grimjim__Magnolia-v2-Gemma2-8k-9B-details.1k-creative-writing-8kt-fineweb-edu-sampleTotal tokens in matching entries: 5_575_157
Average tokens per entry: 5575.16
20k-finewebedu-samples-8ktquantum-circuits-8k
Quantum Circuits 8K Dataset
A synthetic dataset of 8,129 quantum circuit examples for training language models to generate OpenQASM 2.0 code from natural language descriptions.
Quick Stats
Total Samples: 8,129 (description → QASM pairs)
Unique Circuits: 739 base circuits
Categories: 92 distinct quantum circuit types
Qubit Range: 1-9 qubits
Format: OpenQASM 2.0
Augmentation: 11x per circuit (original + 10 paraphrases)
Quality: 100% QASM syntax valid, 0% duplicates… See the full description on the dataset page: https://huggingface.co/datasets/merileijona/quantum-circuits-8k.JiRack-No_Robots_8k-Datasetptbr-deita-8k
PTBR Deita 8k
Portuguese translation of the Deita 8k dataset.
100k-finewebedu-samples-8kt8k_filterd_trainOpenSciDER-SFT-8K
OpenSciDER
This is the model repo for OpenSciDER-SFT-8K, a trajectory dataset curated from SciDER.
This dataset is collected from the inference trajectories of the Qwen3.6-27B and OpenSciDER-27B model on DataSciBench, DS-1000, DS-Bench, and ScienceAgentBench. It also contains the benchmark evaluation trajectories on AI-Idea-Bench, AIRS-Bench, AstroVisBench, DiscoveryBench, MLE-Bench, and SciCode.
There are 3 configurations in this dataset:
openscider: the trajectories of… See the full description on the dataset page: https://huggingface.co/datasets/leonardklin/OpenSciDER-SFT-8K.Math-Japanese-8k算数の文章問題のデータセット
問題文をKimi K2.6 K2.7-codeで作成
回答をGemma4 31Bで作成
AIVision360-8k
Dataset Card for AIVision360-8k
Dataset Description
AIVision360 is the pioneering domain-specific dataset tailor-made for media and journalism, designed expressly for the instruction fine-tuning of Large Language Models (LLMs).The AIVision360-8k dataset is a curated collection sourced from "ainewshub.ie", a platform dedicated to Artificial Intelligence news from quality-controlled publishers. It is designed to provide a comprehensive representation of AI-related… See the full description on the dataset page: https://huggingface.co/datasets/ceadar-ie/AIVision360-8k.CAI-Rev-1-Opus-22K-8K-Subset-Non-Convertedruler2_8kmicrosoft__Phi-3-small-8k-instruct-details
Dataset Card for Evaluation run of microsoft/Phi-3-small-8k-instruct
Dataset automatically created during the evaluation run of model microsoft/Phi-3-small-8k-instruct
The dataset is composed of 34 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/microsoft__Phi-3-small-8k-instruct-details.nyxmed-icd-cpt-8kflashback-samhalle-8k
flashback-samhalle-8k
Swedish Flashback "samhalle" subforum, parsed with --max-chars 8000 into ShareGPT-format chat conversations. Voice-cloning / reactive-agent dataset. Each record is one whole thread (multi-turn conversations); row-level train/validation split has no post-leakage.
