CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01zonghanHZH /UGround-V1-8k UGround-WebHybrid-8K This dataset is a curated 8K-sample subset from the original UGround-V1-Data (Web-Hybrid), as mentioned in our paper. It serves as part of the training corpus for GUI grounding tasks, focusing on diverse web interface screenshots across resolutions and aspect ratios. Paper and Code Paper: ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding Code: https://github.com/zonghanHZH/ZonUI-3B Dataset Details Source:… See the full description on the dataset page: https://huggingface.co/datasets/zonghanHZH/UGround-V1-8k.imageimage-text-to-text1K<n<10K0 likes3.3k downloads1y agoHugging Face02fxmeng /UltraData-SFT-2605-no-think-8k-32k UltraData-SFT-2605 · no_think · 8k–32k A length-filtered subset of the no_think split of openbmb/UltraData-SFT-2605, containing conversations whose token length falls in the 8k–32k range. This is the medium-length tier intended for standard long-context SFT. Two companion tiers were produced from the same source: Dataset Length range Records this repo — fxmeng/UltraData-SFT-2605-no-think-8k-32k 8k–32k tokens 623,421 fxmeng/UltraData-SFT-2605-no-think-32k-200k… See the full description on the dataset page: https://huggingface.co/datasets/fxmeng/UltraData-SFT-2605-no-think-8k-32k.texttext-generation100K<n<1M0 likes3k downloads2mo agoHugging Face03zonghanHZH /AMEX-8k AMEX-8K This dataset is a curated 8K-sample subset from the original AMEX dataset, as mentioned in our paper. It serves as part of the training corpus for GUI grounding tasks, specifically capturing mobile app interfaces across diverse platforms and screen densities. Dataset Details Source: Sampled from AMEX Domain: Mobile GUI screenshots Diversity: Includes a variety of app types and device form factors Use case: GUI grounding pretraining, especially for mobile… See the full description on the dataset page: https://huggingface.co/datasets/zonghanHZH/AMEX-8k.imageimage-text-to-text1K<n<10K1 likes803 downloads1y agoHugging Face04u8nXq3zW /ShowUI-web-8kimage1K<n<10K0 likes384 downloads1y agoHugging Face05GaryYang123 /zh-meme-sft-8k zh-meme-sft-8k 📖 简介 | Introduction zh-meme-sft-8k 是一个高质量的中文互联网梗文化指令微调数据集。该数据集基于抖音、小红书、B站等平台的真实评论互动构建,经过多轮清洗、增强和格式化处理,专门用于训练能够理解和使用网络热梗、具备幽默感的对话模型。 🎯 这个数据集是 Meme-Qwen-7B-Instruct 模型的训练数据,如果你想看微调后的效果,可以直接体验模型! 这个数据集的特点是: 🎯 真实来源:基于真实社交平台的用户互动,保留原本网络表达 🔄 对话结构:包含帖子-评论、评论-回复的完整对话链 🧹 精细清洗:经过多轮规则清洗和LLM增强,去除噪声的同时保留热梗 💬 ChatML格式:标准化为ChatML格式,开箱即用 📊 数据统计 | Data Statistics 数据集 样本数量 占比 训练集 7,377 85% 验证集 868 10% 测试集 435 5% 总计 8… See the full description on the dataset page: https://huggingface.co/datasets/GaryYang123/zh-meme-sft-8k.texttext-generation1K<n<10K84 likes331 downloads5mo agoHugging Face06zonghanHZH /ShowUI-web-8k ShowUI-web-8K This dataset is a curated 8K-sample subset from the original ShowUI-web dataset, as mentioned in our paper. It contributes to the training of GUI grounding models, with a focus on realistic web user interfaces collected from diverse websites. Dataset Details Source: Sampled from ShowUI-web Domain: Web GUI screenshots Diversity: Covers a wide variety of website layouts and components Use case: GUI grounding pretraining for web environments… See the full description on the dataset page: https://huggingface.co/datasets/zonghanHZH/ShowUI-web-8k.imageimage-text-to-text1K<n<10K0 likes314 downloads1y agoHugging Face07kgrabko /JiRack-Magpie-Pro-MT-300K_8k-Datasettext100K<n<1M0 likes173 downloads5mo agoHugging Face08u8nXq3zW /UGround-V1-8kimage1K<n<10K0 likes135 downloads1y agoHugging Face09playwithmino /alimeeting-eval-8k AliMeeting Eval — 8 kHz CH0 clips Unknown-(N) eval clips from AliMeeting Eval (M2MeT / OpenSLR 119), far-field channel 0, resampled to 8 kHz. Manifest Clips manifests/eval_all.jsonl 1280 (full local eval) manifests/eval_n100.jsonl 100 stratified subset manifests/eval_n200.jsonl 200 stratified subset manifests/eval_*mix.jsonl by (N=1\ldots4) Headset s{k}.wav stems (near, TextGrid-gated) are present for the n200 subset (eval_n200_headset.jsonl). Other clips… See the full description on the dataset page: https://huggingface.co/datasets/playwithmino/alimeeting-eval-8k.audioaudio-to-audio1K<n<10K0 likes127 downloads24d agoHugging Face10hw-hwei /MedThoughts-8K MedThoughts-8K English|中文 This dataset is distilled from the full-scale DeepSeek-R1 (671B) in the medical domain. For more detailed information, please refer to our GitHub project MedR1. 1. Original Dataset The data in this dataset is sourced from the US/train partition of MedQA (5 options). 2. Dataset Format The keys in the dataset are explained as follows: "question_id": The unique identifier for the question, "question": The question itself, "options": The… See the full description on the dataset page: https://huggingface.co/datasets/hw-hwei/MedThoughts-8K.textquestion-answering1K<n<10K4 likes118 downloads2y agoHugging Face11raffel36 /benchmark_8k Benchmark 8K Dataset A curated dataset of 1,000 high-quality prompts designed for benchmarking Large Language Model (LLM) performance across various metrics including latency, throughput, and response quality. This dataset features longer, more complex prompts ideal for testing models' capabilities with extended context and detailed analysis tasks. Dataset Overview Size: 100 prompts Format: JSONL (JSON Lines) Average Token Length: Variable (extended context; computed… See the full description on the dataset page: https://huggingface.co/datasets/raffel36/benchmark_8k.texttext-generationn<1K0 likes109 downloads1y agoHugging Face12TigerResearch /tigerbot-gsm-8k-enTigerbot 基于gsm8k数据集加工而来 GSM8K(Grade School Math 8K)是一个包含 8.5K 高质量语言多样化小学数学单词问题的数据集。创建数据集是为了支持对需要多步推理的基本数学问题的问答任务。 原始来源:https://huggingface.co/datasets/gsm8k Usage import datasets ds_sft = datasets.load_dataset('TigerResearch/tigerbot-gsm-8k-en') text1K<n<10K0 likes73 downloads3y agoHugging Face13oddadmix /arabic-rag-chat-8k-eval arabic-rag-chat-8k-eval Per-row evaluation artifacts for the 8,192-token Arabic multi-turn RAG models: the test split, every model's raw replies, every judge verdict, and the rendered report for each. Thirteen judged models, all scored on the same 1,651 prompts by the same judge at temperature 0.0, so the comparison below is like-for-like and can be recomputed offline without a GPU or a judge server. This is the measurement half of oddadmix/100M-8192-Nawah-dsv4; the training… See the full description on the dataset page: https://huggingface.co/datasets/oddadmix/arabic-rag-chat-8k-eval.tabularquestion-answeringn<1K0 likes72 downloads1mo agoHugging Face14ansulev /deepseek-v4-distill-8k 🐳 DeepSeek-V4-Distill-8100x Dataset Summary DeepSeek-V4-Distill-8100x is a supervised fine-tuning dataset for reasoning-oriented distillation. The question prompts come from Jackrong/GLM-5.1-Reasoning-1M-Cleaned, and the answers were generated by the teacher model DeepSeek-V4-Flash. After the cleaning process, the released train split contains 7,716 high-quality JSONL examples. [!NOTE] The answer pool was cleaned to remove real-time questions… See the full description on the dataset page: https://huggingface.co/datasets/ansulev/deepseek-v4-distill-8k.texttext-generation1K<n<10K0 likes49 downloads5mo agoHugging Face15open-llm-leaderboard /grimjim__Magnolia-v2-Gemma2-8k-9B-detailsgated Dataset Card for Evaluation run of grimjim/Magnolia-v2-Gemma2-8k-9B Dataset automatically created during the evaluation run of model grimjim/Magnolia-v2-Gemma2-8k-9B The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/grimjim__Magnolia-v2-Gemma2-8k-9B-details.tabular10K<n<100K0 likes39 downloads2y agoHugging Face16Lambent /1k-creative-writing-8kt-fineweb-edu-sampleTotal tokens in matching entries: 5_575_157 Average tokens per entry: 5575.16 tabular1K<n<10K0 likes38 downloads2y agoHugging Face17Lambent /20k-finewebedu-samples-8kttabular10K<n<100K0 likes37 downloads2y agoHugging Face18merileijona /quantum-circuits-8k Quantum Circuits 8K Dataset A synthetic dataset of 8,129 quantum circuit examples for training language models to generate OpenQASM 2.0 code from natural language descriptions. Quick Stats Total Samples: 8,129 (description → QASM pairs) Unique Circuits: 739 base circuits Categories: 92 distinct quantum circuit types Qubit Range: 1-9 qubits Format: OpenQASM 2.0 Augmentation: 11x per circuit (original + 10 paraphrases) Quality: 100% QASM syntax valid, 0% duplicates… See the full description on the dataset page: https://huggingface.co/datasets/merileijona/quantum-circuits-8k.texttext-generation10K<n<100K1 likes37 downloads6mo agoHugging Face19kgrabko /JiRack-No_Robots_8k-Datasettext1K<n<10K0 likes36 downloads5mo agoHugging Face20botbotrobotics /ptbr-deita-8k PTBR Deita 8k Portuguese translation of the Deita 8k dataset. text1K<n<10K0 likes31 downloads2y agoHugging Face21Lambent /100k-finewebedu-samples-8kttabular100K<n<1M1 likes27 downloads2y agoHugging Face22Hancovirus /8k_filterd_traintext1K<n<10K0 likes25 downloads1y agoHugging Face23leonardklin /OpenSciDER-SFT-8K OpenSciDER This is the model repo for OpenSciDER-SFT-8K, a trajectory dataset curated from SciDER. This dataset is collected from the inference trajectories of the Qwen3.6-27B and OpenSciDER-27B model on DataSciBench, DS-1000, DS-Bench, and ScienceAgentBench. It also contains the benchmark evaluation trajectories on AI-Idea-Bench, AIRS-Bench, AstroVisBench, DiscoveryBench, MLE-Bench, and SciCode. There are 3 configurations in this dataset: openscider: the trajectories of… See the full description on the dataset page: https://huggingface.co/datasets/leonardklin/OpenSciDER-SFT-8K.texttext-generation10K<n<100K0 likes24 downloads4mo agoHugging Face24awakara /Math-Japanese-8k算数の文章問題のデータセット 問題文をKimi K2.6 K2.7-codeで作成 回答をGemma4 31Bで作成 text1K<n<10K0 likes23 downloads28d agoHugging Face25ceadar-ie /AIVision360-8k Dataset Card for AIVision360-8k Dataset Description AIVision360 is the pioneering domain-specific dataset tailor-made for media and journalism, designed expressly for the instruction fine-tuning of Large Language Models (LLMs).The AIVision360-8k dataset is a curated collection sourced from "ainewshub.ie", a platform dedicated to Artificial Intelligence news from quality-controlled publishers. It is designed to provide a comprehensive representation of AI-related… See the full description on the dataset page: https://huggingface.co/datasets/ceadar-ie/AIVision360-8k.textquestion-answering1K<n<10K4 likes22 downloads3y agoHugging Face26NewEden /CAI-Rev-1-Opus-22K-8K-Subset-Non-Convertedtext1K<n<10K0 likes22 downloads9mo agoHugging Face27syuekai /ruler2_8ktext1K<n<10K0 likes22 downloads7mo agoHugging Face28open-llm-leaderboard /microsoft__Phi-3-small-8k-instruct-detailsgated Dataset Card for Evaluation run of microsoft/Phi-3-small-8k-instruct Dataset automatically created during the evaluation run of model microsoft/Phi-3-small-8k-instruct The dataset is composed of 34 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/microsoft__Phi-3-small-8k-instruct-details.tabular10K<n<100K0 likes19 downloads2y agoHugging Face29vineetdaniels /nyxmed-icd-cpt-8ktext1K<n<10K0 likes19 downloads9mo agoHugging Face30jbe /flashback-samhalle-8k flashback-samhalle-8k Swedish Flashback "samhalle" subforum, parsed with --max-chars 8000 into ShareGPT-format chat conversations. Voice-cloning / reactive-agent dataset. Each record is one whole thread (multi-turn conversations); row-level train/validation split has no post-leakage. texttext-generation100K<n<1M0 likes19 downloads4mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.