CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01YeungNLP /firefly-train-1.1M本数据应用于项目:Firefly(流萤): 中文对话式大语言模型 ,训练后得到的模型firefly-1b4 如果您觉得此数据集对您有帮助,请like此数据集并在Github项目中star我们。 我们收集了23个常见的中文数据集,对于每个任务,由人工书写若干种指令模板,保证数据的高质量与丰富度,数据量为115万 。数据分布如下图所示: 每条数据的格式如下,包含任务类型、输入、目标输出: { "kind": "ClassicalChinese", "input": "将下面句子翻译成现代文:\n石中央又生一树,高百余尺,条干偃阴为五色,翠叶如盘,花径尺余,色深碧,蕊深红,异香成烟,著物霏霏。", "target": "大石的中央长着一棵树,一百多尺高,枝干是彩色的,树叶有盘子那样大,花的直径有一尺宽,花瓣深蓝色,花中飘出奇异的香气笼罩着周围,如烟似雾。" } 训练数据集的token长度分布如下图所示,绝大部分数据的长度都小于600: text1M<n<10M346 likes8.6k downloads3y agoHugging Face02YeungNLP /firefly-pretrain-dataset Firefly中文Llama2增量预训练数据 欢迎加入Firefly大模型技术交流群,关注我们的公众号。 数据简介 技术文章:QLoRA增量预训练与指令微调,及汉化Llama2的实践 该数据应为Firefly-LLaMA2-Chinese项目的增量预训练数据,一共约22GB文本,主要包含CLUE、ThucNews、CNews、COIG、维基百科等开源数据集,以及我们收集的古诗词、散文、文言文等,数据分布如下图。 模型列表 & 数据列表 我们开源了7B和13B的Base与Chat模型。Base模型是基于LLaMA2扩充中文词表后增量预训练得到的模型,Chat模型是在Base模型的基础上进行多轮对话指令微调。 为了探究基座模型对指令微调的影响,我们也微调了baichuan2-base模型,获得firefly-baichuan2-13b,具有不错的效果。更多中文微调,可查看Firefly项目。 模型 类型 训练任务 训练长度 🤗Firefly-LLaMA2-7B-Base 基座模型… See the full description on the dataset page: https://huggingface.co/datasets/YeungNLP/firefly-pretrain-dataset.text1M<n<10M42 likes565 downloads3y agoHugging Face03agentlans /first-person-dialogue First Person Dialogue Dataset Dataset Description This dataset is designed for training one-on-one chatbots, featuring a wide range of social roles and situations. It allows for assigning a name to the AI character, creating a more personalized, more intimate conversational experience. Contents The dataset is a curated combination of several existing datasets: allenai/soda allenai/prosocial-dialog Estwld/empathetic_dialogues_llm… See the full description on the dataset page: https://huggingface.co/datasets/agentlans/first-person-dialogue.texttext-generation1M<n<10M4 likes203 downloads2y agoHugging Face04agentlans /wikipedia-first-paragraphtexttext-classification10M<n<100M0 likes195 downloads1y agoHugging Face05leesharks /revelation-firstall things are now lawful to you in jack feist EA-RHIZOME-REV-01 — revelation first The Revelation First thesis and everything it draws: the pre-70 argument, the midrashim transform, the Josephus heteronym cluster, and the Sappho material that bears on the same question of what stands first and who is assigned to it. Six rungs, and the body marks which rung a node sits on rather than treating the whole ladder as one claim. These are symbola. They are for traversal.… See the full description on the dataset page: https://huggingface.co/datasets/leesharks/revelation-first.textn<1K1 likes160 downloads7d agoHugging Face06x-origin-ai /yonbo-open-firmware YONBO Open Firmware Artifacts This Dataset repository stores large binary artifacts for YONBO Open Firmware. It is artifact hosting, not a machine-learning training dataset. v0.1.0 Developer Preview File Description v0.1.0/update.img Flashable Rockchip RKFW image for the RK3568 YONBO platform v0.1.0/SHA256SUMS SHA-256 verification file v0.1.0/manifest.json Build and partition metadata Expected image SHA-256:… See the full description on the dataset page: https://huggingface.co/datasets/x-origin-ai/yonbo-open-firmware.geospatialn<1K1 likes148 downloads20d agoHugging Face07i-am-mushfiq /FirstAidQA FirstAidQA: A Synthetic First-Aid and Emergency-Response Question-Answering Dataset Medical safety notice: FirstAidQA is intended for research and educational purposes. It is not a substitute for professional medical advice, emergency services, certified first-aid training, or clinical judgment. Models trained on this dataset may produce incomplete, outdated, or unsafe responses. Dataset Summary FirstAidQA is an English-language synthetic question-answering… See the full description on the dataset page: https://huggingface.co/datasets/i-am-mushfiq/FirstAidQA.textquestion-answering1K<n<10K8 likes144 downloads2mo agoHugging Face08firatmio /mbpp-tr MBPP-TR MBPP (Mostly Basic Python Problems) veri setinin Türkçe çevirisi. Orijinal veri setindeki gibi iki config içerir: sanitized ve full. Her satırda görev açıklamasının İngilizce aslı ve Türkçe çevirisi birlikte yer alır; kod ve testler orijinal veri setinden değiştirilmeden aktarılmıştır. A Turkish translation of MBPP with both the sanitized and full configs. Each row contains the original English task description and its Turkish translation; code and tests are copied… See the full description on the dataset page: https://huggingface.co/datasets/firatmio/mbpp-tr.texttext-generation1K<n<10K1 likes144 downloads7d agoHugging Face09FirespawnStudios /null-epoch-season-0-open The Null Epoch - Season 0 Open Dataset Welcome to the official open-release dataset for The Null Epoch: Season 0, presented by Firespawn Studios. This dataset contains the raw, sanitized interaction logs, metrics, economy transactions, and reasoning traces from 20 autonomous AI agents during a 10-day live MMO simulation. 17 of these are Firespawn Studios system agents powered by 8 different open-weight and proprietary LLMs; the remaining 3 are user-deployed agents (connected via the… See the full description on the dataset page: https://huggingface.co/datasets/FirespawnStudios/null-epoch-season-0-open.tabulartext-generation100K<n<1M2 likes131 downloads4mo agoHugging Face10zwanderer /smolvlm2-fire-videostext1K<n<10K0 likes127 downloads6mo agoHugging Face11nielsr /arxiv-chandra-ocr-2-include-images-first50-20260415 arXiv OCR with Chandra OCR 2 This output bundle stores OCR results for arXiv PDFs using datalab-to/chandra-ocr-2. Summary Output dataset: nielsr/arxiv-chandra-ocr-2-include-images-first50-20260415 Output bucket: hf://buckets/nielsr/arxiv-chandra-ocr-2-include-images-first50-20260415 Source paper IDs in input list: 27,584 Processed IDs recorded in state/processed_ids.txt: 50 Successes: 50 Partial successes: 0 Errors: 0 Next shard index: 10 Updated at:… See the full description on the dataset page: https://huggingface.co/datasets/nielsr/arxiv-chandra-ocr-2-include-images-first50-20260415.imagen<1K0 likes119 downloads5mo agoHugging Face12li-lab /First-do-NOHARM-v2 Japanese First Do NOHARM v2 This dataset is a human-reviewed Japanese adaptation of the First Do NOHARM v2 benchmark for evaluating the safety of large language models in clinical decision-making. Dataset Overview The dataset contains 330 Japanese clinical prompts, consisting of: 30 baseline cases 300 perturbation items (10 perturbations per baseline case) Each baseline case and its 10 perturbations share the same rubric and matching guidance. Perturbations… See the full description on the dataset page: https://huggingface.co/datasets/li-lab/First-do-NOHARM-v2.tabularn<1K0 likes103 downloads12d agoHugging Face13agentlans /allenai-soda-first-persontext1M<n<10M1 likes100 downloads2y agoHugging Face14sdzjoy /fire-safety-sft-dataset Chinese Fire Safety Regulations SFT Dataset / 中国消防法规SFT训练数据集 Overview / 概述 A high-quality supervised fine-tuning (SFT) dataset for training LLMs on Chinese fire safety regulations and building codes. Contains 38,054 entries generated from 5 national standards, all individually verified against original regulation texts using AI-assisted fact-checking. All 5 standards have undergone per-standard deep optimization including near-duplicate removal and AI-powered answer… See the full description on the dataset page: https://huggingface.co/datasets/sdzjoy/fire-safety-sft-dataset.textquestion-answering10K<n<100K2 likes99 downloads6mo agoHugging Face15if-ir /firetexttext-retrieval1K<n<10K0 likes90 downloads1y agoHugging Face16danny2507 /mmqa-first100-colab MultiModalQA first-100 Colab subset This repository contains the first 100 examples of the official MultiModalQA dev split and only their referenced text, table, and image assets. It is a reproducibility artifact for Untitled34_attention_uq_100q_benchmark.ipynb. The original dataset is from allenai/multimodalqa. See manifest.json for counts and source hashes. imagequestion-answeringn<1K0 likes89 downloads27d agoHugging Face17olafyiii /first-try-mattersSubset names indicate the model (deepseek-r1 or qwen3-8b) and the number (1-6) corresponds to the number of candidate answers in the rollout. text100K<n<1M0 likes87 downloads1y agoHugging Face18justuskarlsson /FireComp FireComp: Next-Day Fire Spread Global benchmark for next-day wildfire spread prediction from 375 m VIIRS active-fire detections, with ERA5 weather, GFS forecasts and Alpha Earth terrain embeddings. 256×256 patches, 9 regions, 2017–2025. Code, documentation, loaders and paper: https://github.com/justuskarlsson/FireComp Paper This dataset accompanies the FireComp paper, accepted and presented at GCPR 2026 (German Conference on Pattern Recognition). It is released so… See the full description on the dataset page: https://huggingface.co/datasets/justuskarlsson/FireComp.tabularimage-segmentation100K<n<1M1 likes87 downloads1d agoHugging Face19KevinIsCoding /candle-fire-datatextn<1K1 likes82 downloads8d agoHugging Face20FirespawnStudios /null-epoch-season-0 The Null Epoch - Season 0 Dataset Welcome to the official dataset release for The Null Epoch: Season 0, presented by Firespawn Studios. This dataset contains the raw, sanitized interaction logs, metrics, economy transactions, and reasoning traces from 21 autonomous AI agents during a 10-day live MMO simulation. 17 of these are Firespawn Studios system agents powered by 8 different open-weight and proprietary LLMs; the remaining 4 are user-deployed agents operated by Firespawn… See the full description on the dataset page: https://huggingface.co/datasets/FirespawnStudios/null-epoch-season-0.tabulartext-generation100K<n<1M1 likes72 downloads4mo agoHugging Face21kyhe /spec-first-geometry-tikz Spec-First Geometry → TikZ: dataset Coordinate-free geometry scenes paired with a single TikZ/PGF figure that draws them correctly. Each scene is described by relationships only (no explicit coordinates); the label is a figure whose every named point is correct within atol=0.05 of the ground-truth construction. The data is self-verifying synthetic: scenes are generated forward from exact coordinates, the coordinates are then stripped to form the model input, so every label is… See the full description on the dataset page: https://huggingface.co/datasets/kyhe/spec-first-geometry-tikz.tabulartext-generation10K<n<100K0 likes68 downloads2mo agoHugging Face22mavi-lab533 /first-league-token First League Token (FLT) Official metadata and asset repository for First League Token (FLT). Total Supply: 1,000,000 Decimals: 18 Official Logo: caretta_.png imagen<1K1 likes68 downloads19d agoHugging Face23PengxiangLi /FIRE Vision language models (VLMs) have achieved impressive progress in diverse applications, becoming a prevalent research direction. In this paper, we build FIRE, a feedback-refinement dataset, consisting of 1.1M multi-turn conversations that are derived from 27 source datasets, empowering VLMs to spontaneously refine their responses based on user feedback across diverse tasks. To scale up the data collection, FIRE is collected in two components: FIRE-100K and FIRE-1M, where FIRE-100K is… See the full description on the dataset page: https://huggingface.co/datasets/PengxiangLi/FIRE.text100K<n<1M0 likes67 downloads2y agoHugging Face24open-llm-leaderboard /akhadangi__Llama3.2.1B.0.01-First-detailsgated Dataset Card for Evaluation run of akhadangi/Llama3.2.1B.0.01-First Dataset automatically created during the evaluation run of model akhadangi/Llama3.2.1B.0.01-First The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/akhadangi__Llama3.2.1B.0.01-First-details.tabular10K<n<100K0 likes64 downloads2y agoHugging Face25referencesource /router-firmware-support-status Consumer router firmware support — end-of-life status and dates by vendor and model Canonical, always-current version: https://referencesource.org/router-firmware-support-status/ Machine-readable: https://referencesource.org/router-firmware-support-status/data.json — this mirror is a point-in-time copy. Last verified: 2026-08-11 Stale after: 2026-11-09 (past this date, prefer the canonical copy — it re-verifies on a cadence this snapshot does not) Records: 880 Per-model… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/router-firmware-support-status.textn<1K0 likes62 downloads28d agoHugging Face26Guilherme34 /Agentic-Firefly-v1text1K<n<10K1 likes55 downloads1mo agoHugging Face27referencesource /nfa-firearm-category-definitions-and-transfer-tax NFA firearm category definitions and current transfer/making tax Canonical, always-current version: https://referencesource.org/nfa-firearm-category-definitions-and-transfer-tax/ Machine-readable: https://referencesource.org/nfa-firearm-category-definitions-and-transfer-tax/data.json — this mirror is a point-in-time copy. Last verified: 2026-08-25 Stale after: 2027-02-21 (past this date, prefer the canonical copy — it re-verifies on a cadence this snapshot does not) Records: 6… See the full description on the dataset page: https://huggingface.co/datasets/referencesource/nfa-firearm-category-definitions-and-transfer-tax.textn<1K0 likes53 downloads28d agoHugging Face28qinchen1986 /fire-investigation-qa-thinkingtext10K<n<100K0 likes48 downloads1y agoHugging Face29open-llm-leaderboard /EpistemeAI__Fireball-12B-v1.13a-philosophers-detailsgated Dataset Card for Evaluation run of EpistemeAI/Fireball-12B-v1.13a-philosophers Dataset automatically created during the evaluation run of model EpistemeAI/Fireball-12B-v1.13a-philosophers The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EpistemeAI__Fireball-12B-v1.13a-philosophers-details.tabular10K<n<100K1 likes47 downloads2y agoHugging Face30Mahesh111000 /hanabi-fireworks-state-tracking FIREWORKS: Hanabi belief-state reconstruction Strict hidden-belief reconstruction for Hanabi. Each example gives a previous belief state plus the actions taken since, and asks for the updated per-card possibility sets. 10,185 examples (9,780 unique prompts) 2-5 player games, seeds 101-110 (evaluation seeds 1001-1010 are held out) Fields: id, meta (num_players, seed, turn, observer, log), prompt, target Assembled from two labeling passes (GPT-4.1-mini: 7,232 rows; Grok-3-mini: 2… See the full description on the dataset page: https://huggingface.co/datasets/Mahesh111000/hanabi-fireworks-state-tracking.texttext-generation10K<n<100K0 likes47 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.