CoolFace
15 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01whoisandy /router-chat-normalized-1m Router Chat Normalized 1M Dataset Description Router Chat Normalized 1M is a multilingual conversational dataset containing 1,353,300 conversations normalized from multiple public chat datasets with automatic language detection. Dataset Structure The dataset contains 2 split(s): train, test. Each conversation includes: conversation_id, messages (list of {role, content} structs), source dataset, detected language, and language confidence score. Source… See the full description on the dataset page: https://huggingface.co/datasets/whoisandy/router-chat-normalized-1m.texttext-generation1M<n<10M0 likes593 downloads4mo agoHugging Face02digi-texx /calib-agentic-normalized calib-agentic (v1) Normalized calibration corpus for post-training quantization. Every row is OpenAI-shaped (messages + tools) regardless of upstream dialect, with provenance in origin and precomputed stats for stratified selection. Schema messages[] — role (system/user/assistant/tool), content, name, tool_call_id, tool_calls[] (name, arguments as a JSON string), reasoning_content tools[] — name, description, parameters (JSON string of the JSON-Schema) origin —… See the full description on the dataset page: https://huggingface.co/datasets/digi-texx/calib-agentic-normalized.text-generation1M<n<10M0 likes94 downloads21d agoHugging Face03jajostrains /Mathlib-Normalized-Sexpr Mathlib Normalized S-Expressions Lean 4 proof states from Mathlib, paired with the tactic applied at each step, in three representations extracted directly from the Lean kernel: Source-faithful S-expressions of the goal and every hypothesis, as Lean elaborated them. Normalized S-expressions of the same state, with stable local-context indices suitable for model input. Annotated tactic syntax -- the original tactic's syntax tree with identifier leaves resolved to the constants… See the full description on the dataset page: https://huggingface.co/datasets/jajostrains/Mathlib-Normalized-Sexpr.tabulartext-generation100K<n<1M0 likes78 downloads26d agoHugging Face04francescortu /DistillDetect-normalized-traces DistillDetect — format-normalized teacher traces Teacher responses from Reference-Based Distillation Detection in LLMs (arXiv:2607.09692), rewritten so that every teacher uses the same output format. 7,918 rows across 8 teacher/prompt-set pairs. Why this exists In the released data each teacher emits a structurally different response, so a student trained on it — and any detector trained to attribute it — can key on surface format instead of the teacher's actual… See the full description on the dataset page: https://huggingface.co/datasets/francescortu/DistillDetect-normalized-traces.texttext-generation1K<n<10K0 likes77 downloads28d agoHugging Face05llami-team /Korean-OpenThoughts-114k-NormalizedKorean-OpenThoughts-114k-Normalized 상세 데이터셋 설명 OpenThoughts-114k-Normalized 데이터셋의 한국어 번역본입니다. OpenAI gpt-4o-mini를 통해 번역됐습니다. Shared by llami-team Language(s) (NLP): Korean Uses 한국어 reasoning 모델 distillation reasoning cold-start 데이터셋 Dataset Structure question: 질문 reasoning: 추론 과정 response: 응답 Dataset Creation [LLAMI Team] (https://llami.net) LLAMI Github lemon-mint Source Data OpenThoughts-114k-Normalized texttext-generation100K<n<1M28 likes75 downloads2y agoHugging Face06cstr /de-wiktionary-sqlite-normalized German Wiktionary - Normalized SQLite Database A fully normalized, production-ready SQLite database of German Wiktionary with complete linguistic information and optimized query performance. 🎯 Key Features ✅ Zero data loss: All information from original Wiktionary preserved ⚡ Lightning-fast queries: Comprehensive indexing (< 5ms typical queries) 🔍 Full grammatical analysis: Complete inflection paradigms, word forms, 185 unique grammatical tags 🔗 Semantic relations:… See the full description on the dataset page: https://huggingface.co/datasets/cstr/de-wiktionary-sqlite-normalized.text-retrieval100K<n<1M0 likes50 downloads10mo agoHugging Face07atrevidasadia /dair-ai-emotion-normalized-instruction-input-output dair-ai emotion | normalized Summary Dataset ID: 143 Type: normalized Rows: 16,000 Source: dair-ai/emotion Dataset Sources #143 dair-ai emotion | normalized [normalized | 16,000 rows] Notes Edited and Exported from the Kitsune Training Suite (Forge) Review the dataset artifact and metadata before publishing. Citation > via dair-ai @inproceedings{saravia-etal-2018-carer, title = "{CARER}: Contextualized Affect… See the full description on the dataset page: https://huggingface.co/datasets/atrevidasadia/dair-ai-emotion-normalized-instruction-input-output.texttext-generation10K<n<100K0 likes44 downloads28d agoHugging Face08EmmaLeonhart /normalized-wikidata Normalized Wikidata A preprocessed text-form view of Wikidata, optimised for training language models or knowledge-graph world models. The goal is a corpus where the semantic content of Wikidata triples comes through cleanly, with the catalog-and-identifier clutter that dominates raw Wikidata by volume stripped out. License inherits from Wikidata: CC-BY-SA 4.0. This dataset is the input to a corresponding series of Loka world-model checkpoints at EmmaLeonhart/loka. Each snapshot… See the full description on the dataset page: https://huggingface.co/datasets/EmmaLeonhart/normalized-wikidata.text-generation1M<n<10M0 likes36 downloads4mo agoHugging Face09kiarashrzg /TinyPersianStories_normalizedtexttext-generation100K<n<1M0 likes26 downloads2y agoHugging Face10deltakitsune /dair-ai-emotion-normalized-instruction-input-output dair-ai emotion | normalized Summary Dataset ID: 143 Type: normalized Rows: 16,000 Source: dair-ai/emotion Dataset Sources #143 dair-ai emotion | normalized [normalized | 16,000 rows] Notes Edited and Exported from the Kitsune Training Suite (Forge) Review the dataset artifact and metadata before publishing. Citation > via dair-ai @inproceedings{saravia-etal-2018-carer, title = "{CARER}: Contextualized Affect Representations for… See the full description on the dataset page: https://huggingface.co/datasets/deltakitsune/dair-ai-emotion-normalized-instruction-input-output.texttext-generation10K<n<100K0 likes26 downloads5mo agoHugging Face11userdavek /Amharic_news_Normalized Dataset Name Amharic news dataset Dataset Details It is a non-normalized version of news dataset crawled from Amharic news websites and from researchers provided in their works. Dataset Description The dataset is collected from different news websites and from different researchers crawled Amharic news dataset from different NLP downstream tasks. News sites like FanaBC, EthiopianReporter, Zehabesha,Esat Amharic, BBC Amharic are the sources for these news data.… See the full description on the dataset page: https://huggingface.co/datasets/userdavek/Amharic_news_Normalized.textsummarization100K<n<1M0 likes24 downloads11mo agoHugging Face12tuandunghcmut /travelplanner-benchmark-normalized TravelPlanner Benchmark (Normalized) Normalized, typed, parquet-first packaging of the TravelPlanner benchmark for planning-centric agent evaluation. Upstream dataset: osunlp/TravelPlanner Upstream code: OSU-NLP-Group/TravelPlanner Paper: TravelPlanner: A Benchmark for Real-World Planning with Language Agents 1) What is included This dataset repo contains: benchmark config (train/validation/test) in typed parquet. reference_entries config: flattened reference-info… See the full description on the dataset page: https://huggingface.co/datasets/tuandunghcmut/travelplanner-benchmark-normalized.tabulartext-generation10K<n<100K0 likes23 downloads7mo agoHugging Face130x7o /dostoevsky_frontier_3k_normalized dostoevsky_frontier_3k_normalized Нормализованная версия 0x7o/dostoevsky_frontier_3k. Нормализация Устранены пунктуационные shortcut-ы, позволяющие модели различать chosen/rejected по артефактам форматирования вместо стиля. Общие (chosen + rejected) ё → е по словарю (книги не используют ё, AI всегда использует — 87.8% accuracy) \xa0 (неразрывный пробел) → обычный пробел … (U+2026) → ... (три точки) – (en dash) → — (em dash) !.. → !..., ?.. → ?... "..." → «...»… See the full description on the dataset page: https://huggingface.co/datasets/0x7o/dostoevsky_frontier_3k_normalized.texttext-generation1K<n<10K1 likes21 downloads7mo agoHugging Face140x7o /dostoevsky_frontier_v2_normalized dostoevsky_frontier_v2_normalized Нормализованная версия 0x7o/dostoevsky_frontier_v2. Нормализация Устранены пунктуационные shortcut-ы, позволяющие модели различать chosen/rejected по артефактам форматирования вместо стиля. Общие (chosen + rejected) ё → е по словарю (книги не используют ё, AI всегда использует — 87.8% accuracy) \xa0 (неразрывный пробел) → обычный пробел … (U+2026) → ... (три точки) – (en dash) → — (em dash) !.. → !..., ?.. → ?... "..." → «...»… See the full description on the dataset page: https://huggingface.co/datasets/0x7o/dostoevsky_frontier_v2_normalized.texttext-generationn<1K0 likes17 downloads7mo agoHugging Face15muammar-jsx /huatuo-family-normalized-combined Huatuo26M-Lite 📚 Table of Contents 🗂 Dataset Description 📝 Dataset Information ℹ️ Data Distribution 📊 Usage 🔧 Citation 📖 Dataset Description 📝 Huatuo26M-Lite is a refined and optimized dataset based on the Huatuo26M dataset, which has undergone multiple purification processes and rewrites. It has more data dimensions and higher data quality. We welcome you to try using it. Dataset Information ℹ️ Dataset Name:… See the full description on the dataset page: https://huggingface.co/datasets/muammar-jsx/huatuo-family-normalized-combined.text-classification100K<n<1M0 likes7 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.