CoolFace
16 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01celsowm /legal_br_sft Legal BR SFT Dataset ⚖️🇧🇷 (Auditado) O Legal BR SFT é um dataset de instruções de alta qualidade focado exclusivamente no Direito Brasileiro. Ele foi projetado para o treinamento de modelos de linguagem (LLMs) através de Supervised Fine-Tuning (SFT). 📊 Estatísticas Auditadas (Regex Refinado) Após auditoria estatística estratificada em 38.153 registros, a distribuição por área do Direito é: Área do Direito Porcentagem Temas Principais Direito Civil 19… See the full description on the dataset page: https://huggingface.co/datasets/celsowm/legal_br_sft.texttext-generation100K<n<1M0 likes61 downloads3mo agoHugging Face02celsowm /jurisprudencias_stftexttext-generation1K<n<10K1 likes60 downloads2y agoHugging Face03celerity-labs /celeritybench-tool-choice Which small model should run a Mac launcher Celeritas is a Spotlight-style launcher that turns what somebody types into tool calls on their own machine. Picking the model to put behind it meant measuring them, and the numbers were going on a public page, so the runs behind them are here. The question is narrow on purpose: for an agent with about thirty tools on a desktop, which model picks the right one? Not reasoning, not code, not knowledge. Tool choice, on short everyday… See the full description on the dataset page: https://huggingface.co/datasets/celerity-labs/celeritybench-tool-choice.tabulartext-generation1K<n<10K0 likes53 downloads7d agoHugging Face04chrishayuk /v11-cells-midtrain-corpus v11 cells mid-training corpus The delegating arm of a paired experiment: teach a 115M model to call an external tool for arithmetic rather than to memorise the answers. Its partner, the maths-only arm, teaches the same model to absorb the arithmetic into its weights instead. Pre-tokenized against the v11 tokenizer (10dd5110…, vocab 71,260), for chrishayuk/v11-tinystories-115m-base. Identity: 2115d6aeff3428e217ef2903a8030facd511dcb00183e9fc3faaf49d01038767 (chuk-datasets… See the full description on the dataset page: https://huggingface.co/datasets/chrishayuk/v11-cells-midtrain-corpus.texttext-generationn<1K0 likes38 downloads2mo agoHugging Face05celsowm /gemini_orpo_dpo_ptbrtexttext-generation10K<n<100K2 likes34 downloads2y agoHugging Face06ce-lery /merged-corpus Merged Corpus Welcome to this repository.This dataset is japanese corpus that includes wiki, wikibooks, wikiversity, cc100, and oscar2109. Getting Started If you want to use this, please run as follows.This process takes about 3 hours. mkdir -p pretrain/input/ cd pretrain/input/ GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/datasets/ce-lery/merged-corpus.git cd merged-corpus git lfs pull bash merge_train.sh texttext-generation10M<n<100M0 likes31 downloads1y agoHugging Face07celsowm /leis_ordinarias_1988_2024textsummarization1K<n<10K2 likes26 downloads2y agoHugging Face08celsowm /valdoria-dpo-qwen35-dataset Valdoria DPO dataset Dataset de preferências conversacional para uso direto com trl.DPOTrainer. Splits train.jsonl: 1.889 pares validation.jsonl: 236 pares test.jsonl: 237 pares Cada linha contém: { "id": "...", "prompt": [{"role": "system", "content": "..."}, {"role": "user", "content": "..."}], "chosen": [{"role": "assistant", "content": "..."}], "rejected": [{"role": "assistant", "content": "..."}], "metadata": {"task_type": "..."… See the full description on the dataset page: https://huggingface.co/datasets/celsowm/valdoria-dpo-qwen35-dataset.texttext-generation1K<n<10K0 likes25 downloads3mo agoHugging Face09CelesteLove /minecraft_qa_es Minecraft Q&A (Spanish) A Spanish, chat-formatted question/answer dataset about Minecraft. Each example is a short conversation with a single user question and a single assistant answer (plus a system prompt). Data format The dataset is provided as JSON Lines (.jsonl): one JSON object per line. Each record has a single key: messages: an array of chat messages, each with: role: one of "system", "user", "assistant" content: the message text Typical structure:… See the full description on the dataset page: https://huggingface.co/datasets/CelesteLove/minecraft_qa_es.textquestion-answering10K<n<100K0 likes24 downloads9mo agoHugging Face10celsowm /modelos_peticoesHTML was converted to Markdown (better for LLMs) texttext-generationn<1K1 likes23 downloads2y agoHugging Face11celsowm /valdoria-sft-qwen35-dataset Dataset Card — Valdoria SFT Pack v3.3.1 Dataset sintético em português para Supervised Fine-Tuning (SFT), baseado em um país fictício chamado República de Valdoria. Formato ChatML (messages), combinando os splits train + validation (2125 exemplos). Descrição Cobertura de tarefas: factual_qa (358) fantasy_boundary (534) refusal (251) unknown_canonical_field (252) negative_case (200) decision_making (191) rule_application (108) classification (131)… See the full description on the dataset page: https://huggingface.co/datasets/celsowm/valdoria-sft-qwen35-dataset.texttext-generation1K<n<10K0 likes18 downloads3mo agoHugging Face12celsowm /verbetes_vademecum_direitoOriginal source: https://vademecumbrasil.com.br/dicionario-juridico/ texttext-generation10K<n<100K0 likes17 downloads2y agoHugging Face13ce-lery /corpus-ja-11b corpus-ja-11b Welcome to this repository.This dataset is japanese corpus that includes wiki, wikibooks, wikiversity, and cc100. Getting Started If you want to use this, please run as follows.This process takes about 3 hours. mkdir -p pretrain/input/ cd pretrain/input/ GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/datasets/ce-lery/corpus-ja-11b.git cd corpus-ja-11b git lfs pull bash merge_train.sh texttext-generation10M<n<100M0 likes17 downloads1y agoHugging Face14dp1812 /celestial-spiritual-ai-dataset CELESTIAL Spiritual AI Training Dataset Overview Comprehensive spiritual AI training dataset with 190+ conversations covering all 16 CELESTIAL spiritual systems. Features Covered Vedic Astrology & Kundli Generation Numerology & Life Path Analysis Divine AI Personas (Krishna, Ganesha, Shiva, Devi, Hanuman, Saraswati) Vastu Shastra & Palmistry Spiritual Guidance & Meditation And 11 more spiritual systems! Quality Metrics Overall Score: 94.9%… See the full description on the dataset page: https://huggingface.co/datasets/dp1812/celestial-spiritual-ai-dataset.texttext-generationn<1K0 likes14 downloads1y agoHugging Face15fatihburakkaragoz /evliya-celebi-seyahatname-ocr Evliya Celebi Seyahatname OCR Corpus OCR-derived text from seven volumes of Evliya Celebi's Seyahatname, packaged for corpus exploration, language modeling, OCR-quality analysis, and historical Ottoman Turkish / Turkish NLP work. Configs pages: one row per OCR page, with page numbers and OCR status. documents: one row per available volume, with page text concatenated. Coverage Available books: 1, 3, 4, 6, 7, 9, 10. Missing from the 1-10 sequence: 2, 5, 8.… See the full description on the dataset page: https://huggingface.co/datasets/fatihburakkaragoz/evliya-celebi-seyahatname-ocr.tabulartext-generation1K<n<10K1 likes14 downloads5mo agoHugging Face16lianghsun /tw-celebrity-chatgated Dataset Card for tw-celebrity-chat 本資料集是以臺灣公眾人物(藝人、團體、運動員、創作者等)相關之維基百科條目為種子(seed),由 LLM 合成的繁體中文「思考+回答」對話集,可用於訓練具備 reasoning 能力的繁中對話模型,特別是針對臺灣演藝/流行文化領域。 Dataset Details Dataset Description 資料集為 reference-based 合成對話:先以維基百科上臺灣公眾人物相關段落作為 seed,再透過 LLM 生成單輪 QA,並在回答前加上 <think>...</think> 思考段落。每筆樣本同時保留: conversations:human/gpt 兩輪結構(gpt 回答中含思考段落); input / output / think:把上述拆解出來的單獨欄位,方便不同訓練框架使用; seed:用來提示 LLM 的原始參考文本。 主題涵蓋臺灣男女歌手、團體(如 183Club、5566、SHE 等)、戲劇演員、棒球與 SBL… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-celebrity-chat.texttext-generation10K<n<100K0 likes4 downloads5mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.