datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
legal_br_sft
Legal BR SFT Dataset ⚖️🇧🇷 (Auditado)
O Legal BR SFT é um dataset de instruções de alta qualidade focado exclusivamente no Direito Brasileiro. Ele foi projetado para o treinamento de modelos de linguagem (LLMs) através de Supervised Fine-Tuning (SFT).
📊 Estatísticas Auditadas (Regex Refinado)
Após auditoria estatística estratificada em 38.153 registros, a distribuição por área do Direito é:
Área do Direito
Porcentagem
Temas Principais
Direito Civil
19… See the full description on the dataset page: https://huggingface.co/datasets/celsowm/legal_br_sft.jurisprudencias_stfceleritybench-tool-choice
Which small model should run a Mac launcher
Celeritas is a Spotlight-style launcher that turns what somebody types into tool
calls on their own machine. Picking the model to put behind it meant measuring
them, and the numbers were going on a public page, so the runs behind them are
here.
The question is narrow on purpose: for an agent with about thirty tools on a
desktop, which model picks the right one? Not reasoning, not code, not
knowledge. Tool choice, on short everyday… See the full description on the dataset page: https://huggingface.co/datasets/celerity-labs/celeritybench-tool-choice.v11-cells-midtrain-corpus
v11 cells mid-training corpus
The delegating arm of a paired experiment: teach a 115M model to call an external
tool for arithmetic rather than to memorise the answers. Its partner, the maths-only
arm, teaches the same model to absorb the arithmetic into its weights instead.
Pre-tokenized against the v11 tokenizer
(10dd5110…, vocab 71,260), for
chrishayuk/v11-tinystories-115m-base.
Identity: 2115d6aeff3428e217ef2903a8030facd511dcb00183e9fc3faaf49d01038767
(chuk-datasets… See the full description on the dataset page: https://huggingface.co/datasets/chrishayuk/v11-cells-midtrain-corpus.gemini_orpo_dpo_ptbrmerged-corpus
Merged Corpus
Welcome to this repository.This dataset is japanese corpus that includes wiki, wikibooks, wikiversity, cc100, and oscar2109.
Getting Started
If you want to use this, please run as follows.This process takes about 3 hours.
mkdir -p pretrain/input/
cd pretrain/input/
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/datasets/ce-lery/merged-corpus.git
cd merged-corpus
git lfs pull
bash merge_train.sh
leis_ordinarias_1988_2024valdoria-dpo-qwen35-dataset
Valdoria DPO dataset
Dataset de preferências conversacional para uso direto com trl.DPOTrainer.
Splits
train.jsonl: 1.889 pares
validation.jsonl: 236 pares
test.jsonl: 237 pares
Cada linha contém:
{
"id": "...",
"prompt": [{"role": "system", "content": "..."}, {"role": "user", "content": "..."}],
"chosen": [{"role": "assistant", "content": "..."}],
"rejected": [{"role": "assistant", "content": "..."}],
"metadata": {"task_type": "..."… See the full description on the dataset page: https://huggingface.co/datasets/celsowm/valdoria-dpo-qwen35-dataset.minecraft_qa_es
Minecraft Q&A (Spanish)
A Spanish, chat-formatted question/answer dataset about Minecraft. Each example is a short conversation with a single user question and a single assistant answer (plus a system prompt).
Data format
The dataset is provided as JSON Lines (.jsonl): one JSON object per line.
Each record has a single key:
messages: an array of chat messages, each with:
role: one of "system", "user", "assistant"
content: the message text
Typical structure:… See the full description on the dataset page: https://huggingface.co/datasets/CelesteLove/minecraft_qa_es.modelos_peticoesHTML was converted to Markdown (better for LLMs)
valdoria-sft-qwen35-dataset
Dataset Card — Valdoria SFT Pack v3.3.1
Dataset sintético em português para Supervised Fine-Tuning (SFT), baseado em um país fictício chamado República de Valdoria. Formato ChatML (messages), combinando os splits train + validation (2125 exemplos).
Descrição
Cobertura de tarefas:
factual_qa (358)
fantasy_boundary (534)
refusal (251)
unknown_canonical_field (252)
negative_case (200)
decision_making (191)
rule_application (108)
classification (131)… See the full description on the dataset page: https://huggingface.co/datasets/celsowm/valdoria-sft-qwen35-dataset.verbetes_vademecum_direitoOriginal source: https://vademecumbrasil.com.br/dicionario-juridico/
corpus-ja-11b
corpus-ja-11b
Welcome to this repository.This dataset is japanese corpus that includes wiki, wikibooks, wikiversity, and cc100.
Getting Started
If you want to use this, please run as follows.This process takes about 3 hours.
mkdir -p pretrain/input/
cd pretrain/input/
GIT_LFS_SKIP_SMUDGE=1 git clone https://huggingface.co/datasets/ce-lery/corpus-ja-11b.git
cd corpus-ja-11b
git lfs pull
bash merge_train.sh
celestial-spiritual-ai-dataset
CELESTIAL Spiritual AI Training Dataset
Overview
Comprehensive spiritual AI training dataset with 190+ conversations covering all 16 CELESTIAL spiritual systems.
Features Covered
Vedic Astrology & Kundli Generation
Numerology & Life Path Analysis
Divine AI Personas (Krishna, Ganesha, Shiva, Devi, Hanuman, Saraswati)
Vastu Shastra & Palmistry
Spiritual Guidance & Meditation
And 11 more spiritual systems!
Quality Metrics
Overall Score: 94.9%… See the full description on the dataset page: https://huggingface.co/datasets/dp1812/celestial-spiritual-ai-dataset.evliya-celebi-seyahatname-ocr
Evliya Celebi Seyahatname OCR Corpus
OCR-derived text from seven volumes of Evliya Celebi's Seyahatname, packaged
for corpus exploration, language modeling, OCR-quality analysis, and historical
Ottoman Turkish / Turkish NLP work.
Configs
pages: one row per OCR page, with page numbers and OCR status.
documents: one row per available volume, with page text concatenated.
Coverage
Available books: 1, 3, 4, 6, 7, 9, 10.
Missing from the 1-10 sequence: 2, 5, 8.… See the full description on the dataset page: https://huggingface.co/datasets/fatihburakkaragoz/evliya-celebi-seyahatname-ocr.tw-celebrity-chat
Dataset Card for tw-celebrity-chat
本資料集是以臺灣公眾人物(藝人、團體、運動員、創作者等)相關之維基百科條目為種子(seed),由 LLM 合成的繁體中文「思考+回答」對話集,可用於訓練具備 reasoning 能力的繁中對話模型,特別是針對臺灣演藝/流行文化領域。
Dataset Details
Dataset Description
資料集為 reference-based 合成對話:先以維基百科上臺灣公眾人物相關段落作為 seed,再透過 LLM 生成單輪 QA,並在回答前加上 <think>...</think> 思考段落。每筆樣本同時保留:
conversations:human/gpt 兩輪結構(gpt 回答中含思考段落);
input / output / think:把上述拆解出來的單獨欄位,方便不同訓練框架使用;
seed:用來提示 LLM 的原始參考文本。
主題涵蓋臺灣男女歌手、團體(如 183Club、5566、SHE 等)、戲劇演員、棒球與 SBL… See the full description on the dataset page: https://huggingface.co/datasets/lianghsun/tw-celebrity-chat.
