CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01Hkang /moda-general-capability-rollouts MODA General Capability Retention Rollouts This dataset contains the raw model generations and evaluation results for the MODA general-capability retention experiments. It covers 16 models, seven benchmarks, 260,592 prompt records, and 2,605,920 stored generations. The evaluation code is pinned to source commit 12ea99b2a57a354f2b7d6792f62a3d9313192fa7. Evaluation protocol Benchmarks: GSM8K, MMLU abstract_algebra, GPQA Diamond, BoolQ, HellaSwag, TruthfulQA, and… See the full description on the dataset page: https://huggingface.co/datasets/Hkang/moda-general-capability-rollouts.tabulartext-generation1M<n<10M0 likes579 downloads1mo agoHugging Face02typhoon-ai /typhoon-s-sovereign-capability-dataset Typhoon-S Training Assets Training and evaluation datasets for Section 3, Thai language models used in the Typhoon-S project. Datasets NitiBench (Legal Domain) nitibench_train_rl.parquet - RL training set (8,211 examples) nitibench_train_pretrain.parquet - Pretrain set (3,648 examples) nitibench_train_sft.parquet - SFT set (3,648 examples) nitibench_test.parquet - Test set (373 examples) (10% of https://huggingface.co/datasets/VISAI-AI/nitibench ccl split)… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/typhoon-s-sovereign-capability-dataset.texttext-generation10K<n<100K0 likes145 downloads8mo agoHugging Face03wannaphong /typhoon-s-sovereign-capability-dataset Typhoon-S Training Assets Training and evaluation datasets for Section 3, Thai language models used in the Typhoon-S project. Datasets NitiBench (Legal Domain) nitibench_train_rl.parquet - RL training set (8,211 examples) nitibench_train_pretrain.parquet - Pretrain set (3,648 examples) nitibench_train_sft.parquet - SFT set (3,648 examples) nitibench_test.parquet - Test set (373 examples) (10% of https://huggingface.co/datasets/VISAI-AI/nitibench ccl split)… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/typhoon-s-sovereign-capability-dataset.texttext-generation10K<n<100K0 likes67 downloads6mo agoHugging Face04SolidSnake123 /nanochat-brevo-capability-data-10x Nanochat Brevo Capability Pilot Brevo presents shuffled dependency records and asks for the complete recursive prerequisite closure in a valid leaf-first order. Training uses project-planning language; validation uses evidence synthesis; test uses build manifests. Eleven deterministic structural styles vary wording, layout, and record order. The latent graph generator and exact validator label every row. No language model generated or labeled the data. Alternative valid orders… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-brevo-capability-data-10x.tabulartext-generation100K<n<1M0 likes38 downloads2mo agoHugging Face05SolidSnake123 /nanochat-brevo-capability-data Nanochat Brevo Capability Pilot Brevo presents shuffled dependency records and asks for the complete recursive prerequisite closure in a valid leaf-first order. Training uses project-planning language; validation uses evidence synthesis; test uses build manifests. Six deterministic structural styles vary wording, layout, and record order. The latent graph generator and exact validator label every row. No language model generated or labeled the data. Alternative valid orders are… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-brevo-capability-data.tabulartext-generation10K<n<100K0 likes28 downloads2mo agoHugging Face06SolidSnake123 /nanochat-brevo-capability-data-v2 Nanochat Brevo Capability Pilot Brevo presents shuffled dependency records and asks for complete recursive prerequisite closures in valid leaf-first orders. Each compact training document reuses one graph for 4 worked questions, increasing answer supervision without repeating the graph. Training uses project-planning language; validation uses evidence synthesis; test uses build manifests. Eleven deterministic structural styles vary wording, layout, and record order. Per-world… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-brevo-capability-data-v2.tabulartext-generation10K<n<100K0 likes16 downloads2mo agoHugging Face07SolidSnake123 /nanochat-brevo-capability-v4-49k-20260714 Nanochat Brevo Capability Pilot Brevo presents shuffled dependency records and asks for complete recursive prerequisite closures in valid leaf-first orders. Each compact training document reuses one graph for 4 worked questions, increasing answer supervision without repeating the graph. Training balances project_plan, build_manifest language; validation uses held-out evidence synthesis; test uses build manifests. Eleven deterministic structural styles vary wording, layout, and… See the full description on the dataset page: https://huggingface.co/datasets/SolidSnake123/nanochat-brevo-capability-v4-49k-20260714.tabulartext-generation10K<n<100K0 likes15 downloads2mo agoHugging Face08ToddLLM /luanti-capability-eval Luanti Capability Evaluation Dataset Description Luanti Capability Evaluation Dataset for Luanti (Minetest) expertise fine-tuning. Dataset Information Size: 60 entries Format: Harmony format for LLM fine-tuning Source: Luanti ContentDB package collection Quality: Filtered and validated Luanti package metadata Usage from datasets import load_dataset dataset = load_dataset("ToddLLM/luanti-capability-eval") print(dataset) Schema Each… See the full description on the dataset page: https://huggingface.co/datasets/ToddLLM/luanti-capability-eval.texttext-generationn<1K0 likes10 downloads1y agoHugging Face09ToddLLM /luanti-capability-training Luanti Capability Training Dataset Description Luanti Capability Training Dataset for Luanti (Minetest) expertise fine-tuning. Dataset Information Size: 600 entries Format: Harmony format for LLM fine-tuning Source: Luanti ContentDB package collection Quality: Filtered and validated Luanti package metadata Usage from datasets import load_dataset dataset = load_dataset("ToddLLM/luanti-capability-training") print(dataset) Schema Each… See the full description on the dataset page: https://huggingface.co/datasets/ToddLLM/luanti-capability-training.texttext-generationn<1K0 likes8 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.