CoolFace
21 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01l3lab /ntp-mathlib miniCTX: Neural Theorem Proving with (Long-)Contexts Lean 4 tactic prediction examples extracted from Mathlib. These examples have not been formatted for instruction tuning (including data splits). Please see l3lab/ntp-mathlib-instruct-* for datasets with instruction tuning examples. Version Generated using ntptoolkit's ntp-training-data. It used the following config for ntp-training-data: { "repo": "https://github.com/leanprover-community/mathlib4", "commit":… See the full description on the dataset page: https://huggingface.co/datasets/l3lab/ntp-mathlib.text100K<n<1M2 likes161 downloads2y agoHugging Face02l3lab /ntp-mathlib-instruct-context miniCTX: Neural Theorem Proving with (Long-)Contexts Lean 4 tactic prediction examples extracted from Mathlib. Examples contain: prompt: instruction, preceding file content, proof state instruction, proof state completion: tactic The file content has been truncated to 1024 tokens. Version Generated using ntptoolkit's ntp-training-data and instruction_tuning.py. It used the following config for ntp-training-data: { "repo":… See the full description on the dataset page: https://huggingface.co/datasets/l3lab/ntp-mathlib-instruct-context.text100K<n<1M1 likes116 downloads2y agoHugging Face03l3lab /ntp-mathlib-instruct-st miniCTX: Neural Theorem Proving with (Long-)Contexts Lean 4 tactic prediction examples extracted from Mathlib. Examples contain: prompt: instruction, proof state completion: tactic Version Generated using ntptoolkit's ntp-training-data and instruction_tuning.py. It used the following config for ntp-training-data: { "repo": "https://github.com/leanprover-community/mathlib4", "commit": "cf8e23a62939ed7cc530fbb68e83539730f32f86", "lean":… See the full description on the dataset page: https://huggingface.co/datasets/l3lab/ntp-mathlib-instruct-st.text100K<n<1M0 likes89 downloads2y agoHugging Face04l3lab /ntp-mathlib-instruct-context-fullproof miniCTX: Neural Theorem Proving with (Long-)Contexts Lean 4 full proof generation examples extracted from Mathlib. Examples contain: prompt: instruction, preceding file content completion: proof The file content has been truncated to 1024 tokens. Version Generated using ntptoolkit's ntp-training-data and instruction_tuning.py. It used the following config for ntp-training-data: { "repo": "https://github.com/leanprover-community/mathlib4", "commit":… See the full description on the dataset page: https://huggingface.co/datasets/l3lab/ntp-mathlib-instruct-context-fullproof.text100K<n<1M1 likes57 downloads2y agoHugging Face05ntphuc149 /ViBidLQA_v1 ViBidLQA: Vietnamese Bidding Legal Question Answering Dataset Summary ViBidLQA is a synthesized Vietnamese legal question-answering dataset built from the Vietnamese Bidding Law. It was created to address the scarcity of large-scale annotated datasets for legal AI in Vietnamese — a low-resource language setting. The dataset contains 3,013 QA pairs generated automatically using a large language model (Gemini) and verified by two domain experts. You may also like the new… See the full description on the dataset page: https://huggingface.co/datasets/ntphuc149/ViBidLQA_v1.textquestion-answering1K<n<10K1 likes56 downloads5mo agoHugging Face06ntphuc149 /ViBidLQAgated 📘 ViBidLQA — Vietnamese Bidding Law QA Dataset ViBidLQA is a high-quality Vietnamese legal QA dataset specifically curated from the Vietnamese Bidding Law (No. 22/2023/QH15) and its Implementation Decree (No. 24/2024/NĐ-CP). This dataset was developed to support both extractive and abstractive QA tasks in the legal domain, and serves as a robust benchmark for evaluating Vietnamese Legal QA systems. 🔍 Motivation Although existing datasets like ALQAC have been… See the full description on the dataset page: https://huggingface.co/datasets/ntphuc149/ViBidLQA.textquestion-answering1K<n<10K4 likes28 downloads5mo agoHugging Face07ntphiep /vit5-tst-data-sino-vietnamesetext100K<n<1M1 likes27 downloads1y agoHugging Face08ardauzunoglu /dclm-fineweb-ntp-eval-3k DCLM and FineWeb-Edu NTP Evaluation Retained 3K This repository contains two 3,000-document subsets retained by the next-token prediction (NTP) ordering analysis: dclm: 3,000 unique documents selected from dclm_new_eval50k.jsonl, originally sampled from mlfoundations/dclm-baseline-1.0. fineweb_edu: 3,000 unique documents selected from fineweb_edu_new_eval50k.jsonl, originally sampled from HuggingFaceFW/fineweb-edu (sample-10BT). Load either configuration with: from datasets… See the full description on the dataset page: https://huggingface.co/datasets/ardauzunoglu/dclm-fineweb-ntp-eval-3k.tabular1K<n<10K0 likes27 downloads2mo agoHugging Face09ntphuc149 /ViLegalMCQgated ViLegalMCQ ViLegalMCQ is a Vietnamese legal Multiple Choice Question Answering (MCQ) dataset released alongside the ViLegalLM suite. It is synthetically generated from the ALQAC legal corpus using Qwen3-8B with human filtering, providing training data for context-based legal MCQ tasks. Paper: ViLegalLM: Language Models for Vietnamese Legal Text — Read paper Resources: GitHub | ViLegalBERT | ViLegalQwen2.5-1.5B-Base | ViLegalQwen3-1.7B-Base Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/ntphuc149/ViLegalMCQ.textquestion-answering10K<n<100K2 likes25 downloads3mo agoHugging Face10ntphuc149 /ViSpanExtractQA 📘 ViSpanExtractQA — Vietnamese Span-based QA Benchmark ViSpanExtractQA is a consolidated Vietnamese QA dataset designed for span-based extractive question answering tasks. It aggregates and harmonizes multiple high-quality resources to create a diverse, multilingual, and robust benchmark for Vietnamese QA systems. 🔍 Motivation While several QA datasets exist in Vietnamese, most are limited in size, scope, or consistency. ViSpanExtractQA aims to bridge this gap by… See the full description on the dataset page: https://huggingface.co/datasets/ntphuc149/ViSpanExtractQA.text100K<n<1M0 likes20 downloads1y agoHugging Face11ntphiep /vit5-tst-data-coarsetext100K<n<1M0 likes19 downloads1y agoHugging Face12cosmos1030 /qwen3-4b-s70pct-ntpkd-only-ot3-rollouts-longtextn<1K0 likes15 downloads1mo agoHugging Face13cosmos1030 /qwen3-4b-s70pct-ntpkdopkd-ot3-rollouts-longtextn<1K0 likes14 downloads1mo agoHugging Face14Hao-Chen /ntp_temples1text1K<n<10K0 likes13 downloads2y agoHugging Face15ntphuc149 /ViLegalTextgated ViLegalTexts: A 16GB Vietnamese Legal Pre-training Corpus This repository contains the pre-training corpus used to train ViLegalLM. The corpus was crawled from four publicly available Vietnamese legal repositories. For full details, please refer to the paper: Read paper 📋 Overview Property Value Language Vietnamese Domain Legal Corpus size 16 GB Format .zip containing multiple .txt files Sources 4 public Vietnamese legal repositories… See the full description on the dataset page: https://huggingface.co/datasets/ntphuc149/ViLegalText.text100K<n<1M4 likes11 downloads3mo agoHugging Face16ntphuc /bimnext_math_vi_1textn<1K0 likes10 downloads2y agoHugging Face17ntphiep /vit5-tst-data-formaltext100K<n<1M0 likes10 downloads1y agoHugging Face18cosmos1030 /qwen3-1.7b-s70pct-ntpkd-only-ot3-rollouts-longtextn<1K0 likes10 downloads1mo agoHugging Face19ntphuc149 /ViLegalTFgated ViLegalTF ViLegalTF is a Vietnamese legal True/False Question Answering (TF) dataset released alongside the ViLegalLM suite. It is synthetically generated from the ALQAC legal corpus using Qwen3-8B with human filtering, providing training data for context-based legal true/false judgment tasks. Paper: ViLegalLM: Language Models for Vietnamese Legal Text — Read paper Resources: GitHub | ViLegalBERT | ViLegalQwen2.5-1.5B-Base | ViLegalQwen3-1.7B-Base Dataset Summary… See the full description on the dataset page: https://huggingface.co/datasets/ntphuc149/ViLegalTF.textquestion-answering10K<n<100K2 likes8 downloads3mo agoHugging Face20ntphiep /vit5-tst-data-casualtext100K<n<1M0 likes7 downloads1y agoHugging Face21cosmos1030 /qwen3-1.7b-s70pct-ntpkdopkd-ot3-rollouts-longtextn<1K0 likes6 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.