CoolFace
12 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01wytro /Know-Your-Sourcestabulartext-generation10M<n<100M0 likes1.2k downloads1mo agoHugging Face02alwaysgood /financial-english-source-corpus Financial English Source Corpus This dataset is a filtered, fuzzy-deduplicated English source-text corpus for financial-domain language-model training and translation-data generation. This version preserves the final pre-split source rows. Derived 1280-token split versions are available separately: financial-english-source-corpus-qwen35-1280 financial-english-source-corpus-gemma4-e2b-1280 Dataset Rows below are uploaded train rows before source-length splitting.… See the full description on the dataset page: https://huggingface.co/datasets/alwaysgood/financial-english-source-corpus.tabulartext-generation1M<n<10M0 likes547 downloads13d agoHugging Face03NuBerea /source-classificationsgated NuBerea Source Gold Set Curated source-critical classifications for the Hebrew Bible, New Testament, and Septuagint — the classical concerns of source criticism (documentary strata in the Old Testament, corpus structure in the New Testament, translation traditions in the Septuagint) expressed as structured, verse-level data, together with statistical validation summaries and characteristic-vocabulary ("hallmark") term lists. This dataset is part of the NuBerea curated corpus… See the full description on the dataset page: https://huggingface.co/datasets/NuBerea/source-classifications.tabulartext-generation10K<n<100K0 likes430 downloads2mo agoHugging Face04lapa-llm /classifier_source Dataset Card for Lapa High Quality Pretraining Dataset Dataset Description Dataset Summary This dataset is a random sample of both https://huggingface.co/datasets/lapa-llm/pretraining-lower-quality and https://huggingface.co/datasets/lapa-llm/pretraining-high-quality to transfer classifiers from English language to Ukrainian.It was used to transfer the following models from this collection https://huggingface.co/collections/lapa-llm/lapa-v012-pretraining:… See the full description on the dataset page: https://huggingface.co/datasets/lapa-llm/classifier_source.tabulartext-generation1M<n<10M0 likes356 downloads10mo agoHugging Face05alwaysgood /financial-english-source-corpus-qwen35-1280 Financial English Source Corpus Qwen35 1280 This dataset is a filtered, fuzzy-deduplicated English source-text corpus for financial-domain language-model training and translation-data generation. The uploaded Parquet files are already prepared with the 1280-token source split used by the downstream training pipeline. This split version is derived from the pre-split Financial English Source Corpus by applying sentence-boundary splitting with the qwen3.5 tokenizer.… See the full description on the dataset page: https://huggingface.co/datasets/alwaysgood/financial-english-source-corpus-qwen35-1280.tabulartext-generation1M<n<10M0 likes265 downloads3mo agoHugging Face06eewer /swerebench-traces-raw-source-verification-enhanced-20260617 SWE-rebench Raw Source Verification Enhanced 20260617 This is a private raw source dataset for building refined mini-swe-agent SFT datasets. It is intentionally not tokenized and intentionally preserves source data plus metadata for downstream filtering, masking, weighting, and audit. Do not treat every row as a clean endpoint solve. Download The full dataset directory is uploaded as a single compressed archive: hf download… See the full description on the dataset page: https://huggingface.co/datasets/eewer/swerebench-traces-raw-source-verification-enhanced-20260617.tabulartext-generationn<1K0 likes76 downloads3mo agoHugging Face07AmanPriyanshu /tool-reasoning-sft-RESEARCH-rlvr-env-retrieval-source Tool-Reasoning SFT — RLVR Retrieval Source Trajectories 156,381 multi-turn agentic retrieval trajectories across three document corpora, in a strict reasoning + tool-call format with validated FSM transitions. Each trajectory records a model searching a corpus, opening documents, and citing relevant passages to answer a question. Author: Aman Priyanshu Source Environments Trajectories were collected against three RLVR retrieval environments from the FORMAT: Search -… See the full description on the dataset page: https://huggingface.co/datasets/AmanPriyanshu/tool-reasoning-sft-RESEARCH-rlvr-env-retrieval-source.tabulartext-generation100K<n<1M0 likes72 downloads6mo agoHugging Face08KuanKuanKuan /falsifyrl-source FalsifyRL Reward-Hacking Falsification FalsifyRL is a synthetic, executable benchmark for identifying and repairing proxy-reward failures in embodied multi-agent reinforcement learning. Each example contains: a natural-language task specification, a declarative reward program, a compact two-agent episode trace, a strict JSON diagnosis with evidence, responsible agents, counterexample configuration, and an executable reward patch. Dataset design The dataset… See the full description on the dataset page: https://huggingface.co/datasets/KuanKuanKuan/falsifyrl-source.tabulartext-classification1K<n<10K0 likes58 downloads2mo agoHugging Face09essobi /dclm-crossover-source DCLM Cross-Over Source Subset of DCLM-Baseline selected for synthetic augmentation with format-aware prompt routing. Selection Picked every 3th shard (9313 of 27938 shards) Word count filter: 50-8000 Per-site cap: 10,000 Format detection: skip prompts that duplicate native document format Stats Metric Value Source docs scanned 54,947,699 Selected 54,017,165 Total words 44,119,449,000 Avg words/doc 816 Length filtered 930,534… See the full description on the dataset page: https://huggingface.co/datasets/essobi/dclm-crossover-source.tabulartext-generation100M<n<1B1 likes56 downloads5mo agoHugging Face10haowu89 /math-ai-bench-sources-latest math-ai-bench-sources-latest This dataset is an updated aggregated multi-trajectory benchmark built from the latest parallelthinking_benchmark files under /scratch/haowu/datasets/datasets/parallelthinking_benchmark_latest. It follows the same high-level format as haowu89/math-ai-bench-sources, but it is a newer version with: updated benchmark composition updated model set aligned question coverage across all included models Included Models Qwen2.5-1.5B-Instruct… See the full description on the dataset page: https://huggingface.co/datasets/haowu89/math-ai-bench-sources-latest.tabulartext-generation1K<n<10K0 likes29 downloads6mo agoHugging Face11haowu89 /math-ai-bench-sources math-ai-bench-sources This dataset contains math_ai_parallelthinking_benchmark.jsonl, built for comparing multiple reasoning trajectories across models on the same set of questions. File math_ai_parallelthinking_benchmark.jsonl Data Construction The benchmark is built from subsets of zechen-nlp/math-ai-bench (including gpqa) and distilled with the following 3 models: Qwen_Qwen2.5-1.5B-Instruct Qwen_Qwen3-4B-Nothinking Qwen_Qwen3-4B-Thinking For each model… See the full description on the dataset page: https://huggingface.co/datasets/haowu89/math-ai-bench-sources.tabulartext-generation1K<n<10K0 likes17 downloads7mo agoHugging Face12kaushik-systalyze /customer-transcript-source Customer Transcript Source Curated customer-support and transcript-analytics prompts mapped to a single fixed "analyze this transcript -> compact JSON" prompt, for benchmarking batched offline LLM inference on realistic workloads. Motivation and intended use This dataset provides a realistic transcript-analytics workload for batched offline-inference experiments: throughput benchmarking and predicted-vs-observed throughput validation. Rows carry token accounting… See the full description on the dataset page: https://huggingface.co/datasets/kaushik-systalyze/customer-transcript-source.tabulartext-generation1K<n<10K0 likes12 downloads3mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.