CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HiTZ /BertaQA Dataset Card for BertaQA BertaQA is a trivia dataset comprising 4,756 multiple-choice trivia questions, with one single correct answer and 2 additional distractors. Crucially, questions are distributed between local and global topics. Whereas answering questions in the latter group requires general world knowledge, local questions require specific knowledge about the Basque Country and its culture. Additionally, questions are classified into eight categories, namely Basque and… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/BertaQA.tabularquestion-answering10K<n<100K1 likes623 downloads2y agoHugging Face02m-a-p /FineFineWeb-bert-seeddata FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus arXiv: Coming Soon Project Page: Coming Soon Blog: Coming Soon Data Statistics Domain (#tokens/#samples) Iteration 1 Tokens Iteration 2 Tokens Iteration 3 Tokens Total Tokens Iteration 1 Count Iteration 2 Count Iteration 3 Count Total Count aerospace 5.77B 261.63M 309.33M 6.34B 9100000 688505 611034 10399539 agronomy 13.08B 947.41M 229.04M 14.26B 15752828 2711790 649404 19114022 artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-bert-seeddata.texttext-classification1M<n<10M2 likes378 downloads2y agoHugging Face03BertilBraun /TinyPython TinyPython Tasks TinyPython is a synthetic Python dataset inspired by the idea behind TinyStories: if the data distribution is narrow, clean, and high quality, even very small language models can learn useful structure. Instead of broad repository code or competitive-programming solutions, TinyPython focuses on short natural-language programming tasks paired with complete, typed, standalone Python functions. The goal is to provide a compact instruction-to-code corpus for… See the full description on the dataset page: https://huggingface.co/datasets/BertilBraun/TinyPython.texttext-generation1M<n<10M1 likes93 downloads3mo agoHugging Face04BertilBraun /voice-light-tool-use-synthetic Voice Light Teacher-Led Tool-Use Synthetic This repository contains the current canonical synthetic source dataset for Voice Light's conversational tool-use fine-tuning. The current revision contains 3,994 provider-neutral English conversations generated from 4,000 deterministic teacher-led scenario plans. Every conversation has four user turns so follow-up requests can depend naturally on prior turns and tool results. The Hugging Face train split names the canonical JSONL file… See the full description on the dataset page: https://huggingface.co/datasets/BertilBraun/voice-light-tool-use-synthetic.texttext-generation1K<n<10K0 likes73 downloads2mo agoHugging Face05mmdjiji /bert-chinese-idiomsFor the detail, see github:mmdjiji/bert-chinese-idioms. preprocess.js is a Node.JS script to generate the data for training the language model. text10K<n<100K4 likes60 downloads4y agoHugging Face06toksuitebackup /bert-base-multilingual-cased-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model. The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl). tabular10M<n<100M0 likes58 downloads10mo agoHugging Face07ssurface /hallucination-bert-spans Hallucination BERT Span Dataset Flat, one-row-per-span dataset intended for span/token-classification (BIO-tagging style) hallucination detection over agent tool-calling traces, derived from the same judging pipeline as the reasoning-distillation set in this collection. File ds_bert_spans_full.jsonl — 11,942 rows. Already self-contained — no join needed. Each row is one hallucinated span: span (verbatim text), type (taxonomy label), avg_iou / exact / n_judges… See the full description on the dataset page: https://huggingface.co/datasets/ssurface/hallucination-bert-spans.tabular10K<n<100K0 likes56 downloads2mo agoHugging Face08Lo /clip-bert-data CLIP-BERT training data This data was used to train the CLIP-BERT model first described in this paper. The dataset is based on text and images from MS COCO, SBU Captions, Visual Genome QA and Conceptual Captions. The image features have been extracted using the CLIP model openai/clip-vit-base-patch32 available on Huggingface. text1M<n<10M1 likes54 downloads4y agoHugging Face09YesaOuO /TEKGEN-Sentence-BERTtext10K<n<100K0 likes35 downloads2y agoHugging Face10Idlebeing /Bert-senttext10K<n<100K0 likes34 downloads1y agoHugging Face11Linkhero2 /bertopic-conflictos-chile-v17-academic 🏆 BERTopic v17 - "The Academic" Research-Backed Solutions Applied: ClassTfidfTransformer: bm25_weighting + reduce_frequent_words ngram_range=(1, 2) - entity fusions handle compounds PartOfSpeech for valid Spanish noun phrases Chained: POS → KeyBERT → MMR(diversity=1.0) top_n_words=5 Results: K Natural: 51 topics Total docs: 3,268 textn<1K0 likes34 downloads9mo agoHugging Face12Dudeman523 /Bert-Rustbusters-Relevance Laser Cleaning Query Relevance Dataset Overview This dataset was created for training text classification models to identify customer queries relevant to laser cleaning services. It contains a comprehensive collection of text examples labeled for relevance to laser cleaning, enabling automated triage of customer inquiries for laser cleaning businesses. Files The dataset is available in multiple formats: full_dataset.jsonl - Complete dataset in JSONL format… See the full description on the dataset page: https://huggingface.co/datasets/Dudeman523/Bert-Rustbusters-Relevance.textn<1K0 likes28 downloads2y agoHugging Face13AlexSham /DaNetQA_for_BERTRussianNLP/russian_super_glue/DaNetQA train and val, with token [SEP] text1K<n<10K0 likes24 downloads2y agoHugging Face14AlexSham /TERRA_for_BERTlabel 1 = entailnment label 0 = not entailnment Russian super glue TERRA train and val text1K<n<10K0 likes24 downloads2y agoHugging Face15as-cle-bert /saccaromyces-cerevisiae-basetextn<1K1 likes22 downloads2y agoHugging Face16JWBickel /KJVPericopeTopics_bertopicWith bertopic (https://maartengr.github.io/BERTopic/), I ran a dataset of pericopes which covers the entire Bible. The pericope became the topic, and under each heading, 3 verses were selected as representative. A useful feature of this is that the representative verses are guaranteed to come from the section of Scripture that's connected with the pericope, which gives much better quality than if they were, for instance, semantically chosen from a vector database of the entire Bible. texttext-classification1K<n<10K0 likes19 downloads3y agoHugging Face17as-cle-bert /genetics-arxiv-wiki Dataset Card for Dataset Name Small genetics-related text dataset based on 23200 ArXiv abstact records and 111 Wikipedia pages. Dataset Details Dataset Description Dataset was produced using the python scripts you will find in this GitHub repository. It represents a collection of genetics-related text data taken from ArXiv abstracts dataset and Wikipedia. Dataset holds a total of 23311 text records, 23200 of which belonging to categories q-bio.BM, q-bio.GN… See the full description on the dataset page: https://huggingface.co/datasets/as-cle-bert/genetics-arxiv-wiki.text10K<n<100K2 likes19 downloads3y agoHugging Face18Linkhero2 /bertopic-conflictos-chile-v18-pure 🏆 BERTopic v18 - "The Pure" Hypothesis Entity fusions created artificial high-frequency tokens that dominated n-grams, causing redundancy. Changes from v17: REMOVED all entity fusions (text is now natural) RESTORED ngram_range=(1, 4) to capture full phrases PartOfSpeech + MMR(0.8) representation chain top_n_words=7 Expected Result: Keywords like "tribunal ambiental antofagasta" instead of "tribunalambiental antofagasta" Results: K… See the full description on the dataset page: https://huggingface.co/datasets/Linkhero2/bertopic-conflictos-chile-v18-pure.textn<1K0 likes18 downloads9mo agoHugging Face19bert-ka /Turkish-Municipality-Instruction-Tuning-Datasettexttext-generationn<1K0 likes18 downloads1mo agoHugging Face20BertilBraun /competency-extraction-dpo-v2 Competency Extraction DPO Dataset Overview The Competency Extraction DPO Dataset is a specialized dataset designed to extract competency profiles from scientific publications. The dataset focuses on identifying and structuring competencies found within the abstracts of academic papers. It contains a total of 6,179 samples and provides valuable insights into automated competency extraction. Dataset Format The dataset follows the standard DPO dataset format with… See the full description on the dataset page: https://huggingface.co/datasets/BertilBraun/competency-extraction-dpo-v2.textfeature-extraction1K<n<10K0 likes16 downloads2y agoHugging Face21seabelial /category-training-berttext10K<n<100K0 likes16 downloads1y agoHugging Face22rausch /prepared_context_4_experiments_old_train_bert-base-uncasedtextn<1K0 likes11 downloads2y agoHugging Face23BertilBraun /competency-extraction-dpo Competency Extraction DPO Dataset Overview The Competency Extraction DPO Dataset is a specialized dataset designed to extract competency profiles from scientific publications. The dataset focuses on identifying and structuring competencies found within the abstracts of academic papers. It contains a total of 6,179 samples and provides valuable insights into automated competency extraction. Dataset Format The dataset follows the standard DPO dataset format with… See the full description on the dataset page: https://huggingface.co/datasets/BertilBraun/competency-extraction-dpo.textfeature-extraction1K<n<10K0 likes11 downloads2y agoHugging Face24open-llm-leaderboard /Jimmy19991222__llama-3-8b-instruct-gapo-v2-bert_f1-beta10-gamma0.3-lr1.0e-6-scale-log-detailsgated Dataset Card for Evaluation run of Jimmy19991222/llama-3-8b-instruct-gapo-v2-bert_f1-beta10-gamma0.3-lr1.0e-6-scale-log Dataset automatically created during the evaluation run of model Jimmy19991222/llama-3-8b-instruct-gapo-v2-bert_f1-beta10-gamma0.3-lr1.0e-6-scale-log The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Jimmy19991222__llama-3-8b-instruct-gapo-v2-bert_f1-beta10-gamma0.3-lr1.0e-6-scale-log-details.tabular10K<n<100K0 likes8 downloads2y agoHugging Face25Linkhero2 /bertopic-conflictos-chile-v21-masterpiece 🏆 BERTopic v21 - THE MASTERPIECE Mejoras sobre v20: SmartDeduplicator (0.7): Más permisivo, conserva "Argentina", "transfronterizo" Safety Net: Garantiza mínimo 6 keywords por topic Stopwords SOTA: Sin meses, años, nombres personales Visualizaciones: Barcharts y mapas HTML automáticos Métricas: K Natural: 50 Total docs: 3,268 Timestamp: 2026-01-13 17:07:41.278922 Archivos: bertopic_topic_explanations_v21.xlsx: Keywords por K… See the full description on the dataset page: https://huggingface.co/datasets/Linkhero2/bertopic-conflictos-chile-v21-masterpiece.tabularn<1K0 likes8 downloads8mo agoHugging Face26n0w0f /MatText_Robocrys_bert_scaleupn<1K0 likes8 downloads7mo agoHugging Face27JGlang /BERTsequenceClasstext1K<n<10K0 likes7 downloads3y agoHugging Face28quipohealth /berttext1K<n<10K0 likes6 downloads3y agoHugging Face29ICKD /sst5-berttabular10K<n<100K0 likes6 downloads2y agoHugging Face30open-llm-leaderboard /Jimmy19991222__llama-3-8b-instruct-gapo-v2-bert-f1-beta10-gamma0.3-lr1.0e-6-1minus-rerun-detailsgated Dataset Card for Evaluation run of Jimmy19991222/llama-3-8b-instruct-gapo-v2-bert-f1-beta10-gamma0.3-lr1.0e-6-1minus-rerun Dataset automatically created during the evaluation run of model Jimmy19991222/llama-3-8b-instruct-gapo-v2-bert-f1-beta10-gamma0.3-lr1.0e-6-1minus-rerun The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task. The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Jimmy19991222__llama-3-8b-instruct-gapo-v2-bert-f1-beta10-gamma0.3-lr1.0e-6-1minus-rerun-details.tabular10K<n<100K0 likes6 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.