datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BertaQA
Dataset Card for BertaQA
BertaQA is a trivia dataset comprising 4,756 multiple-choice trivia questions, with one single correct answer and 2 additional distractors. Crucially, questions are distributed between local and global topics. Whereas answering questions in the latter group requires general world knowledge, local questions require specific knowledge about the Basque Country and its culture. Additionally, questions are classified into eight categories, namely Basque and… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/BertaQA.FineFineWeb-bert-seeddata
FineFineWeb: A Comprehensive Study on Fine-Grained Domain Web Corpus
arXiv: Coming Soon
Project Page: Coming Soon
Blog: Coming Soon
Data Statistics
Domain (#tokens/#samples)
Iteration 1 Tokens
Iteration 2 Tokens
Iteration 3 Tokens
Total Tokens
Iteration 1 Count
Iteration 2 Count
Iteration 3 Count
Total Count
aerospace
5.77B
261.63M
309.33M
6.34B
9100000
688505
611034
10399539
agronomy
13.08B
947.41M
229.04M
14.26B
15752828
2711790
649404
19114022
artistic… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/FineFineWeb-bert-seeddata.TinyPython
TinyPython Tasks
TinyPython is a synthetic Python dataset inspired by the idea behind TinyStories: if the data distribution is narrow, clean, and high quality, even very small language models can learn useful structure.
Instead of broad repository code or competitive-programming solutions, TinyPython focuses on short natural-language programming tasks paired with complete, typed, standalone Python functions.
The goal is to provide a compact instruction-to-code corpus for… See the full description on the dataset page: https://huggingface.co/datasets/BertilBraun/TinyPython.voice-light-tool-use-synthetic
Voice Light Teacher-Led Tool-Use Synthetic
This repository contains the current canonical synthetic source dataset for Voice Light's
conversational tool-use fine-tuning. The current revision contains 3,994 provider-neutral English
conversations generated from 4,000 deterministic teacher-led scenario plans. Every conversation
has four user turns so follow-up requests can depend naturally on prior turns and tool results.
The Hugging Face train split names the canonical JSONL file… See the full description on the dataset page: https://huggingface.co/datasets/BertilBraun/voice-light-tool-use-synthetic.bert-chinese-idiomsFor the detail, see github:mmdjiji/bert-chinese-idioms.
preprocess.js is a Node.JS script to generate the data for training the language model.
bert-base-multilingual-cased-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model.
The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).
hallucination-bert-spans
Hallucination BERT Span Dataset
Flat, one-row-per-span dataset intended for span/token-classification
(BIO-tagging style) hallucination detection over agent tool-calling traces,
derived from the same judging pipeline as the reasoning-distillation set in
this collection.
File
ds_bert_spans_full.jsonl — 11,942 rows. Already self-contained — no join
needed. Each row is one hallucinated span: span (verbatim text), type
(taxonomy label), avg_iou / exact / n_judges… See the full description on the dataset page: https://huggingface.co/datasets/ssurface/hallucination-bert-spans.clip-bert-data
CLIP-BERT training data
This data was used to train the CLIP-BERT model first described in this paper.
The dataset is based on text and images from MS COCO, SBU Captions, Visual Genome QA and Conceptual Captions.
The image features have been extracted using the CLIP model openai/clip-vit-base-patch32 available on Huggingface.
TEKGEN-Sentence-BERTBert-sentbertopic-conflictos-chile-v17-academic
🏆 BERTopic v17 - "The Academic"
Research-Backed Solutions Applied:
ClassTfidfTransformer: bm25_weighting + reduce_frequent_words
ngram_range=(1, 2) - entity fusions handle compounds
PartOfSpeech for valid Spanish noun phrases
Chained: POS → KeyBERT → MMR(diversity=1.0)
top_n_words=5
Results:
K Natural: 51 topics
Total docs: 3,268
Bert-Rustbusters-Relevance
Laser Cleaning Query Relevance Dataset
Overview
This dataset was created for training text classification models to identify customer queries relevant to laser cleaning services. It contains a comprehensive collection of text examples labeled for relevance to laser cleaning, enabling automated triage of customer inquiries for laser cleaning businesses.
Files
The dataset is available in multiple formats:
full_dataset.jsonl - Complete dataset in JSONL format… See the full description on the dataset page: https://huggingface.co/datasets/Dudeman523/Bert-Rustbusters-Relevance.DaNetQA_for_BERTRussianNLP/russian_super_glue/DaNetQA train and val, with token [SEP]
TERRA_for_BERTlabel 1 = entailnment
label 0 = not entailnment
Russian super glue TERRA train and val
saccaromyces-cerevisiae-baseKJVPericopeTopics_bertopicWith bertopic (https://maartengr.github.io/BERTopic/), I ran a dataset of pericopes which covers the entire Bible.
The pericope became the topic, and under each heading, 3 verses were selected as representative.
A useful feature of this is that the representative verses are guaranteed to come from the section of Scripture
that's connected with the pericope, which gives much better quality than if they were, for instance, semantically
chosen from a vector database of the entire Bible.
genetics-arxiv-wiki
Dataset Card for Dataset Name
Small genetics-related text dataset based on 23200 ArXiv abstact records and 111 Wikipedia pages.
Dataset Details
Dataset Description
Dataset was produced using the python scripts you will find in this GitHub repository.
It represents a collection of genetics-related text data taken from ArXiv abstracts dataset and Wikipedia.
Dataset holds a total of 23311 text records, 23200 of which belonging to categories q-bio.BM, q-bio.GN… See the full description on the dataset page: https://huggingface.co/datasets/as-cle-bert/genetics-arxiv-wiki.bertopic-conflictos-chile-v18-pure
🏆 BERTopic v18 - "The Pure"
Hypothesis
Entity fusions created artificial high-frequency tokens that dominated n-grams, causing redundancy.
Changes from v17:
REMOVED all entity fusions (text is now natural)
RESTORED ngram_range=(1, 4) to capture full phrases
PartOfSpeech + MMR(0.8) representation chain
top_n_words=7
Expected Result:
Keywords like "tribunal ambiental antofagasta" instead of "tribunalambiental antofagasta"
Results:
K… See the full description on the dataset page: https://huggingface.co/datasets/Linkhero2/bertopic-conflictos-chile-v18-pure.Turkish-Municipality-Instruction-Tuning-Datasetcompetency-extraction-dpo-v2
Competency Extraction DPO Dataset
Overview
The Competency Extraction DPO Dataset is a specialized dataset designed to extract competency profiles from scientific publications. The dataset focuses on identifying and structuring competencies found within the abstracts of academic papers. It contains a total of 6,179 samples and provides valuable insights into automated competency extraction.
Dataset Format
The dataset follows the standard DPO dataset format with… See the full description on the dataset page: https://huggingface.co/datasets/BertilBraun/competency-extraction-dpo-v2.category-training-bertprepared_context_4_experiments_old_train_bert-base-uncasedcompetency-extraction-dpo
Competency Extraction DPO Dataset
Overview
The Competency Extraction DPO Dataset is a specialized dataset designed to extract competency profiles from scientific publications. The dataset focuses on identifying and structuring competencies found within the abstracts of academic papers. It contains a total of 6,179 samples and provides valuable insights into automated competency extraction.
Dataset Format
The dataset follows the standard DPO dataset format with… See the full description on the dataset page: https://huggingface.co/datasets/BertilBraun/competency-extraction-dpo.Jimmy19991222__llama-3-8b-instruct-gapo-v2-bert_f1-beta10-gamma0.3-lr1.0e-6-scale-log-details
Dataset Card for Evaluation run of Jimmy19991222/llama-3-8b-instruct-gapo-v2-bert_f1-beta10-gamma0.3-lr1.0e-6-scale-log
Dataset automatically created during the evaluation run of model Jimmy19991222/llama-3-8b-instruct-gapo-v2-bert_f1-beta10-gamma0.3-lr1.0e-6-scale-log
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Jimmy19991222__llama-3-8b-instruct-gapo-v2-bert_f1-beta10-gamma0.3-lr1.0e-6-scale-log-details.bertopic-conflictos-chile-v21-masterpiece
🏆 BERTopic v21 - THE MASTERPIECE
Mejoras sobre v20:
SmartDeduplicator (0.7): Más permisivo, conserva "Argentina", "transfronterizo"
Safety Net: Garantiza mínimo 6 keywords por topic
Stopwords SOTA: Sin meses, años, nombres personales
Visualizaciones: Barcharts y mapas HTML automáticos
Métricas:
K Natural: 50
Total docs: 3,268
Timestamp: 2026-01-13 17:07:41.278922
Archivos:
bertopic_topic_explanations_v21.xlsx: Keywords por K… See the full description on the dataset page: https://huggingface.co/datasets/Linkhero2/bertopic-conflictos-chile-v21-masterpiece.MatText_Robocrys_bert_scaleupBERTsequenceClassbertsst5-bertJimmy19991222__llama-3-8b-instruct-gapo-v2-bert-f1-beta10-gamma0.3-lr1.0e-6-1minus-rerun-details
Dataset Card for Evaluation run of Jimmy19991222/llama-3-8b-instruct-gapo-v2-bert-f1-beta10-gamma0.3-lr1.0e-6-1minus-rerun
Dataset automatically created during the evaluation run of model Jimmy19991222/llama-3-8b-instruct-gapo-v2-bert-f1-beta10-gamma0.3-lr1.0e-6-1minus-rerun
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Jimmy19991222__llama-3-8b-instruct-gapo-v2-bert-f1-beta10-gamma0.3-lr1.0e-6-1minus-rerun-details.
