datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
BertaQA
Dataset Card for BertaQA
BertaQA is a trivia dataset comprising 4,756 multiple-choice trivia questions, with one single correct answer and 2 additional distractors. Crucially, questions are distributed between local and global topics. Whereas answering questions in the latter group requires general world knowledge, local questions require specific knowledge about the Basque Country and its culture. Additionally, questions are classified into eight categories, namely Basque and… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/BertaQA.fineweb2-ro-bert
FineWeb2-Ro-BERT
FineWeb2-Ro-BERT is a large-scale pretraining dataset in the Romanian language. The data is derived from FineWeb2 and annotated using a bert architecture for signals such as educational quality or topic. More details can be found here.
Key Features
Massive Scale: Contains approximately 54.1M rows (documents or sequences), providing comprehensive linguistic coverage for training robust Romanian embeddings and encoders.
Usage
You… See the full description on the dataset page: https://huggingface.co/datasets/surogate/fineweb2-ro-bert.wm_benchmark_test_uploadfineweb2-ro-bert
FineWeb2-Ro-BERT
FineWeb2-Ro-BERT is a large-scale pretraining dataset in the Romanian language. The data is derived from FineWeb2 and annotated using a bert architecture for signals such as educational quality or topic. More details can be found here.
Key Features
Massive Scale: Contains approximately 54.1M rows (documents or sequences), providing comprehensive linguistic coverage for training robust Romanian embeddings and encoders.
Usage
You can load… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-Ro/fineweb2-ro-bert.magic-bert-dataset
Magic-BERT Binary File Classification Dataset
This dataset contains tokenized binary file samples for training MIME type classification models.
Each sample is a 64KB chunk from the beginning of a file, tokenized using a byte-level BPE tokenizer.
Splits
Split
Samples
Train
37,111
Validation
4,591
Test
4,748
Features
Each sample contains:
Feature
Type
Description
blake2b
string
Content hash (unique sample ID)
mime_type
string
MIME… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/magic-bert-dataset.aio-passages-bpr-bert-base-japanese-v3
Dataset Card for llm-book/aio-passages-bert-base-japanese-v3-bpr
書籍『大規模言語モデル入門』で使用する、「AI王」コンペティションのパッセージデータセットに BPR によるパッセージの埋め込みを適用したデータセットです。
llm-book/aio-passages のデータセットに対して、llm-book/bert-base-japanese-v3-bpr-passage-encoder によるパッセージのバイナリベクトルが embeddings フィールドに追加されています。
Licence
本データセットで利用している Wikipedia のコンテンツは、クリエイティブ・コモンズ表示・継承ライセンス 3.0 (CC BY-SA 3.0) および GNU 自由文書ライセンス (GFDL) の下に配布されているものです。
job_listing_german_cleaned_bert
Dataset Card for "job_listing_german_cleaned_bert"
More Information needed
jawiki-paragraphs-unsup-simcse-bert-base-japanese-v3
Dataset Card for "jawiki-paragraphs-unsup-simcse-bert-base-japanese-v3"
More Information needed
bert-base-multilingual-cased-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model.
The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl).
hallucination-bert-spans
Hallucination BERT Span Dataset
Flat, one-row-per-span dataset intended for span/token-classification
(BIO-tagging style) hallucination detection over agent tool-calling traces,
derived from the same judging pipeline as the reasoning-distillation set in
this collection.
File
ds_bert_spans_full.jsonl — 11,942 rows. Already self-contained — no join
needed. Each row is one hallucinated span: span (verbatim text), type
(taxonomy label), avg_iou / exact / n_judges… See the full description on the dataset page: https://huggingface.co/datasets/ssurface/hallucination-bert-spans.marc
MARC: Metaphor Abstraction and Reasoning Corpus
What This Is
MARC identifies puzzles where figurative language and visual examples are genuinely complementary: the model fails given examples alone, fails given the metaphor alone, but succeeds when both are presented together. We call this the MARC property. The corpus provides 78 MARC-verified puzzles with 1,230 domain-diverse figurative descriptions and complete behavioral trial data for three language models.
Suppose… See the full description on the dataset page: https://huggingface.co/datasets/bertybaums/marc.Elite-Chess-Games
Elite Chess Data
Processed and added version of this dataset
marc2
MARC2: Metaphor Abstraction and Reasoning Corpus v2
MARC2 extends the MARC-from-LARC methodology to the ARC-AGI2 dataset. It provides a corpus of figurative language puzzles where metaphorical descriptions help AI models solve abstract reasoning tasks they cannot solve from examples alone.
The MARC Property
A task has the MARC property (for a given model) when:
Examples alone fail — the model cannot solve the task from input/output examples
Figurative description… See the full description on the dataset page: https://huggingface.co/datasets/bertybaums/marc2.market-positivity-bert-tokenizedtest_split_with_embeddings_bert_base_portuguese
Dataset Card for "test_split_with_embeddings_bert_base_portuguese"
More Information needed
clincal_bert_large_code_mappingyelp-bert-scaledBERT_SBERT_embeddings_SAEtrec-bert-scaledtest-datasetft_bert_benchmark14D4T-chunked-bert-FTil_gymThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.1",
"robot_type": null,
"total_episodes": 30,
"total_frames": 1120,
"total_tasks": 1,
"total_videos": 60,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:30"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Bertsin/il_gym.bertopic-conflictos-chile-v22-complete
🏆 BERTopic v22 - THE COMPLETE MASTERPIECE
Mejoras sobre v21:
MODELO GUARDADO: SafeTensors para reutilización
TODAS LAS VISUALIZACIONES: DataMapPlot, Hierarchy, Heatmap, etc.
MÉTRICAS MATEMÁTICAS: Coherence (UMass, NPMI), Diversity, Silhouette
TOPICS OVER TIME: Análisis temporal
HIERARCHICAL TOPICS: Árbol jerárquico
GET_DOCUMENT_INFO: Metadata completa
REDUCE_OUTLIERS: Reducción inteligente
Métricas Matemáticas:
Topic Diversity: 0.6995305164319249… See the full description on the dataset page: https://huggingface.co/datasets/Linkhero2/bertopic-conflictos-chile-v22-complete.retrieval_verification_bm25_bert
Dataset Card for "retrieval_verification_bm25_bert"
More Information needed
Bertsafety-qa-bert-dataset
Safety QA Dataset
Dataset Description
There are two dataset that is publicaly available dataset from Mine Safety and Health Administration (MSHA). The 'seed_annotated_data.csv' dataset contains seed annotated data where the answer to the safety related questions are annotated in the accident narratives for initial training. The main 'training data.csv' data is used during the active learning (AL) process for question answering tasks in occupational safety and health… See the full description on the dataset page: https://huggingface.co/datasets/adanish91/safety-qa-bert-dataset.Evaluation_google-bert-bert-large-uncasedBERT_SBERT_PALM_embeddingsSAEMildsum-BERT-below-threshold
