CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01HiTZ /BertaQA Dataset Card for BertaQA BertaQA is a trivia dataset comprising 4,756 multiple-choice trivia questions, with one single correct answer and 2 additional distractors. Crucially, questions are distributed between local and global topics. Whereas answering questions in the latter group requires general world knowledge, local questions require specific knowledge about the Basque Country and its culture. Additionally, questions are classified into eight categories, namely Basque and… See the full description on the dataset page: https://huggingface.co/datasets/HiTZ/BertaQA.tabularquestion-answering10K<n<100K1 likes613 downloads2y agoHugging Face02surogate /fineweb2-ro-bert FineWeb2-Ro-BERT FineWeb2-Ro-BERT is a large-scale pretraining dataset in the Romanian language. The data is derived from FineWeb2 and annotated using a bert architecture for signals such as educational quality or topic. More details can be found here. Key Features Massive Scale: Contains approximately 54.1M rows (documents or sequences), providing comprehensive linguistic coverage for training robust Romanian embeddings and encoders. Usage You… See the full description on the dataset page: https://huggingface.co/datasets/surogate/fineweb2-ro-bert.tabular10M<n<100M0 likes314 downloads27d agoHugging Face03BertG666 /wm_benchmark_test_uploadimage1K<n<10K2 likes228 downloads1y agoHugging Face04OpenLLM-Ro /fineweb2-ro-bert FineWeb2-Ro-BERT FineWeb2-Ro-BERT is a large-scale pretraining dataset in the Romanian language. The data is derived from FineWeb2 and annotated using a bert architecture for signals such as educational quality or topic. More details can be found here. Key Features Massive Scale: Contains approximately 54.1M rows (documents or sequences), providing comprehensive linguistic coverage for training robust Romanian embeddings and encoders. Usage You can load… See the full description on the dataset page: https://huggingface.co/datasets/OpenLLM-Ro/fineweb2-ro-bert.tabular10M<n<100M2 likes164 downloads10mo agoHugging Face05mjbommar /magic-bert-dataset Magic-BERT Binary File Classification Dataset This dataset contains tokenized binary file samples for training MIME type classification models. Each sample is a 64KB chunk from the beginning of a file, tokenized using a byte-level BPE tokenizer. Splits Split Samples Train 37,111 Validation 4,591 Test 4,748 Features Each sample contains: Feature Type Description blake2b string Content hash (unique sample ID) mime_type string MIME… See the full description on the dataset page: https://huggingface.co/datasets/mjbommar/magic-bert-dataset.tabulartext-classification10K<n<100K0 likes135 downloads10mo agoHugging Face06llm-book /aio-passages-bpr-bert-base-japanese-v3 Dataset Card for llm-book/aio-passages-bert-base-japanese-v3-bpr 書籍『大規模言語モデル入門』で使用する、「AI王」コンペティションのパッセージデータセットに BPR によるパッセージの埋め込みを適用したデータセットです。 llm-book/aio-passages のデータセットに対して、llm-book/bert-base-japanese-v3-bpr-passage-encoder によるパッセージのバイナリベクトルが embeddings フィールドに追加されています。 Licence 本データセットで利用している Wikipedia のコンテンツは、クリエイティブ・コモンズ表示・継承ライセンス 3.0 (CC BY-SA 3.0) および GNU 自由文書ライセンス (GFDL) の下に配布されているものです。 tabular1M<n<10M1 likes114 downloads3y agoHugging Face07serbog /job_listing_german_cleaned_bert Dataset Card for "job_listing_german_cleaned_bert" More Information needed tabular100K<n<1M1 likes105 downloads3y agoHugging Face08llm-book /jawiki-paragraphs-unsup-simcse-bert-base-japanese-v3 Dataset Card for "jawiki-paragraphs-unsup-simcse-bert-base-japanese-v3" More Information needed tabular1M<n<10M1 likes84 downloads3y agoHugging Face09toksuitebackup /bert-base-multilingual-cased-toksuite-detokenizedTraining data of the model detokenized in the exact order seen by the model. The training data is partitioned into 8 chunks (chunk-0 through chunk-7), based on the GPU rank that generated the data. Each chunk contains detokenized text files in JSON Lines format (.jsonl). tabular10M<n<100M0 likes58 downloads10mo agoHugging Face10ssurface /hallucination-bert-spans Hallucination BERT Span Dataset Flat, one-row-per-span dataset intended for span/token-classification (BIO-tagging style) hallucination detection over agent tool-calling traces, derived from the same judging pipeline as the reasoning-distillation set in this collection. File ds_bert_spans_full.jsonl — 11,942 rows. Already self-contained — no join needed. Each row is one hallucinated span: span (verbatim text), type (taxonomy label), avg_iou / exact / n_judges… See the full description on the dataset page: https://huggingface.co/datasets/ssurface/hallucination-bert-spans.tabular10K<n<100K0 likes56 downloads2mo agoHugging Face11bertybaums /marc MARC: Metaphor Abstraction and Reasoning Corpus What This Is MARC identifies puzzles where figurative language and visual examples are genuinely complementary: the model fails given examples alone, fails given the metaphor alone, but succeeds when both are presented together. We call this the MARC property. The corpus provides 78 MARC-verified puzzles with 1,230 domain-diverse figurative descriptions and complete behavioral trial data for three language models. Suppose… See the full description on the dataset page: https://huggingface.co/datasets/bertybaums/marc.tabularvisual-question-answering10K<n<100K0 likes47 downloads6mo agoHugging Face12Q-bert /Elite-Chess-Games Elite Chess Data Processed and added version of this dataset tabular100K<n<1M1 likes44 downloads2y agoHugging Face13bertybaums /marc2 MARC2: Metaphor Abstraction and Reasoning Corpus v2 MARC2 extends the MARC-from-LARC methodology to the ARC-AGI2 dataset. It provides a corpus of figurative language puzzles where metaphorical descriptions help AI models solve abstract reasoning tasks they cannot solve from examples alone. The MARC Property A task has the MARC property (for a given model) when: Examples alone fail — the model cannot solve the task from input/output examples Figurative description… See the full description on the dataset page: https://huggingface.co/datasets/bertybaums/marc2.tabulartext-generation10K<n<100K0 likes41 downloads5mo agoHugging Face14leetdavid /market-positivity-bert-tokenizedtabular1K<n<10K0 likes40 downloads5y agoHugging Face15iara-project /test_split_with_embeddings_bert_base_portuguese Dataset Card for "test_split_with_embeddings_bert_base_portuguese" More Information needed tabular100K<n<1M0 likes39 downloads3y agoHugging Face16mozay22 /clincal_bert_large_code_mappingtabular1K<n<10K1 likes35 downloads3y agoHugging Face17ICKD /yelp-bert-scaledtabular100K<n<1M0 likes33 downloads2y agoHugging Face18v1ctor10 /BERT_SBERT_embeddings_SAEtabular10K<n<100K0 likes30 downloads2y agoHugging Face19ICKD /trec-bert-scaledtabular1K<n<10K0 likes28 downloads2y agoHugging Face20Q-bert /test-datasettabular1K<n<10K0 likes26 downloads3y agoHugging Face21rachid16 /ft_bert_benchmark1tabular1K<n<10K0 likes26 downloads2y agoHugging Face22SOMIL366 /4D4T-chunked-bert-FTtabular10M<n<100M0 likes26 downloads4mo agoHugging Face23Bertsin /il_gymThis dataset was created using LeRobot. Dataset Structure meta/info.json: { "codebase_version": "v2.1", "robot_type": null, "total_episodes": 30, "total_frames": 1120, "total_tasks": 1, "total_videos": 60, "total_chunks": 1, "chunks_size": 1000, "fps": 10, "splits": { "train": "0:30" }, "data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet", "video_path":… See the full description on the dataset page: https://huggingface.co/datasets/Bertsin/il_gym.tabularrobotics1K<n<10K0 likes25 downloads11mo agoHugging Face24Linkhero2 /bertopic-conflictos-chile-v22-complete 🏆 BERTopic v22 - THE COMPLETE MASTERPIECE Mejoras sobre v21: MODELO GUARDADO: SafeTensors para reutilización TODAS LAS VISUALIZACIONES: DataMapPlot, Hierarchy, Heatmap, etc. MÉTRICAS MATEMÁTICAS: Coherence (UMass, NPMI), Diversity, Silhouette TOPICS OVER TIME: Análisis temporal HIERARCHICAL TOPICS: Árbol jerárquico GET_DOCUMENT_INFO: Metadata completa REDUCE_OUTLIERS: Reducción inteligente Métricas Matemáticas: Topic Diversity: 0.6995305164319249… See the full description on the dataset page: https://huggingface.co/datasets/Linkhero2/bertopic-conflictos-chile-v22-complete.tabularn<1K0 likes25 downloads8mo agoHugging Face25nikchar /retrieval_verification_bm25_bert Dataset Card for "retrieval_verification_bm25_bert" More Information needed tabular10K<n<100K0 likes24 downloads3y agoHugging Face26TechTrekIndia01 /Berttabular100K<n<1M0 likes24 downloads3y agoHugging Face27adanish91 /safety-qa-bert-dataset Safety QA Dataset Dataset Description There are two dataset that is publicaly available dataset from Mine Safety and Health Administration (MSHA). The 'seed_annotated_data.csv' dataset contains seed annotated data where the answer to the safety related questions are annotated in the accident narratives for initial training. The main 'training data.csv' data is used during the active learning (AL) process for question answering tasks in occupational safety and health… See the full description on the dataset page: https://huggingface.co/datasets/adanish91/safety-qa-bert-dataset.tabularquestion-answering1K<n<10K0 likes23 downloads11mo agoHugging Face28ainewtrend07 /Evaluation_google-bert-bert-large-uncasedtabular10K<n<100K0 likes22 downloads1y agoHugging Face29v1ctor10 /BERT_SBERT_PALM_embeddingsSAEtabular10K<n<100K0 likes21 downloads2y agoHugging Face30priyankrathore /Mildsum-BERT-below-thresholdtabularn<1K0 likes21 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.