BERT
Datasets
All datasets matching “BERT”multilingual-pl-bertAttribution: Wikipedia.org
bert-128-grouped
Dataset Card for "bert-128-grouped"
More Information needed
conceptual-captions-12m-webdataset-bertsSimpleSafetyTestsmc4-es-sampled50 million documents in Spanish extracted from mC4 applying perplexity sampling via mc4-sampling: "https://huggingface.co/datasets/bertin-project/mc4-sampling". Please, refer to BERTIN Project. The original dataset is the Multlingual Colossal, Cleaned version of Common Crawl's web crawl corpus (mC4), based on the Common Crawl dataset: "https://commoncrawl.org", and processed by AllenAI.BERTJudge-Dataset
BERTJudge-Dataset
Dataset Description
BERTJudge-Dataset is the training dataset used for developing BERTJudge models, as introduced in the paper BERT-as-a-Judge: A Robust Alternative to Lexical Methods for Efficient Reference-Based LLM Evaluation. It comprises question–candidate–reference triplets generated by 36 recent open-weight, instruction-tuned models across 7 established tasks, and synthetically annotated using nvidia/Llama-3_3-Nemotron-Super-49B-v1_5.
The dataset… See the full description on the dataset page: https://huggingface.co/datasets/artefactory/BERTJudge-Dataset.
