datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mbert-mrtydi-corpusmedhie-tokenized-dataset-mBERTmbert-mrtydisindhi-gold-corpus-mlm-tokenized-mbertmBERT-large-claim-agent-v10
mBERT-large Claim Agent — Training Dataset v10
Sentence-level binary classification data used to fine-tune mBERT-large for claim detection in
medical-aesthetics promotional material. A claim is a statement of product efficacy, safety,
indication, or market performance that requires substantiation against an approved claims matrix.
Schema
column
type
description
id
int
Unique row id, 0..4717
sentence
str
The extracted sentence
label
int
1 = claim, 0… See the full description on the dataset page: https://huggingface.co/datasets/Inabia-AI/mBERT-large-claim-agent-v10.linguistic_representation_mBERTThis dataset obtains genealogical and typological information for the 104 languages used for pre-training of the language model multilingual BERT (Devlin et al., 2019).
The genealogical information covers the language family and the genus for each language.
For typological description of the pre-training languages, 36 features from WALS (Dryer & Haspelmath, 2013) were used.
The information provided here can be used, among other things, to investigate how the pre-training corpus is structured… See the full description on the dataset page: https://huggingface.co/datasets/MayaGalvez/linguistic_representation_mBERT.tokenized_sefaria_dataset_mbertkgz_dataset_chunked_medium_mbertmedhie-tokenized-dataset-roots_ar_b50-mBERTmedhie-tokenized-dataset-sefaria-mBERTmedhie-tokenized-dataset-roots_ar-mBERTmbert-base-cased-NER-NL-legislation-refs-data
Dataset description
This dataset was created for fine-tuning the model mbert-base-cased-NER-NL-legislation-refs and consists of 512 token long examples which each contain one or more legislation references. These examples were created from a weakly labelled corpus of Dutch case law which was scraped from Linked Data Overheid, pre-tokenized and labelled (biluo_tags_from_offsets) through spaCy and further tokenized through applying Hugging Face's AutoTokenizer.from_pretrained() for… See the full description on the dataset page: https://huggingface.co/datasets/romjansen/mbert-base-cased-NER-NL-legislation-refs-data.medhie-tokenized-dataset-roots_ar_f50-mBERTtokenized_roots_dataset_mbertmbert-circuit-outputsresults_mbert_basembert_reward_dpoed_model_ckpt20_c4_lowq_200m2b_subsample20m_grpo_prompt
