datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
enrun-emails-token-classificationtask388_torque_token_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task388_torque_token_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task388_torque_token_classification.token-classification-checkpoint-downloadsglobalise_NER_token_classification_dataset
Dataset Card for Dataset Name
The globalise_NER_token_classification dataset is a fine-grained dataset for the training of token-classification NER models on Dutch East-India Company texts (17th to 18th century).
Dataset Details
Dataset Description
The dataset provides 15 fine-grained labels detailing activities and people of the Dutch East-India Company (VOC), and can be used to train NER token-classification models for the
period 17th-18th century and the… See the full description on the dataset page: https://huggingface.co/datasets/globalise/globalise_NER_token_classification_dataset.classification_token_propagandatoken_classification_19thJulyinappropriateness-token-classification-binarized-multi-reftoken_classification_ratnakar_1300token-classification-tutorial
Token Classification Tutorial Dataset
Dataset Description
This dataset contains predicted probabilities for token classification used in the cleanlab tutorial: Token Classification.
The dataset demonstrates how to use cleanlab to identify and correct label issues in token classification datasets, such as Named Entity Recognition (NER) tasks where each token in a sequence is assigned a class label.
Dataset Summary
Task: Token classification / Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/Cleanlab/token-classification-tutorial.token-classification-brand
Dataset Card for "token-classification-brand"
More Information needed
claims_token_classificationtoken_classification_datasetinappropriateness-token-classification-binarizedgec-token-classificationautotrain-data-test-token-classification
AutoTrain Dataset for project: test-token-classification
Dataset Description
This dataset has been automatically processed by AutoTrain for project test-token-classification.
Languages
The BCP-47 code for the dataset's language is en.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"tokens": [
"I",
"will",
"be",
"traveling",
"to",
"Tokyo",
"next"… See the full description on the dataset page: https://huggingface.co/datasets/PhaniManda/autotrain-data-test-token-classification.bert_token_classificationautotrain-data-demo-on-token-classification
AutoTrain Dataset for project: demo-on-token-classification
Dataset Description
This dataset has been automatically processed by AutoTrain for project demo-on-token-classification.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"tokens": [
"I",
"will",
"be",
"traveling",
"to",
"Tokyo",
"next"… See the full description on the dataset page: https://huggingface.co/datasets/PhaniManda/autotrain-data-demo-on-token-classification.inappropriateness-token-classification-multi-reftoken-classification-japanese-search-local-cuisine料理を検索するための質問文と、質問文に含まれる検索検索用キーワードの情報を持ったデータセットです
固有表現の種類は以下の4つです。
AREA: 都道府県/地方
TYPE: 種類
SZN: 季節
INGR: 食材
GitHub
untokenized_dataset_list.ipynb(データセットの作成に使ったノートブック)
このデータセットを使った言語モデルのファインチューニングと、ファインチューニングした言語モデルを使ったアプリのコードもこのリポジトリにあります
詳細情報
Qiita
mock_token_classification_datasetcord-v2-token-classificationbert_token_classification_augmentedinappropriateness-token-classificationsangapac-token-classificationautotrain-data-token-classification
AutoTrain Dataset for project: token-classification
Dataset Description
This dataset has been automatically processed by AutoTrain for project token-classification.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"tokens": [
"Pd",
"has",
"been",
"regarded",
"as",
"one",
"of",
"the"… See the full description on the dataset page: https://huggingface.co/datasets/QNN/autotrain-data-token-classification.FAKEINVEST_TOKENCLASSIFICATIONflan_combined_task388_torque_token_classificationtoken_classification_datasets
