dragonscale-ai/kniv-corpus-en
kniv-corpus-en A multi-domain English NLP corpus with four annotation layers: Named Entity Recognition (18 types), POS tagging (17 UPOS tags), dependency parsing, and dialog act classification (9 types). All data is commercially licensed (CC BY-SA 4.0 compatible). Built for training kniv multi-task NLP models that power the uniko cognitive memory system. Quick Start from datasets import load_dataset # Load the full corpus via HuggingFace ds =… See the full description on the dataset page: https://huggingface.co/datasets/dragonscale-ai/kniv-corpus-en.
Upload prepared/kniv-deberta-cascade/docred_dev.json with huggingface_hub
Upload prepared/kniv-deberta-cascade/docred_train.json with huggingface_hub
Upload prepared/kniv-deberta-cascade/nli_dev.json with huggingface_hub
Upload prepared/kniv-deberta-cascade/nli_train.json with huggingface_hub
Upload prepared/kniv-deberta-cascade/intent_dev.json with huggingface_hub
Upload prepared/kniv-deberta-cascade/intent_train.json with huggingface_hub
Upload prepared/kniv-deberta-cascade/squad_dev.json with huggingface_hub
Upload prepared/kniv-deberta-cascade/squad_train.json with huggingface_hub
Upload prepared/kniv-deberta-cascade/truecase_dev.json with huggingface_hub
Upload prepared/kniv-deberta-cascade/truecase_train.json with huggingface_hub
Upload prepared/kniv-deberta-cascade/punct_dev.json with huggingface_hub
Upload prepared/kniv-deberta-cascade/punct_train.json with huggingface_hub
Upload prepared/kniv-deberta-cascade/lemma_lookup.json with huggingface_hub
Upload prepared/kniv-deberta-cascade/lemma_morph_vocabs.json with huggingface_hub
Upload prepared/kniv-deberta-cascade/ud_test_extended.json with huggingface_hub
Upload prepared/kniv-deberta-cascade/ud_dev_extended.json with huggingface_hub
Upload prepared/kniv-deberta-cascade/ud_train_extended.json with huggingface_hub
Delete prepared/kniv-deberta-cascade/cls_swda_mrda_train.json with huggingface_hub
Upload prepared/kniv-deberta-cascade/cls_sgd_mwoz_train.json with huggingface_hub
Upload prepared/kniv-deberta-cascade/cls_balanced_train.json with huggingface_hub
Upload prepared/kniv-deberta-cascade/dep_silver_spacy.json with huggingface_hub
Upload prepared/kniv-deberta-cascade/srl_allennlp_silver.json with huggingface_hub
Upload folder using huggingface_hub
Upload prepared/kniv-deberta-cascade/srl_silver_dannashao.json with huggingface_hub
Upload folder using huggingface_hub
Add SRL training data (PropBank + MASC + QA-SRL + silver, 68K train)
Upload prepared/kniv-deberta-cascade/cls_swda_mrda_train.json with huggingface_hub
Upload prepared/kniv-deberta-cascade/srl_full_train.json with huggingface_hub
Upload prepared/kniv-deberta-cascade/ner_spanmarker_dev.json with huggingface_hub
Upload prepared/kniv-deberta-cascade/ner_spanmarker_train.json with huggingface_hub
Upload benchmarks/conll2003_test.json with huggingface_hub
Upload benchmarks/ontonotes5_test.json with huggingface_hub
Upload prepared/kniv-deberta-cascade/label_vocabs.json with huggingface_hub
Upload prepared/kniv-deberta-cascade/srl_test.json with huggingface_hub
Upload prepared/kniv-deberta-cascade/srl_dev.json with huggingface_hub
Upload prepared/kniv-deberta-cascade/srl_train.json with huggingface_hub
Upload prepared/kniv-deberta-cascade/posner_small_train.json with huggingface_hub
Upload prepared/kniv-deberta-cascade/ud_ner_train.json with huggingface_hub
Upload prepared/kniv-deberta-cascade/fewnerd_train.json with huggingface_hub
Upload prepared/kniv-deberta-cascade/label_vocabs.json with huggingface_hub
Upload prepared/kniv-deberta-cascade/ud_test.json with huggingface_hub
Upload prepared/kniv-deberta-cascade/ud_dev.json with huggingface_hub
Upload prepared/kniv-deberta-cascade/ud_train.json with huggingface_hub
Upload prepared/kniv-deberta-cascade/posner_dev.json with huggingface_hub
Upload prepared/kniv-deberta-cascade/posner_train.json with huggingface_hub
Upload prepared/kniv-deberta-cascade/ner_test.json with huggingface_hub
Upload prepared/kniv-deberta-cascade/ner_dev.json with huggingface_hub
Upload prepared/kniv-deberta-cascade/ner_train.json with huggingface_hub
Rewrite README: collection details, pipeline, gold filtering methodology
Add gold-filtered prepared training data (45K NER + UD)
