datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
eai-taxonomy-math-w-fm-classify-behaviors
🧮 EAI Taxonomy Math w/ Behavioral Classifications (10K Sample)
A 10,000 document sample from EssentialAI/eai-taxonomy-math-w-fm enhanced with 4 behavioral reasoning classifications using GPT-4.1-mini.
Behavioral Classifications
Structured behavioral analysis following the approach from cognitive-behaviors:
backtracking_json: Identifies reasoning that backtracks or revisits earlier steps
backward_chaining_json: Detects goal-oriented reasoning working backwards… See the full description on the dataset page: https://huggingface.co/datasets/nlile/eai-taxonomy-math-w-fm-classify-behaviors.classify-embeddings-sf-baseball-marinasmyers_briggs_text_classifytask517_emo_classify_emotion_of_dialogue
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task517_emo_classify_emotion_of_dialogue
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task517_emo_classify_emotion_of_dialogue.ecqa_explanation_classifydomain_field_classifying_cleaned
domain_field_classifyingcleaned
This is a derived, leakage-cleaned copy of the public
wantRP/domain_field_classifying dataset. It contains the two shougang
variants:
shougang-10000/
shougang-5415/
For each variant, all.json, train.json, val.json, and test.json are
provided.
Cleaning provenance
Rows whose content fingerprint matched any row in the three splits of the
shougang-native-release-b SFT release were removed from every split. The
fingerprint is the exact… See the full description on the dataset page: https://huggingface.co/datasets/xianyu9n/domain_field_classifying_cleaned.dwd-hf-classify-1
DWD HF CLASSIFY 1
Dataset gathered from a 20m anntena from slovakia, recieving dwd (militarry weather info)
Curently super small, recieving tooks a lot of time at 50 baud (50bps)
It was made by using a known list of types and qualities and training a small ml to help together with some rules hardcoded to classify the type and quality.
If there are issues please report them and i will try improoving the ml and rules.
TYPE
= garbage / broken line
= metadata… See the full description on the dataset page: https://huggingface.co/datasets/simonko912/dwd-hf-classify-1.filtered_sky_code_8k_math_10k_rubric_evidence_classifyreddit-suicidal-classify-kagglemultilang-classify-dataset-02
Dataset Card for Multilingual Language Detection
Dataset Details
This dataset is a comprehensive resource for multilingual text classification, specifically designed for language identification. It contains over 100,000 text samples from 36 different languages, sourced from various public datasets and meticulously cleaned for machine learning applications.
The primary goal of this dataset is to train models that can accurately predict the language of a given text snippet.… See the full description on the dataset page: https://huggingface.co/datasets/minhleduc/multilang-classify-dataset-02.OR-bench-classify
Operations Research Question Classification Dataset
This dataset contains operations research (OR) questions classified into different categories with detailed reasoning.
Categories
The dataset includes questions classified into the following categories:
Linear Programming (LP): Both objective function and constraints are linear, decision variables are continuous
Integer Programming (IP): Both objective function and constraints are linear, all decision variables must be… See the full description on the dataset page: https://huggingface.co/datasets/yilingwang/OR-bench-classify.classify_alignment_faking_human_labelsecqa_classify_94translation-classifytask1618_cc_alligned_classify_tel_eng
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1618_cc_alligned_classify_tel_eng
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1618_cc_alligned_classify_tel_eng.reddit-suicidal-classify-trainsecurity-classify-data
数据介绍
选择了clouditera/security-paper-datasets 这个数据集的前20w行数据,用AI给数据打标签,选择了187270的与网络安全相关的数据。
在普通的数据集中,也选择18w的数据,数据来源为zhwiki与enwiki,各选择9w条。
用bert-base-chinese微调了模型,地址为
llama2_classifying_and_explainning_v5
Dataset Card for "llama2_classifying_and_explainning_v5"
More Information needed
vdo_format_classifytask375_classify_type_of_sentence_in_debate
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task375_classify_type_of_sentence_in_debate
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task375_classify_type_of_sentence_in_debate.ecqa_classify_5abstract-discipline-classifyrlv_cdip_classify_qamyers_briggs_blog_classifyllama2_classifying_and_explainning
Dataset Card for "llama2_classifying_and_explainning"
More Information needed
llama2_classifying_and_explainning_v3
Dataset Card for "llama2_classifying_and_explainning_v3"
More Information needed
global_classify_v2
Dataset Card for "global_classify_v2"
More Information needed
myers_briggs_description_classify
myers brigs self-identity descriptions
synthetically generated using:
multiple models for variety of responses (llama3.2-1b, llama3.3-70b, nous-hermes, gpt4-o-mini, mistral, phi3.5, and lexi-llama)
randomization of type for better model response in concurrent inference (although sorted before upload, so should be shuffled at minimum before training)
parsing for redundancy and then consolidated
a collection of descriptive statements that should describe one personality type over… See the full description on the dataset page: https://huggingface.co/datasets/jtatman/myers_briggs_description_classify.classify-dataset-pocVulnerability-Classify-TP-FP
