datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
task066_timetravel_binary_consistency_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task066_timetravel_binary_consistency_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task066_timetravel_binary_consistency_classification.ko-voicephishing-binary-classificationtask1559_blimp_binary_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1559_blimp_binary_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1559_blimp_binary_classification.toxicity-multilingual-binary-classification-datasetThis dataset is a comprehensive collection designed to aid in the development of robust and nuanced models for identifying toxic language across multiple languages, while critically distinguishing it from expressions related to mental health, specifically depression. It synthesizes content from three existing public datasets (ToxiGen, TextDetox, and Mental Health - Depression) with a newly generated synthetic dataset (ToxiLLaMA). The creation process involved careful collection, extensive… See the full description on the dataset page: https://huggingface.co/datasets/malexandersalazar/toxicity-multilingual-binary-classification-dataset.task609_sbic_potentially_offense_binary_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task609_sbic_potentially_offense_binary_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task609_sbic_potentially_offense_binary_classification.task607_sbic_intentional_offense_binary_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task607_sbic_intentional_offense_binary_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task607_sbic_intentional_offense_binary_classification.toxicity-multilingual-binary-classification-datasetThis dataset is a comprehensive collection designed to aid in the development of robust and nuanced models for identifying toxic language across multiple languages, while critically distinguishing it from expressions related to mental health, specifically depression. It synthesizes content from three existing public datasets (ToxiGen, TextDetox, and Mental Health - Depression) with a newly generated synthetic dataset (ToxiLLaMA). The creation process involved careful collection, extensive… See the full description on the dataset page: https://huggingface.co/datasets/hsbharadwaj/toxicity-multilingual-binary-classification-dataset.skin-lesion-HM10000-binary-classificationt
DermAI Workshop: Balanced Binary Skin Lesion Dataset
Dataset Description
This dataset is a modified version of the HAM10000 ("Human Against Machine with 10000 training images") dataset. It has been prepared specifically for educational purposes for the "DermAI: Building an AI-Powered Skin Lesion Classifier" workshop at the UTD Biotech Club.
The original dataset contains 7 classes of skin lesions. This version has been processed into binary classes: 'Benign' and… See the full description on the dataset page: https://huggingface.co/datasets/preetsojitra/skin-lesion-HM10000-binary-classificationt.task1493_bengali_geopolitical_hate_speech_binary_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1493_bengali_geopolitical_hate_speech_binary_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1493_bengali_geopolitical_hate_speech_binary_classification.task608_sbic_sexual_offense_binary_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task608_sbic_sexual_offense_binary_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task608_sbic_sexual_offense_binary_classification.task1490_bengali_personal_hate_speech_binary_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1490_bengali_personal_hate_speech_binary_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1490_bengali_personal_hate_speech_binary_classification.brain-tumor-binary-classification-dataset
Dataset Card for "brain-tumor-binary-classification-dataset"
More Information needed
task1492_bengali_religious_hate_speech_binary_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1492_bengali_religious_hate_speech_binary_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1492_bengali_religious_hate_speech_binary_classification.mazes-binary-classificationtask1548_wiqa_binary_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1548_wiqa_binary_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1548_wiqa_binary_classification.mimic-cxr-binary-classification-cied-detection-2-5Ktask1491_bengali_political_hate_speech_binary_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1491_bengali_political_hate_speech_binary_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1491_bengali_political_hate_speech_binary_classification.task1560_blimp_binary_classification
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1560_blimp_binary_classification
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1560_blimp_binary_classification.CIA-Drug_Trafficking-Binary_Classification-Koreanultrafeedback-binary-classificationThis dataset is derived from argilla/ultrafeedback-binarized-preferences-cleaned using the followig processing:
import random
import pandas as pd
from datasets import Dataset, load_dataset
from sklearn.model_selection import GroupKFold
data_df = load_dataset("argilla/ultrafeedback-binarized-preferences-cleaned", split="train").to_pandas()
rng = random.Random(42)
rng = random.Random(42)
def get_assistant_text(messages):
t = ""for msg in messages:
if msg["role"] == "assistant":… See the full description on the dataset page: https://huggingface.co/datasets/rbiswasfc/ultrafeedback-binary-classification.flan_combined_task1490_bengali_personal_hate_speech_binary_classificationnci-binary-classificationflan_combined_task1491_bengali_political_hate_speech_binary_classificationskin-lesion-HM10000-binary-classification-4K-subset
DermAI Workshop: Balanced Binary Skin Lesion Dataset
Dataset Description
This dataset is a modified and balanced subset of the HAM10000 ("Human Against Machine with 10000 training images") dataset. It has been prepared specifically for educational purposes for the "DermAI: Building an AI-Powered Skin Lesion Classifier" workshop at the UTD Biotech Club.
The original dataset contains 7 classes of skin lesions. This version has been processed into two balanced, binary… See the full description on the dataset page: https://huggingface.co/datasets/preetsojitra/skin-lesion-HM10000-binary-classification-4K-subset.mazes-binary-classification-simplifiedhd-bert-voicephishing-binary-classification-ver4reward-bench-binary-classificationRepurposed allenai/reward-bench dataset for binary classification task, using the following script:
import random
import pandas as pd
from datasets import Dataset, load_dataset
from sklearn.model_selection import GroupKFold
data_df = load_dataset("allenai/reward-bench", split="raw").to_pandas()
rng = random.Random(43)
examples = []
for idx, row in data_df.iterrows():
if rng.random() > 0.5:
response_a = row["chosen"]
response_b = row["rejected"]
label = 0… See the full description on the dataset page: https://huggingface.co/datasets/rbiswasfc/reward-bench-binary-classification.reddit_binary_classification_completionko-voicephishing-binary-classification-ver2qwen3_0.6b-rlvr_task1548_wiqa_binary_classification
