datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
split-text-quickmt-train.zh-enspilit text (sentence)
iemocap_audio_text_splitted
Dataset Card for "iemocap_audio_text_splitted"
More Information needed
vi-text_corpus-dantri.com.vn-splittedtext_and_concat_image_hf_version_epoch_1_with_prefix_with_exist_split_fixed_best_of_16_CoTundl_text_split
Dataset Card for "undl_text_split"
More Information needed
isogram-ai-text-detection-splits
Isogram AI Text Detection Permissive Splits
This dataset contains train/validation/test splits for binary AI-generated text detection.
It is built from sources whose dataset-level licenses were checked as permissive or
public-domain-compatible.
Schema
text: essay text.
label: 0 for human-written text, 1 for AI-generated text.
source_dataset: upstream dataset identifier.
source_detail: source label retained from the upstream data.
source_license: row-level… See the full description on the dataset page: https://huggingface.co/datasets/sinatras/isogram-ai-text-detection-splits.BLIP3o-Pretrain-Long-Caption-Text-Splits_only_thought_text_hf_version_epoch_1_with_prefix_with_exist_split_fixed_best_of_16_imageswikipedia_20220301.simple_sentence_split_text_has_at_least_5_wordstext-splitter-alpacahttps://huggingface.co/datasets/mhenrichsen/context-aware-splits-english
HTML_CSS_CodeDataSet_100k_text_only_splitpt_text_split_1haitian-tags-and-text-generated-splitDay_if_sentient_beings_SPLITED_BY_CAT_IM_SIGN_DEPTH_TEXT_47
_train_collect_cot_only_thought_text_hf_version_epoch_1_with_prefix_with_exist_split_fixedDay_if_sentient_beings_SPLITED_BY_ZHONGLI_IM_SIGN_DEPTH_TEXT_47
split_text_datasetDay_if_sentient_beings_SPLITED_BY_XIAO_IM_SIGN_DEPTH_TEXT_47
tasariv_splits_transcibed_filtered-tags-and-text-generated
