datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
falcon_medquad_generated_texts_verify_claims_deepseek_balanced_3000llm-generated-textsThis dataset is composed of parallel texts, generated by LLMs and written by human authors. The methodology for constructing the is based on the [1] and uses prompts from [2].
The dataset comprises of powerful LLMs generations, 21'000 in total. Used LLMs:
GPT4 Turbo 2024-04-09: https://platform.openai.com/docs/models/gpt-4-turbo-and-gpt-4
GPT4 Omni: https://openai.com/index/hello-gpt-4o
Claude 3 Opus: https://www.anthropic.com/news/claude-3-family
Llama3 70B: https://llama.meta.com/llama3/… See the full description on the dataset page: https://huggingface.co/datasets/artnitolog/llm-generated-texts.generated-phi-format-text-3ai-generated-texts
Spanish DPO Preference Pairs for Detector Evasion
Preference pairs for DPO fine-tuning of Qwen/Qwen2.5-0.5B-Instruct against the Oculus multilingual AI text detector on Spanish academic abstracts. Repository id: pymlex/ai-generated-texts.
Dataset size
Statistic
Count
Train abstracts processed
8891
DPO pairs retained
6396
Pairs skipped by logit margin
2495
Empty paraphrase pairs
0
Logit margin threshold: absolute gap at least 1.… See the full description on the dataset page: https://huggingface.co/datasets/pymlex/ai-generated-texts.ai-generated-text-classification
Dataset Card for "ai-generated-text-classification"
More Information needed
generated-phi-format-texthuman-ai-generated-text
Dataset Card for human-ai-generated-text
This dataset has been created with distilabel.
Dataset Summary
This dataset contains a pipeline.yaml which can be used to reproduce the pipeline that generated it in distilabel using the distilabel CLI:
distilabel pipeline run --config "https://huggingface.co/datasets/ardavey/human-ai-generated-text/raw/main/pipeline.yaml"
or explore the configuration:
distilabel pipeline info --config… See the full description on the dataset page: https://huggingface.co/datasets/ardavey/human-ai-generated-text.chatbot-arena-ja-calm2-7b-chat-experimental_deduped_add_generated_textgenerated-phi-format-text-2la-speech-and-text-generated-countrydiffusion-generated-text
Diffusion-Generated Text
This dataset contains 16,791 question-response pairs generated by a diffusion language model. It is released to support research on diffusion-generated language and machine-generated text detection.
Dataset schema
Column
Type
Description
question
string
Input question or prompt.
dLLM_response
string
Response produced by the diffusion language model.
Loading
from datasets import load_dataset
dataset =… See the full description on the dataset page: https://huggingface.co/datasets/paoche11/diffusion-generated-text.Generated-texts-Llama-2-7b-hf-arc_challengevizuara_gpt_generated_textfalcon_wmt19_deen_generated_texts_verify_claims_deepseekfalcon_mmlu_generated_texts_verify_claims_deepseek_balanced_3000Generated-texts-Llama-2-7b-hf-reward_benchla-speech-and-text-generated-bakfalcon_sciq_generated_texts_verify_claims_deepseekfalcon_sciq_generated_texts_verify_claims_deepseek_balanced_3000falcon_samsum_generated_texts_verify_claims_deepseekfalcon_gsm8k_generated_texts_verify_claims_deepseekGenerated-texts-Mistral-7B-v0.1-arc_challengebayly-tags-and-text-generated-perufalcon_sciq_generated_textslibritts-r-tags-and-text-generatedfalcon_mmlu_generated_texts_verify_claims_deepseekfalcon_gsm8k_generated_texts_verify_claims_deepseek_balanced_3000falcon_coqa_generated_texts_verify_claims_deepseek_balanced_3000Generated-texts-Llama-2-7b-hf-arc_challengefalcon_medquad_generated_texts_verify_claims_deepseek
