llm-generated
human-vs-llm-generated-text-detection-distilbertllm-detect-ai-generated-text-bertLLM_generated_text_detectorhuman-vs-llm-generated-text-detection-distilbertdistilbert-llm-generated-essay-classifierhuman-vs-llm-generated-text-detection-distilberthuman-vs-llm-generated-text-detection-distilbertLLM-Generated-Summaries-tdidf-baseline
ivypanda-llm-generated-essays
AI-Generated Essays Dataset
This dataset contains AI-generated academic essays created using the models:
Mistral 7B Instruct v0.2 (Q5_K_M quantized)
Temperature: 0.7
Max tokens: 4096
Top-p: 0.9 (default)
Top-k: 40 (default)
Repeat penalty: 1.1 (default)
Context window: 32768 tokens
Llama 3 13B Instruct v0.1 (Q5_K_M quantized)
Temperature: 0.7
Max tokens: 4096
Top-p: 0.9 (default)
Top-k: 40 (default)
Repeat penalty: 1.1 (default)
Context window: 8192 tokens
DeepSeek-V3.2
API… See the full description on the dataset page: https://huggingface.co/datasets/artfultom/ivypanda-llm-generated-essays.llm-generated-essayLLM-generated-emoji-descriptions
Emoji Metadata Dataset
Overview
The LLM Emoji Dataset is a comprehensive collection of enriched semantic descriptions for emojis, generated using Meta AI's Llama-3-8B model. This dataset aims to provide semantic context for each emoji, enhancing their usability in various NLP applications, especially those requiring semantic search. The LLM Emoji Dataset was used to build a multilingual search engine for emojies, which you can interact with using this online Streamlit… See the full description on the dataset page: https://huggingface.co/datasets/badrex/LLM-generated-emoji-descriptions.38k-zh-yue-translation-llm-generatedThis dataset consists of Chinese (Simplified) to Cantonese translation pairs generated using large language models (LLMs) and translated by Google Palm2. The dataset aims to provide a collection of translated sentences for training and evaluating Chinese (Simplified) to Cantonese translation models.
The dataset creation process involved two main steps:
LLM Sentence Generation: ChatGPT, a powerful LLM, was utilized to generate 10 sentences for each term pair. These sentences were generated in… See the full description on the dataset page: https://huggingface.co/datasets/hon9kon9ize/38k-zh-yue-translation-llm-generated.llm-generated-textsThis dataset is composed of parallel texts, generated by LLMs and written by human authors. The methodology for constructing the is based on the [1] and uses prompts from [2].
The dataset comprises of powerful LLMs generations, 21'000 in total. Used LLMs:
GPT4 Turbo 2024-04-09: https://platform.openai.com/docs/models/gpt-4-turbo-and-gpt-4
GPT4 Omni: https://openai.com/index/hello-gpt-4o
Claude 3 Opus: https://www.anthropic.com/news/claude-3-family
Llama3 70B: https://llama.meta.com/llama3/… See the full description on the dataset page: https://huggingface.co/datasets/artnitolog/llm-generated-texts.royal_society_corpus_LLM_generated_metadata
llm-detect-ai-generated-text-datasetteLLM-Agents_Generated_Consensus_Statementsrepro-black-box-detection-llm-generated-text-gjsrepro-telescope-improving-zero-shot-detection-of-llm-generated-content-by-measuring-token-repetirepro-from-llm-generated-conjectures-to-lean-formalizations-automated-polynomial-inequality-prov
