datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ivypanda-llm-generated-essays
AI-Generated Essays Dataset
This dataset contains AI-generated academic essays created using the models:
Mistral 7B Instruct v0.2 (Q5_K_M quantized)
Temperature: 0.7
Max tokens: 4096
Top-p: 0.9 (default)
Top-k: 40 (default)
Repeat penalty: 1.1 (default)
Context window: 32768 tokens
Llama 3 13B Instruct v0.1 (Q5_K_M quantized)
Temperature: 0.7
Max tokens: 4096
Top-p: 0.9 (default)
Top-k: 40 (default)
Repeat penalty: 1.1 (default)
Context window: 8192 tokens
DeepSeek-V3.2
API… See the full description on the dataset page: https://huggingface.co/datasets/artfultom/ivypanda-llm-generated-essays.llm-generated-essayLLM-generated-emoji-descriptions
Emoji Metadata Dataset
Overview
The LLM Emoji Dataset is a comprehensive collection of enriched semantic descriptions for emojis, generated using Meta AI's Llama-3-8B model. This dataset aims to provide semantic context for each emoji, enhancing their usability in various NLP applications, especially those requiring semantic search. The LLM Emoji Dataset was used to build a multilingual search engine for emojies, which you can interact with using this online Streamlit… See the full description on the dataset page: https://huggingface.co/datasets/badrex/LLM-generated-emoji-descriptions.38k-zh-yue-translation-llm-generatedThis dataset consists of Chinese (Simplified) to Cantonese translation pairs generated using large language models (LLMs) and translated by Google Palm2. The dataset aims to provide a collection of translated sentences for training and evaluating Chinese (Simplified) to Cantonese translation models.
The dataset creation process involved two main steps:
LLM Sentence Generation: ChatGPT, a powerful LLM, was utilized to generate 10 sentences for each term pair. These sentences were generated in… See the full description on the dataset page: https://huggingface.co/datasets/hon9kon9ize/38k-zh-yue-translation-llm-generated.llm-generated-textsThis dataset is composed of parallel texts, generated by LLMs and written by human authors. The methodology for constructing the is based on the [1] and uses prompts from [2].
The dataset comprises of powerful LLMs generations, 21'000 in total. Used LLMs:
GPT4 Turbo 2024-04-09: https://platform.openai.com/docs/models/gpt-4-turbo-and-gpt-4
GPT4 Omni: https://openai.com/index/hello-gpt-4o
Claude 3 Opus: https://www.anthropic.com/news/claude-3-family
Llama3 70B: https://llama.meta.com/llama3/… See the full description on the dataset page: https://huggingface.co/datasets/artnitolog/llm-generated-texts.royal_society_corpus_LLM_generated_metadatallm-generated-repair-data-comboaugmented_dataset_llm_generated_NER
📚 Augmented LLM-Generated NER Dataset for Scholarly Text
🧠 Dataset Summary
This dataset contains synthetically generated academic text tailored for Named Entity Recognition (NER) in the software engineering domain. The synthetic data augments scholarly writing using large language models (LLMs), with entity consistency maintained via token preservation.
The dataset is generated by merging and rephrasing pairs of annotated sentences from scholarly papers using… See the full description on the dataset page: https://huggingface.co/datasets/psresearch/augmented_dataset_llm_generated_NER.Reindex-Then-Adapt-LLM-Generated-DataICD-10-LLM-generated-Synthetic-Circulatory-System-I00-I99
MedGemma ICD-10 Clinical Notes Dataset — Circulatory System
Synthetic clinical notes generated by MedGemma-4B-IT for fine-tuning ICD-10-CM diagnosis code prediction models. Focused on Chapter 9: Diseases of the Circulatory System (I00-I99).
Dataset Summary
Split
Examples
Unique ICD-10 Codes
Train
6,275
1,255
Each example is a realistic clinical note paired with its ICD-10-CM diagnosis code, formatted as a chat conversation for instruction fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/singhankit16/ICD-10-LLM-generated-Synthetic-Circulatory-System-I00-I99.ICD-10-LLM-generated-Synthetic-Clinical-Note-G00-G99
MedGemma ICD-10 Clinical Notes Dataset
Synthetic clinical notes (english) generated by MedGemma-4B-IT for fine-tuning ICD-10-CM diagnosis code prediction models. Focused on Chapter 6: Diseases of the Nervous System (G00-G99).
Dataset Summary
Split
Examples
Unique ICD-10 Codes
Train
3,325
665
Eval
250
50
Each example is a realistic clinical note paired with its ICD-10-CM diagnosis code, formatted as a chat conversation for instruction fine-tuning.… See the full description on the dataset page: https://huggingface.co/datasets/singhankit16/ICD-10-LLM-generated-Synthetic-Clinical-Note-G00-G99.test_generated_seqLLM_Generated_Summaries_Dataset
BBC News Summaries Database
A small multilingual dataset of BBC News articles paired with two kinds of short summaries: human-written ones taken from XL-Sum, and machine-written ones I generated myself by prompting five different LLMs.
I built this for my TYP to compare how AI summaries stack up against the human reference across a handful of languages.
What's in here
5 languages: English (en), Spanish (es), French (fr), Arabic (ar), Mandarin Chinese (zh).
For… See the full description on the dataset page: https://huggingface.co/datasets/Ef05/LLM_Generated_Summaries_Dataset.SRS_document_LLM_generated_datasetgenerated_seqstrain_generated_seqasyncapi-llm_generated-datasetLLM-Generated_News_DatasetLLM_generated_Amharic_QA_for_Family_Code_of_Ethiopiagenerated_train_diff_v2LLMgenerated_fictive_medical_report_and_summaries_with_omissions_label_Fr_Healthcare
🏥 French Synthetic Medical Reports and Summaries with Omission Labels (Fr-Healthcare)
This dataset contains fictitious French medical reports, each paired with a summary and a binary label indicating whether the summary omits relevant factual content. It is designed solely for evaluating factual consistency and omission detection in Natural Language Processing, particularly in the medical domain. We must emphasize that all names, identifiers, dates, medical information, and any… See the full description on the dataset page: https://huggingface.co/datasets/AchOk78/LLMgenerated_fictive_medical_report_and_summaries_with_omissions_label_Fr_Healthcare.generated_train_diffdeepmath_data_generated_by_rlt_modelpayment-related-llm-generatedllm-generated-train-filesllm_generated_data_hw1
