datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
flux_generatedgenerated-csvsarc-voicesamples-generatedgenerated-imagesgenerated_chat_0.4M
Generated_Chat_0.4M
内容
包含约40万条由BELLE项目生成的个性化角色对话数据,包含角色介绍。
注意:此数据集是由ChatGPT产生的,未经过严格校验,题目或解题过程可能包含错误。使用过程中请注意这一点。
样例
{
"instruction": "于以下角色信息完成一段对话\nAlice是一个年轻的律师,她正在为一个涉及濒危物种保护的案件做准备。Bob是一位当地的居民,他所在的社区经常受到该濒危物种的干扰。\n",
"input": "",
"output": "\nAlice: 你好Bob,我是濒危物种保护案的律师,很高兴见到你。我们前几天发给你的文件你看到了吗?\nBob: 看到了,你们说要禁止我们在该物种的常住地采集蘑菇,这很不公平。\nAlice: 保护濒危物种对整个生态系统都很重要,如果我们的举动损害了这个物种的生存,那么整个生态系统都将遭受损失。\nBob: 我理解您的立场,但是我们一直以来都依靠这个物种来维持我们的经济生活,现在我们要怎么办?\nAlice:… See the full description on the dataset page: https://huggingface.co/datasets/BelleGroup/generated_chat_0.4M.Generated_summariesgenerated-videos200k_HEAVY_gpt4o-description-gpt4omini-code_generated_problemsHere is the dataset of ~100k synthetic data generated by 162 seeds.
We generate the dataset with the following steps and two approaches:
Generate ~110k descriptions by GPT4o.
Approach 1: Generate ~110k codes follow each description by GPT4o-mini.
Approach 2: Generate ~110k codes follow each description by GPT4o-mini and suggest it to use specific library functions.
Run the ~220k codes and do auto-filtering.
Get the final ~200k legitimate ARC-like tasks with examples.
human_ai_generated_text
Human or AI-Generated Text
The data can be valuable for educators, policymakers, and researchers interested in the evolving education landscape, particularly in detecting or identifying texts written by Humans or Artificial Intelligence systems.
File Name
model_training_dataset.csv
File Structure
id: Unique identifier for each record.
human_text: Human-written content.
ai_text: AI-generated texts.
instructions: Description of the task given to both Humans and… See the full description on the dataset page: https://huggingface.co/datasets/dmitva/human_ai_generated_text.gpt-oss120b-generated-perfectblendmsmarco_generated_question_runsgenerate-readme-eval
Generate README Eval
The generate-readme-eval is a dataset (train split) and benchmark (test split) to evaluate the effectiveness of LLMs
when summarizing entire GitHub repos in form of a README.md file. The datset is curated from top 400 real Python repositories
from GitHub with at least 1000 stars and 100 forks. The script used to generate the dataset can be found here.
For the dataset we restrict ourselves to GH repositories that are less than 100k tokens in size to allow us to… See the full description on the dataset page: https://huggingface.co/datasets/patched-codes/generate-readme-eval.Generated-LoRA-Input-Images-for-Mitigating-Biasarabic-generated-abstracts
Arabic Machine-Generated Text Dataset
This dataset contains machine-generated Arabic text across multiple generation methods, and Large Language Model (LLMs).
It was created as part of the research paper "Arabic machine-generated text detection: Stylometric analysis and cross-model evaluation" (https://www.sciencedirect.com/science/article/abs/pii/S0957417425042599).
📋 Dataset Overview
The dataset addresses the need for comprehensive Arabic machine-generated text… See the full description on the dataset page: https://huggingface.co/datasets/KFUPM-JRCAI/arabic-generated-abstracts.generated-finemath-585942-node-0generated-finemath-292968-node-2Hy-Generated-audio-data-with-cv20.0
Hy-Generated Audio Data with CV20.0
This dataset provides Armenian speech data consisting of both real and generated audio clips.
The train, test, and eval splits are derived from the Common Voice 20.0 Armenian dataset.
The generated split contains 100,000 high-quality clips synthesized using a fine-tuned F5-TTS model, covering 404 equal distribution of synthetic voices.
📊 Dataset Statistics
Split
# Clips
Duration (hours)
train
9,300
13.53
test
5,818… See the full description on the dataset page: https://huggingface.co/datasets/ErikMkrtchyan/Hy-Generated-audio-data-with-cv20.0.generated-finemath-292968-node-3generated-finemath-322376-node-0AI-and-Human-Generated-Text
AI & Human Generated Text
I am Using this dataset for AI Text Detection for https://exnrt.com.
Check Original DataSet GitHub Repository Here: https://github.com/panagiotisanagnostou/AI-GA
Description
The AI-GA dataset, short for Artificial Intelligence Generated Abstracts, comprises abstracts and titles. Half of these abstracts are generated by AI, while the remaining half are original. Primarily intended for research and experimentation in natural language… See the full description on the dataset page: https://huggingface.co/datasets/Ateeqq/AI-and-Human-Generated-Text.dbpedia-entity-generated-queries
Dataset Card for BEIR Benchmark
Dataset Summary
BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks:
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018
Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus
News Retrieval: TREC-NEWS, Robust04
Argument Retrieval: Touche-2020, ArguAna
Duplicate Question Retrieval: Quora, CqaDupstack
Citation-Prediction: SCIDOCS
Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/dbpedia-entity-generated-queries.Hy-Generated-audio-data-2
Hy-Generated Audio Data 2
This dataset provides Armenian speech data consisting of generated audio clips and is addition to this dataset.
The generated split contains 137,419 high-quality clips synthesized using a fine-tuned F5-TTS model, covering 404 equal distribution of synthetic voices.
📊 Dataset Statistics
Split
# Clips
Duration (hours)
generated
137,419
173.76
Total duration: ~173 hours
🛠️ Loading the Dataset
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/ErikMkrtchyan/Hy-Generated-audio-data-2.generated-finemath-292968-node-1task1729_personachat_generate_next
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task1729_personachat_generate_next
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative Instructions on 1600+ NLP Tasks}… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task1729_personachat_generate_next.climate-fever-generated-queries
Dataset Card for BEIR Benchmark
Dataset Summary
BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks:
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018
Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus
News Retrieval: TREC-NEWS, Robust04
Argument Retrieval: Touche-2020, ArguAna
Duplicate Question Retrieval: Quora, CqaDupstack
Citation-Prediction: SCIDOCS
Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/climate-fever-generated-queries.ghanaian-corpus-generate
Ghanaian Corpus — Generated Content-Writer Prompts
Description
Synthetic content-writer prompts generated from 534,294 Ghanaian corpus chunks using Qwen2.5-0.5B-Instruct via vLLM.
Each row contains:
source_type: Category of the source document (e.g., parliament, news, academic)
source: Name or identifier of the source document
page_range: Page range within the source
text: The original corpus chunk text (minimum 4 sentences, each with 6+ unique alphabetic words)… See the full description on the dataset page: https://huggingface.co/datasets/ghanaopenai/ghanaian-corpus-generate.magpie-qwen2.5-pro-1m-v0.1-Qwen3-235B-A22B-Instruct-2507-FP8-generatedperfectblend-Qwen3-235B-A22B-Instruct-2507-FP8-generatednq-generated-queries
Dataset Card for BEIR Benchmark
Dataset Summary
BEIR is a heterogeneous benchmark that has been built from 18 diverse datasets representing 9 information retrieval tasks:
Fact-checking: FEVER, Climate-FEVER, SciFact
Question-Answering: NQ, HotpotQA, FiQA-2018
Bio-Medical IR: TREC-COVID, BioASQ, NFCorpus
News Retrieval: TREC-NEWS, Robust04
Argument Retrieval: Touche-2020, ArguAna
Duplicate Question Retrieval: Quora, CqaDupstack
Citation-Prediction: SCIDOCS
Tweet… See the full description on the dataset page: https://huggingface.co/datasets/BeIR/nq-generated-queries.human-vs-Ai-generated-dataset
