Plagiarism
python_plagiarism_code_dataset
Python Plagiarism Code Dataset
Overview
This dataset contains pairs of Python code samples with varying degrees of similarity, designed for training and evaluating plagiarism detection systems. The dataset was created using Large Language Models (LLMs) to generate synthetic code variations at different transformation levels, simulating real-world plagiarism scenarios in an academic context.
Purpose
The dataset addresses the limitations of existing code… See the full description on the dataset page: https://huggingface.co/datasets/nop12/python_plagiarism_code_dataset.bangla-plagiarism-datasetBangla Plagiarism Datasetpan-plagiarism-corpus-2011plagiarism
