CoolFace
Datasetpublic

siddqamar/GMO-Myths-and-Truths

Dataset Card for GMO Myths and Truths (NLP Classification) Dataset Summary This dataset contains a structured collection of claims and evidence-based findings regarding Genetically Modified Organisms (GMOs). The data was extracted and adapted from the technical report: "GMO Myths and Truths: An evidence-based examination of the claims made for the safety and efficacy of genetically modified crops" (Version 1.3a, June 2012). It is designed for binary text… See the full description on the dataset page: https://huggingface.co/datasets/siddqamar/GMO-Myths-and-Truths.

sourceHugging Facecc-by-4.0updated 5mo agoView on Hugging Face
0likes20downloads
Dataset Card

Dataset Card for GMO Myths and Truths (NLP Classification)

Dataset Summary

This dataset contains a structured collection of claims and evidence-based findings regarding Genetically Modified Organisms (GMOs). The data was extracted and adapted from the technical report: "GMO Myths and Truths: An evidence-based examination of the claims made for the safety and efficacy of genetically modified crops" (Version 1.3a, June 2012).

It is designed for binary text classification, sentiment analysis, and semantic search tasks within the biotechnology domain.

Data Structure

The dataset is organized into two primary subsets:

  1. 1.Pure: Direct extractions of paired Myth/Truth statements from the source document (94 balanced samples).
  2. 2.Augmented: A robust training set expanded via Linguistic Data Augmentation (including synonym substitution, structural variation, and contextual wrapping) to improve model generalization (500+ balanced samples).

Label Taxonomy

Statements are categorized into a binary format suitable for machine learning training:

LabelDesignationDescription
0MythRepresents industry-proponent claims or marketing arguments as identified by the source report.
1TruthRepresents the evidence-based findings, scientific rebuttals, or safety data reported by the authors.

Source Attribution

  • —Report: GMO Myths and Truths
  • —Authors: Michael Antoniou (PhD), Claire Robinson (MPhil), John Fagan (PhD)
  • —Publisher: Earth Open Source
  • —Release Date: June 2012
  • —Version: 1.3a

Curation & Methodology

Technical Curator Statement

This dataset was technically curated and structured from the original PDF source for the purpose of Natural Language Processing (NLP) research and model benchmarking. The curation process involved automated text extraction, de-duplication, and linguistic transformation to adapt the source material into a machine-readable format.

Liability & Intent

The curator of this dataset is a technical developer and not a subject matter expert in molecular genetics or agricultural biotechnology. This project is a technical implementation intended for training text-classification models (such as all-MiniLM-L6-v2) and does not represent a personal or scientific endorsement of the claims contained within. For all scientific or policy inquiries, users must refer to the primary authors listed in the source section.

Supported Tasks

  • —Binary Text Classification: Distinguishing between proponent claims and evidence-based report findings.
  • —Semantic Similarity: Benchmarking embedding models on specialized scientific and industrial terminology.
  • —Data Augmentation Research: Evaluating the performance of models trained on linguistically varied synthetic samples.

Licensing

This dataset is a derivative work based on the Earth Open Source publication. It is provided for research, benchmarking, and non-commercial educational use. For permissions regarding the original text or redistribution of the core findings, please consult the original authors.


How to Load in Python

python
from datasets import load_dataset

# To load the augmented dataset for model training
dataset = load_dataset("siddqamar/GMO-Myths-and-Truths", split="augmented")

# To load the pure, direct extractions
pure_claims = load_dataset("siddqamar/GMO-Myths-and-Truths", split="pure")

license: cc-by-4.0 ---