geogpt
Datasets
All datasets matching “geogpt”GeoGPT-CoT-QA
GeoGPT-CoT-QA Dataset: A Large-scale Geoscience Chain-of-Thought QA Dataset for Supervised Fine-Tuning of LLMs
1. Dataset Description
We introduce GeoGPT-CoT-QA Dataset, a large-scale synthetic question–answer (QA) corpus enriched with chain-of-thought (CoT) reasoning traces, developed to support supervised fine-tuning (SFT) of geoscience reasoning models. The GeoGPT-R1-Preview is specifically fine-tuned using this dataset to enhance its geoscience reasoning capabilities.… See the full description on the dataset page: https://huggingface.co/datasets/GeoGPT-Research-Project/GeoGPT-CoT-QA.GeoGPT4V-1.0GeoGPT-QA
GeoGPT-QA Dataset: A Large-scale Geoscience QA Dataset for Supervised Fine-tuning of LLMs
1. Dataset Description
We introduce GeoGPT-QA Dataset, a large-scale synthetic question–answer (QA) corpus developed to support supervised fine-tuning (SFT) of geoscience foundation models.
The dataset is derived from open-access geoscience publications distributed under the CC BY license. Using an automated data synthesis pipeline, we generated professional QA pairs from article… See the full description on the dataset page: https://huggingface.co/datasets/GeoGPT-Research-Project/GeoGPT-QA.GeoGPT_Training_Data_from_Open-Access_Papers
Description
This dataset lists the publishers and journals that have released open access geoscience papers used for GeoGPT training. It also explains how GeoGPT filters and selects content based on licensing terms to ensure compliance. The dataset includes papers published under various open access licenses, among which those licensed under CC BY and CC BY-NC have been used for training. In total, we have collected approximately 280,000 such papers from 15 publishers and… See the full description on the dataset page: https://huggingface.co/datasets/GeoGPT-Research-Project/GeoGPT_Training_Data_from_Open-Access_Papers.GeoRAG-QA
GeoRAG-QA: A Benchmark Test Set for Geoscience Information Retrieval
1. Dataset Description
We introduce GeoRAG-QA, a curated test set designed for evaluating retrieval-augmented generation (RAG) systems and information retrieval approaches in the geoscience domain.
GeoRAG-QA was constructed using the Test Set Generation module from RAGAS (Es et al., 2023). The dataset consists of automatically generated QA items based on open-access geoscience publications (see GeoGPT… See the full description on the dataset page: https://huggingface.co/datasets/GeoGPT-Research-Project/GeoRAG-QA.GeoGPT_Training_Data_from_Geoscience_Subset_of_CommonCrawl
Description
This dataset is a geoscience-specific subset of CommonCrawl used for GeoGPT training. CommonCrawl is a free and open repository of web crawl data with over 250 billion web pages and is widely used by leading large language models. We apply data mining algorithms to extract geoscience-related content from this vast dataset.
This dataset comprises 12,414,268 samples, each containing the following metadata to trace the data source within CommonCrawl:
id (string):… See the full description on the dataset page: https://huggingface.co/datasets/GeoGPT-Research-Project/GeoGPT_Training_Data_from_Geoscience_Subset_of_CommonCrawl.
