CoolFace
Datasetpublic

THULab/greatplains_multisource_2000_2024

Great Plains Multisource NDVI–Climate Time Series (TsFile) Apache TsFile version of TaoDerong/GreatPlains-Multisource-2000-2024. Overview 8-day aligned multivariate time series for the southern Great Plains grassland region of the United States (~105°W–95°W, 32°N–40°N) over 2000–2024. The region is drought-sensitive, grassland-dominated, and highly responsive to precipitation anomalies. The dataset integrates multi-source observations: MODIS Terra NDVI (8-day… See the full description on the dataset page: https://huggingface.co/datasets/THULab/greatplains_multisource_2000_2024.

sourceHugging Facecc-by-4.0updated 20d agoView on Hugging Face
0likes32downloads
Dataset Card

Great Plains Multisource NDVI–Climate Time Series (TsFile)

Apache TsFile version of `TaoDerong/GreatPlains-Multisource-2000-2024`.

Overview

8-day aligned multivariate time series for the southern Great Plains grassland region of the United States (~105°W–95°W, 32°N–40°N) over 2000–2024. The region is drought-sensitive, grassland-dominated, and highly responsive to precipitation anomalies. The dataset integrates multi-source observations: MODIS Terra NDVI (8-day composites), CHIRPS daily precipitation, and ERA5-Land daily surface soil moisture, 2 m air temperature, and potential evapotranspiration, aligned to the NDVI time step for unified modeling input.

Schema (TsFile structure)

  • Time (INT64, milliseconds) — the 8-day composite date.
  • ndvi (FIELD, FLOAT) — MODIS Terra NDVI (8-day composite).
  • precip_8d_sum_mm (FIELD, FLOAT) — 8-day summed precipitation.
  • soil_moisture_8d_mean (FIELD, FLOAT) — 8-day mean surface soil moisture.
  • temp_8d_mean_C (FIELD, FLOAT) — 8-day mean 2 m air temperature.
  • pet_8d_sum_mm (FIELD, FLOAT) — 8-day summed potential evapotranspiration.

Usage

Install the Apache TsFile Python SDK (pip install tsfile) and read a converted file:

python
from pathlib import Path
from tsfile import TsFileReader

path = Path("greatplains_multisource_2000_2024.tsfile")
with TsFileReader(str(path)) as reader:
    schemas = reader.get_all_table_schemas()
    print("tables:", list(schemas))
    table_name = next(iter(schemas))
    table = schemas[table_name]
    columns = [column.get_column_name() for column in table.get_columns()]
    print("columns:", columns)
    field_names = [
        column.get_column_name()
        for column in table.get_columns()
        if column.get_column_name() not in {"Time", "time"}
    ]
    if field_names:
        with reader.query_table(table_name, field_names[:3], batch_size=1024) as result:
            batch = result.read_arrow_batch()
            if batch is not None:
                print(batch.to_pandas().head())

Source & license

  • Original dataset: https://huggingface.co/datasets/TaoDerong/GreatPlains-Multisource-2000-2024
  • Author / publisher: TaoDerong
  • License: CC-BY-4.0