THULab/greatplains_multisource_2000_2024
Great Plains Multisource NDVI–Climate Time Series (TsFile) Apache TsFile version of TaoDerong/GreatPlains-Multisource-2000-2024. Overview 8-day aligned multivariate time series for the southern Great Plains grassland region of the United States (~105°W–95°W, 32°N–40°N) over 2000–2024. The region is drought-sensitive, grassland-dominated, and highly responsive to precipitation anomalies. The dataset integrates multi-source observations: MODIS Terra NDVI (8-day… See the full description on the dataset page: https://huggingface.co/datasets/THULab/greatplains_multisource_2000_2024.
Great Plains Multisource NDVI–Climate Time Series (TsFile)
Apache TsFile version of `TaoDerong/GreatPlains-Multisource-2000-2024`.
Overview
8-day aligned multivariate time series for the southern Great Plains grassland region of the United States (~105°W–95°W, 32°N–40°N) over 2000–2024. The region is drought-sensitive, grassland-dominated, and highly responsive to precipitation anomalies. The dataset integrates multi-source observations: MODIS Terra NDVI (8-day composites), CHIRPS daily precipitation, and ERA5-Land daily surface soil moisture, 2 m air temperature, and potential evapotranspiration, aligned to the NDVI time step for unified modeling input.
Schema (TsFile structure)
- Time (INT64, milliseconds) — the 8-day composite date.
- ndvi (FIELD, FLOAT) — MODIS Terra NDVI (8-day composite).
- precip_8d_sum_mm (FIELD, FLOAT) — 8-day summed precipitation.
- soil_moisture_8d_mean (FIELD, FLOAT) — 8-day mean surface soil moisture.
- temp_8d_mean_C (FIELD, FLOAT) — 8-day mean 2 m air temperature.
- pet_8d_sum_mm (FIELD, FLOAT) — 8-day summed potential evapotranspiration.
Usage
Install the Apache TsFile Python SDK (pip install tsfile) and read a converted file:
from pathlib import Path
from tsfile import TsFileReader
path = Path("greatplains_multisource_2000_2024.tsfile")
with TsFileReader(str(path)) as reader:
schemas = reader.get_all_table_schemas()
print("tables:", list(schemas))
table_name = next(iter(schemas))
table = schemas[table_name]
columns = [column.get_column_name() for column in table.get_columns()]
print("columns:", columns)
field_names = [
column.get_column_name()
for column in table.get_columns()
if column.get_column_name() not in {"Time", "time"}
]
if field_names:
with reader.query_table(table_name, field_names[:3], batch_size=1024) as result:
batch = result.read_arrow_batch()
if batch is not None:
print(batch.to_pandas().head())Source & license
- Original dataset: https://huggingface.co/datasets/TaoDerong/GreatPlains-Multisource-2000-2024
- Author / publisher: TaoDerong
- License: CC-BY-4.0
