THULab/supply_chain_customer
Supply_Chain_Customer (TsFile) Apache TsFile version of the Supply_Chain_Customer sub-dataset of Real-TSF/TIME-ProcessedCSV. Overview Supply_Chain_Customer is one of the processed time-series collections bundled in TIME-ProcessedCSV, a multi-domain repository of cleaned CSV series (energy, transport, weather, finance, health and more), each with a timestamp column and one or more measurement columns per file. Source files: 1 CSV file(s) under… See the full description on the dataset page: https://huggingface.co/datasets/THULab/supply_chain_customer.
SupplyChainCustomer (TsFile)
Apache TsFile version of the Supply_Chain_Customer sub-dataset of `Real-TSF/TIME-ProcessedCSV`.
Overview
Supply_Chain_Customer is one of the processed time-series collections bundled in TIME-ProcessedCSV, a multi-domain repository of cleaned CSV series (energy, transport, weather, finance, health and more), each with a timestamp column and one or more measurement columns per file.
- Source files: 1 CSV file(s) under
https://huggingface.co/datasets/Real-TSF/TIME-ProcessedCSV/tree/main/Supply_Chain_Customer - Converted rows: 72,252 (long format: one row per measurement)
- Data files:
['supply_chain_customer.tsfile']
Schema (TsFile structure)
- Time (INT64, milliseconds) — the source
timestampcolumn. - freq (TAG) — source sampling-frequency directory (e.g.
H,15T,D). - series (TAG) — source file stem (e.g.
item0,NRSROT). - channel (TAG) — source measurement column name.
- measurement (FIELD, FLOAT) — the measurement value.
Usage
Install the Apache TsFile Python SDK (pip install tsfile) and read a converted file:
from pathlib import Path
from tsfile import TsFileReader
path = Path("supply_chain_customer.tsfile")
with TsFileReader(str(path)) as reader:
schemas = reader.get_all_table_schemas()
print("tables:", list(schemas))
table_name = next(iter(schemas))
table = schemas[table_name]
columns = [column.get_column_name() for column in table.get_columns()]
print("columns:", columns)
field_names = [
column.get_column_name()
for column in table.get_columns()
if column.get_column_name() not in {"Time", "time"}
]
if field_names:
with reader.query_table(table_name, field_names[:3], batch_size=1024) as result:
batch = result.read_arrow_batch()
if batch is not None:
print(batch.to_pandas().head())Source & license
- Original dataset: https://huggingface.co/datasets/Real-TSF/TIME-ProcessedCSV/tree/main/SupplyChainCustomer
- Bundle: https://huggingface.co/datasets/Real-TSF/TIME-ProcessedCSV
- License: apache-2.0
