gwiezdny/MTBench_finance_news
MTBench: A Multimodal Time Series Benchmark MTBench (Huggingface, Github, Arxiv) is a suite of multimodal datasets for evaluating large language models (LLMs) in temporal and cross-modal reasoning tasks across finance and weather domains. Each benchmark instance aligns high-resolution time series (e.g., stock prices, weather data) with textual context (e.g., news articles, QA prompts), enabling research into temporally grounded and multimodal understanding. 🏦 Labled… See the full description on the dataset page: https://huggingface.co/datasets/gwiezdny/MTBench_finance_news.
MTBench: A Multimodal Time Series Benchmark
MTBench (Huggingface, Github, Arxiv) is a suite of multimodal datasets for evaluating large language models (LLMs) in temporal and cross-modal reasoning tasks across finance and weather domains.
Each benchmark instance aligns high-resolution time series (e.g., stock prices, weather data) with textual context (e.g., news articles, QA prompts), enabling research into temporally grounded and multimodal understanding.
🏦 Labled Finance News Dataset
We collect and clean financial news from sources including GlobeNews, MarketWatch, SeekingAlpha, Zacks, Invezz, Quartz (QZ), PennyStocks, and Benzinga, covering May 2021 to September 2023.
Each article is parsed to extract the title, context, associated stock names, and publishing date. We retain a curated subset of 20,000 articles, ensuring balanced length distribution. GPT-4o is used to annotate each article with:
- News type
- Temporal effect range
- Sentiment polarity
🔖 News Label Taxonomy
A. News Type
- Market News & Analysis
- Macro & Economic: Interest rates, inflation, geopolitical events
- Stock Market Updates: Indices, sector performance
- Company-Specific: Earnings, M&A, leadership changes
- Investment & Stock Analysis
- Fundamental: Earnings, revenue, P/E ratios
- Technical: Patterns, indicators (RSI, moving averages)
- Recommendations: Analyst upgrades/downgrades, price targets
- Trading & Speculative
- Options & Derivatives: Futures, strategies
- Penny Stocks: Micro-cap, high-risk investments
- Short Selling: Squeezes, manipulation, regulatory topics
B. Temporal Impact
- Retrospective
- Short-Term: ≤ 3 months
- Medium-Term: 3–12 months
- Long-Term: > 1 year
- Present-Focused
- Real-Time: Intraday updates
- Recent Trends: Ongoing market behavior
- Forward-Looking
- Short-Term Outlook: 3–6 months
- Medium-Term: 6 months – 2 years
- Long-Term: > 2 years
C. Sentiment
- Positive
- Bullish, Growth-Oriented, Upbeat Reactions
- Neutral
- Balanced, Mixed Outlook, Speculative
- Negative
- Bearish, Risk Alerts, Market Panic
📦 Other MTBench Datasets
🔹 Finance Domain
- `MTBench_finance_news` 20,000 articles with URL, timestamp, context, and labels
- `MTBench_finance_stock` Time series of 2,993 stocks (2013–2023)
- `MTBench_finance_aligned_pairs_short` 2,000 news–series pairs
- Input: 7 days @ 5-min
- Output: 1 day @ 5-min
- `MTBench_finance_aligned_pairs_long` 2,000 news–series pairs
- Input: 30 days @ 1-hour
- Output: 7 days @ 1-hour
- `MTBench_finance_QA_short` 490 multiple-choice QA pairs
- Input: 7 days @ 5-min
- Output: 1 day @ 5-min
- `MTBench_finance_QA_long` 490 multiple-choice QA pairs
- Input: 30 days @ 1-hour
- Output: 7 days @ 1-hour
🔹 Weather Domain
- `MTBench_weather_news` Regional weather event descriptions
- `MTBench_weather_temperature` Meteorological time series from 50 U.S. stations
- `MTBench_weather_aligned_pairs_short` Short-range aligned weather text–series pairs
- `MTBench_weather_aligned_pairs_long` Long-range aligned weather text–series pairs
- `MTBench_weather_QA_short` Short-horizon QA with aligned weather data
- `MTBench_weather_QA_long` Long-horizon QA for temporal and contextual reasoning
🧠 Supported Tasks
MTBench supports a wide range of multimodal and temporal reasoning tasks, including:
- 📈 News-aware time series forecasting
- 📊 Event-driven trend analysis
- ❓ Multimodal question answering (QA)
- 🔄 Text-to-series correlation analysis
- 🧩 Causal inference in financial and meteorological systems
📄 Citation
If you use MTBench in your work, please cite:
@article{chen2025mtbench,
title={MTBench: A Multimodal Time Series Benchmark for Temporal Reasoning and Question Answering},
author={Chen, Jialin and Feng, Aosong and Zhao, Ziyu and Garza, Juan and Nurbek, Gaukhar and Qin, Cheng and Maatouk, Ali and Tassiulas, Leandros and Gao, Yifeng and Ying, Rex},
journal={arXiv preprint arXiv:2503.16858},
year={2025}
}
