datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
FRED-SemBench
FRED-SemBench
FRED-SemBench is a 200-question benchmark candidate for evaluating whether LLM
agents retrieve macroeconomic answers with the intended concept, series,
transformation, unit, observation period, and data-vintage semantics.
This dataset accompanies the FinNLP 2026 paper
“FRED-SemBench: Evaluating Semantic Reliability in LLM Access to
Macroeconomic Data”
by Wilson Wang, Chandler Han, and Peter Zhang (Kairos-AI).
Status and scope
50 independently… See the full description on the dataset page: https://huggingface.co/datasets/wangjinh/FRED-SemBench.Instruction-Tuning-with-GPT-4-RedPajama-Chat
Instruction Tuning with GPT 4 RedPajama-Chat
This dataset has been converted from the Instruction-Tuning-with-GPT-4 dataset for the purpose of fine-tuning the RedPajama-INCITE-Chat-3B-v1 model.
About Instruction-Tuning-with-GPT-4
English Instruction-Following Data generated by GPT-4 using Alpaca prompts for fine-tuning LLMs.
Usage and License Notices
The data is intended and licensed for research use only. The dataset is CC BY NC 4.0 (allowing only… See the full description on the dataset page: https://huggingface.co/datasets/Fredithefish/Instruction-Tuning-with-GPT-4-RedPajama-Chat.Qwen3.5-reasoning-700x
Dataset Card (Qwen3.5-reasoning-700x)
Dataset Summary
Qwen3.5-reasoning-700x is a high-quality distilled dataset.
This dataset uses the high-quality instructions constructed by Alibaba-Superior-Reasoning-Stage2 as the seed question set. By calling the latest Qwen3.5-27B full-parameter model on the Alibaba Cloud DashScope platform as the teacher model, it generates high-quality responses featuring long-text reasoning processes (Chain-of-Thought). It covers several major… See the full description on the dataset page: https://huggingface.co/datasets/freddm/Qwen3.5-reasoning-700x.
