datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MacroLens
MacroLens
A benchmarking corpus for contextual financial reasoning under macroeconomic scenarios across 4,416 U.S. small- and micro-cap equities (2021-01-04 — 2026-03-31). MacroLens unifies seven tasks over a single point-in-time panel: contextual time-series forecasting, public valuation, financial-statement generation, scenario-conditioned return forecasting, private-company valuation, generator evaluation from natural-language descriptions, and real-estate valuation.
Task… See the full description on the dataset page: https://huggingface.co/datasets/DeepAuto-AI/MacroLens.aloha_static_battery_ep005_009This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "unknown",
"total_episodes": 5,
"total_frames": 3000,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 50,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/macrodata/aloha_static_battery_ep005_009.aloha_static_battery_ep000_004This dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v3.0",
"robot_type": "unknown",
"total_episodes": 5,
"total_frames": 3000,
"total_tasks": 1,
"chunks_size": 1000,
"data_files_size_in_mb": 100,
"video_files_size_in_mb": 200,
"fps": 50,
"splits": {
"train": "0:5"
},
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/macrodata/aloha_static_battery_ep000_004.macro-mdsfi-etf-macro-signal-master-dataOddBenchMacro-Dataset
MACRO: Advancing Multi-Reference Image Generation with Structured Long-Context Data
MACRO is a large-scale benchmark and training dataset for multi-reference image generation. It covers four task categories and four image-count brackets, providing both training splits and a curated evaluation benchmark.
Dataset Summary
Task
Train samples (per category)
Eval samples (per category)
Customization
1-3: 20,000 / 4-5: 20,000 / 6-7: 30,000 / ≥8: 30,000
250 each… See the full description on the dataset page: https://huggingface.co/datasets/Azily/Macro-Dataset.WGO-Bench
WGO-Bench: What's Going On Benchmark
WGO-Bench is a small, manually annotated benchmark for evaluating how well vision-language models can turn robot and egocentric manipulation videos into timestamped subtask annotations.
Each row contains one video episode, a high-level task instruction, and gold subtask segments with start time, end time, and a concise action label. The benchmark is designed for two related tasks:
Boundary detection: predict where one meaningful manipulation… See the full description on the dataset page: https://huggingface.co/datasets/macrodata/WGO-Bench.code-parrot-github-code
GitHub Code Dataset
Dataset Description
The GitHub Code dataset consists of 115M code files from GitHub in 32 programming languages with 60 extensions totaling in 1TB of data. The dataset was created from the public GitHub dataset on Google BiqQuery.
How to use it
The GitHub Code dataset is a very large dataset so for most use cases it is recommended to make use of the streaming API of datasets. You can load and iterate through the dataset with the following… See the full description on the dataset page: https://huggingface.co/datasets/macrocosm-os/code-parrot-github-code.SynBenchmacrobench-bittensor-01OvRBenchimagesmacrossdelta
Bangumi Image Base of Macross Delta
This is the image base of bangumi Macross Delta, we detected 45 characters, 4504 images in total. The full dataset is here.
Please note that these image bases are not guaranteed to be 100% cleaned, they may be noisy actual. If you intend to manually train models using this dataset, we recommend performing necessary preprocessing on the downloaded dataset to eliminate potential noisy samples (approximately 1% probability).
Here is the… See the full description on the dataset page: https://huggingface.co/datasets/BangumiBase/macrossdelta.MacroLens
MacroLens
A benchmarking corpus for contextual financial reasoning under macroeconomic scenarios across 4,416 U.S. small- and micro-cap equities (2021-01-04 — 2026-03-31). MacroLens unifies seven tasks over a single point-in-time panel: contextual time-series forecasting, public valuation, financial-statement generation, scenario-conditioned return forecasting, private-company valuation, generator evaluation from natural-language descriptions, and real-estate valuation.
Task… See the full description on the dataset page: https://huggingface.co/datasets/macrolens/MacroLens.nras-cypa-macrocyclic-glues-GA-II
NRAS–Cyclophilin A Macrocyclic Glue Designs (GA-II)
Why this target matters. NRAS-mutant melanoma has no approved targeted therapy and poor outcomes once immunotherapy fails; RAS(ON) tri-complex glues are among the very few mechanisms that engage NRAS at all.
180 small molecules generated de novo by the Technetium TC-43.ai engine (GA-II), conditioned on the NRAS·Cyclophilin A protein–protein interface, with macrocyclic ring closure imposed during generation.
Each molecule was… See the full description on the dataset page: https://huggingface.co/datasets/Tc-43/nras-cypa-macrocyclic-glues-GA-II.p2-etf-fpcr-macro-resultsSiDoLa-NS-Macro-mSC
SiDoLa-NS-Macro-mCB
https://sidolans01.mgifive.org/
Dataset Summary
This dataset contains high-resolution microscopy images of central nervous system (CNS) spinal cord sections, together with their corresponding labels for training segmentation and detection models. The dataset is primarily intended for training and benchmarking deep learning pipelines (e.g., YOLO, SAHI, SAM-based workflows).
In addition to the raw images and labels, some dataset folders also contain:… See the full description on the dataset page: https://huggingface.co/datasets/FIVE-MGI/SiDoLa-NS-Macro-mSC.arxiv_abstractsAll 2.3 million papers in the Arxiv, embedded via abstract with the InstructorXL model.
No claims are made about the copyright or license of contained materials. We assume no responsibilty for and are not liable under any circumstances for damages. Use at your own risk.
Good luck, have fun.
task705_mmmlu_answer_generation_high_school_macroeconomics
Dataset Card for Natural Instructions (https://github.com/allenai/natural-instructions) Task: task705_mmmlu_answer_generation_high_school_macroeconomics
Additional Information
Citation Information
The following paper introduces the corpus in detail. If you use the corpus in published work, please cite it:
@misc{wang2022supernaturalinstructionsgeneralizationdeclarativeinstructions,
title={Super-NaturalInstructions: Generalization via Declarative… See the full description on the dataset page: https://huggingface.co/datasets/Lots-of-LoRAs/task705_mmmlu_answer_generation_high_school_macroeconomics.global-macro-trade-indicators-dataset
Global Macro & Trade Indicators
Unified long-format macroeconomic and trade indicators from World Bank WDI, OECD and Eurostat official statistics, with IMF WEO and UN Comtrade derived metrics (no restricted raw values).
Part of the DataForge Open Data program — full production
packages, free for academic and personal use. Canonical dataset page:
https://data.zalize.com/datasets/global-macro-trade-indicators-dataset
Packages in this repo
Package
Tier
Rows… See the full description on the dataset page: https://huggingface.co/datasets/zalizedata/global-macro-trade-indicators-dataset.certvas-af-macro-sample
AF-MACRO — sample package
License-clean African & emerging-markets macroeconomic indicators, cited to their published source series, provenance-documented. See PROVENANCE.md and LICENSE.txt.
Quickstart
import duckdb, pandas as pd
df = duckdb.sql("select * from 'data/gold_macro.parquet'").df() # or read_csv
df = pd.read_parquet('data/gold_macro.parquet')
Excel: open any data/*.csv directly. Fields are documented in DICTIONARY.md; coverage and freshness in… See the full description on the dataset page: https://huggingface.co/datasets/267Certvas/certvas-af-macro-sample.p2-etf-transfer-entropy-macro-resultsSiDoLa-NS-Macro-mCB
SiDoLa-NS-Macro-mCB
https://sidolans01.mgifive.org/
Dataset Summary
This dataset contains high-resolution microscopy images of central nervous system (CNS) coronal sections, together with their corresponding labels for training segmentation and detection models. The dataset is primarily intended for training and benchmarking deep learning pipelines (e.g., YOLO, SAHI, SAM-based workflows).
In addition to the raw images and labels, some dataset folders also contain:… See the full description on the dataset page: https://huggingface.co/datasets/FIVE-MGI/SiDoLa-NS-Macro-mCB.macrodata_aloha_static_battery_ep005_009
ALOHA Static Battery Episodes 005-009 TsFile
Apache TsFile representation of macrodata/aloha_static_battery_ep005_009, a LeRobot v3 ALOHA manipulation dataset.
Source
Repository owner: Macrodata Labs
Repository contributor and uploader: hynky
License: Apache-2.0
Task: Place the battery into the slot of the remote controller.
Split: train; 5 episodes; 3,000 frame rows; 1 task; 50 fps
Numeric source layout: data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet… See the full description on the dataset page: https://huggingface.co/datasets/THULab/macrodata_aloha_static_battery_ep005_009.csc_eval_public
csc_eval_public
一、测评数据说明
1.1 测评数据来源
1.gen_de3.json(5545): '的地得'纠错, 由人民日报/学习强国/chinese-poetry等高质量数据人工生成;
2.lemon_v2.tet.json(1053): relm论文提出的数据, 多领域拼写纠错数据集(7个领域), ; 包括game(GAM), encyclopedia (ENC), contract (COT), medical care(MEC), car (CAR), novel (NOV), and news (NEW)等领域;
3.acc_rmrb.tet.json(4636): 来自NER-199801(人民日报高质量语料);
4.acc_xxqg.tet.json(5000): 来自学习强国网站的高质量语料;
5.gen_passage.tet.json(10000): 源数据为qwen生成的好词好句, 由几乎所有的开源数据汇总的混淆词典生成;… See the full description on the dataset page: https://huggingface.co/datasets/Macropodus/csc_eval_public.macroecon
MacroEconHorizon (TsFile)
Apache TsFile version of bkoyuncu/MacroEcon.
Overview
MacroEconHorizon is a curated collection of macroeconomic time series spanning 63
countries (plus the euro area and South Africa, 64 country files in total) from
1970 to 2023. It is designed to support nowcasting, forecasting, and scenario
analysis for machine-learning researchers and economic policy makers. Indicators
cover GDP, inflation, unemployment rates, commodity prices… See the full description on the dataset page: https://huggingface.co/datasets/THULab/macroecon.MacroBench
MacroBench: A Novel Testbed for Web Automation Scripts via Large Language Models
Dataset Description
MacroBench is a code-first benchmark that evaluates whether Large Language Models can synthesize reusable browser-automation programs (macros) from natural-language goals by reading HTML/DOM and emitting Selenium code.
Quick Links
Paper: arXiv:2510.04363
GitHub: MacroBench Repository
Dataset Files
The dataset includes the following files in the… See the full description on the dataset page: https://huggingface.co/datasets/hyunjun1121/MacroBench.cl-macros
Common Lisp Macro Transformations
A fine-tuning dataset for training models to generate Common Lisp macros.
Each example is a call-form, macro-definition, and expanded-form triple.
Summary
4,267 examples from 120+ Common Lisp libraries
Split: 2,985 train / 637 validation / 645 test
Format: JSONL with instruction, input, output, category, technique, complexity, quality_score
Mean quality score: 0.80
Sources: Let Over Lambda, On Lisp, Alexandria, Serapeum, Iterate… See the full description on the dataset page: https://huggingface.co/datasets/j14i/cl-macros.china-macro-trade-monthly-dataset
China Macro & Trade Monthly Indicators — IMF/World Bank Basis
Monthly China macroeconomic and goods-trade indicators built from the IMF Data Portal (SDMX 2.1) and World Bank Open Data APIs: CPI with COICOP breakdown, production index, goods trade totals, exchange rates, international reserves, plus monthly goods exports/imports against 200+ partner economies and 17 World Bank WDI annual context indicators — as published by the sources, no imputation.
Part of the DataForge Open… See the full description on the dataset page: https://huggingface.co/datasets/zalizedata/china-macro-trade-monthly-dataset.
