datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gw-tti-playground-queueorena-cold-storage
orena-cold-storage
ORena / FOCUS MICCAI 2026 — DiffAtlas 冷存储归档。本地集群空间释放后迁移至此长期保存。
总览:20,390 文件 / 1,495.75 GiB
归档时间:2026-08-30
可见性:public
完整性:全部 2,750 个 .pt 权重 LFS SHA256 校验有效(0 空 / 0 损坏 / 0 重复)
内容索引
Model/ — 模型权重(2,589 文件 / 1,486.46 GiB)
子项目(实验)
文件数
大小
TS_train_80percent_base_10w
784
450.13 GiB
DiffAtlas_TotalSegmentator
500
287.07 GiB
TS_train_80percent
450
258.37 GiB
DiffAtlas_MMWHS-MRI_full
426
244.59 GiB… See the full description on the dataset page: https://huggingface.co/datasets/gwd200/orena-cold-storage.GWHD
GWHD 2021: Global Wheat Head Dataset (Object Detection)
Unofficial redistribution of the Global Wheat Head Dataset (GWHD) 2021 competition release, reformatted into a standardized YOLO-compatible directory layout.
Disclaimer
This repository is not an official release of the Global Wheat Head Dataset.
GWHD was created by the Global Wheat Head Detection consortium — a multi-institution, multi-country collaboration (David, Serouart, Madec, et al.) — who… See the full description on the dataset page: https://huggingface.co/datasets/dronefreak/GWHD.STdata1GWFSS-competition
Introduction
Competition Page
If you want any update on the Global Wheat Dataset Community, go on https://www.global-wheat.com/
Wheat is a cornerstone of global food security, serving as a dietary staple for billions of people worldwide. Detailed analysis of wheat plants can help scientists and farmers cultivate healthier, more resilient, and more productive crops. The Global Wheat Full Semantic Segmentation (GWFSS) task aims to perform pixel-level segmentation of plant components… See the full description on the dataset page: https://huggingface.co/datasets/XIANG-Shuai/GWFSS-competition.gwhd2021_set0
Global Wheat Head Detection (GWHD) dataset
Cite as:
David, Etienne et al. (2020). Global Wheat Head Detection (GWHD) dataset: a large and diverse dataset of high-resolution RGB-labelled images to develop and benchmark wheat head detection methods. Plant Phenomics, 2020. Science Partner Journal.DOI: 10.5281/zenodo.5092309
License: CC-BY-4.0
GWFSSThis is a version of the GWFSS-competition repo that is usable as a standalone huggingface dataset.
The content is identical to the zips in the GWFSS-competition files tab.
m3guo-game-stagesGWFSS_v1.0This is the labelled split of GWFSS.
The unlabelled split is available at https://www.research-collection.ethz.ch/handle/20.500.11850/734546.
The benchmark model is available at: https://huggingface.co/GlobalWheat/GWFSS_model_v1.0.
Related Paper
This dataset is associated with the following paper:
The Global Wheat Full Semantic Organ Segmentation (GWFSS) Dataset
https://doi.org/10.1016/j.plaphe.2025.100084
spectro_caption_dataset
Dataset Card for "spectro_caption_dataset"
More Information needed
gwhd2021_augmentedset1
Global Wheat Head Detection (GWHD) dataset
Cite as:
David, Etienne et al. (2020). Global Wheat Head Detection (GWHD) dataset: a large and diverse dataset of high-resolution RGB-labelled images to develop and benchmark wheat head detection methods. Plant Phenomics, 2020. Science Partner Journal.DOI: 10.5281/zenodo.5092309
License: CC-BY-4.0
Augmentations utilizadas
HorizontalFlip(p=0.5)
VerticalFlip(p=0.2)
Affine(p=0.7, balanced_scale=False, border_mode=0, fill=0.0… See the full description on the dataset page: https://huggingface.co/datasets/claytonsds/gwhd2021_augmentedset1.corpusphilo-v2
Corpus de copies de philosophie
Chaque copie possède son propre dossier, nommé avec le sujet en premier, puis le concours, l’année lorsqu’elle est connue et la note.
source.pdf : copie originale.
images/ : pages numérotées selon le PDF, y compris les pages vides conservées pour archivage.
transcriptions_md/ : toutes les transcriptions Markdown, y compris les versions indépendantes de chaque modèle. Les anciens exports et leurs variantes sont également à la racine de ce dossier… See the full description on the dataset page: https://huggingface.co/datasets/GwendalTsang/corpusphilo-v2.car-classification-image-datasetarticraft-vl-7k-128k
Articraft VL 7k 128k
This dataset contains Articraft visual agentic SFT samples for Megatron-SWIFT.
Files
train.jsonl: training split.
test.jsonl: validation split.
train_index.jsonl, test_index.jsonl: metadata/index rows.
metadata.json: split and filtering metadata.
excluded_over_128k.jsonl: rows excluded by the 128k token filter.
images/*.png: rendered Blender images referenced by the JSONL rows.
Format
Each JSONL row contains:
messages: ms-swift… See the full description on the dataset page: https://huggingface.co/datasets/gwc000/articraft-vl-7k-128k.sample_datasetstgfn-repro-posterCORPUS
Corpus de copies de philosophie
Chaque copie possède son propre dossier, nommé avec le sujet en premier, puis le concours, l’année lorsqu’elle est connue et la note.
source.pdf : copie originale.
images/ : pages numérotées selon le PDF, y compris les pages vides conservées pour archivage.
La … — Claude ….md : transcriptions indépendantes de chaque modèle.
metadata.json : source, empreintes, pages transcrites ou exclues, modèles et paramètres.
provenance/ : réponses brutes par… See the full description on the dataset page: https://huggingface.co/datasets/GwendalTsang/CORPUS.gwalther-handwriting-ground-truth
Gwalther's Latin Handwriting Ground Truth
Dataset Description
The Gwalther's Latin Handwriting Ground Truth dataset provides resources for Handwritten Text Recognition (HTR) systems, specifically focused on a single hand from the Reformation era.
The dataset contains ground truth for the handwriting of Ruolph Gwalther (1519-1586), extracted from his work Lateinische Gedichte, which accumulated his writings between 1540 and 1580.
Key Figures
Metric
Value… See the full description on the dataset page: https://huggingface.co/datasets/SSamDav/gwalther-handwriting-ground-truth.gwanbo-digital-render
Gwanbo Digital Re-render (관보 디지털 재렌더)
대한민국 관보(官報) 1994–2000 PETY 스캔본을 인식+재렌더(recognition + re-render)
파이프라인으로 디지털 본 품질로 재조판한 데이터셋.
생성 방식
스캔 PDF(250dpi) → Chandra OCR-2 → Qwen3.6-27B 후처리(마크다운)
→ fontsr.render_digital:
· gwanbo_ocr_error_patterns.json 오류패턴 교정(애급→예금, 소재→소계, 조홍→조흥 …)
· 내용 기반 컬럼폭 표 레이아웃
· UnBatang(은바탕 = 한양 신명조 계보, HWP 3.0~2000 기본 본문 글꼴) 렌더
스캔 글리프를 복원하는 SR과 달리, 인식된 텍스트를 깨끗한 명조 폰트로 재조판해
'디지털 본에 준하는' 완벽한 글리프를 생성한다. OCR 혼동 오류까지 교정된… See the full description on the dataset page: https://huggingface.co/datasets/yakdoli/gwanbo-digital-render.MTBench_finance_news
MTBench: A Multimodal Time Series Benchmark
MTBench (Huggingface, Github, Arxiv) is a suite of multimodal datasets for evaluating large language models (LLMs) in temporal and cross-modal reasoning tasks across finance and weather domains.
Each benchmark instance aligns high-resolution time series (e.g., stock prices, weather data) with textual context (e.g., news articles, QA prompts), enabling research into temporally grounded and multimodal understanding.
🏦 Labled… See the full description on the dataset page: https://huggingface.co/datasets/gwiezdny/MTBench_finance_news.roleplay-characters
Dataset Card for "roleplay-characters"
More Information needed
gwhdboysboys <3
twitter-Starry698-2026.02.27-2027292506421932078-gWPr2fnLVqTkT47V-part1twitter-dyf798798-2024.09.02-1830587173235724769-p7hkV_GW66nk0OHE-part1MME
Evaluation Dataset for MME
discrete_landcoverher2_3danalytic_gw_book
