8-k
Datasets
All datasets matching “8-k”8KDehazeThe Full version of the datasets 8KDehaze
MINI version : https://huggingface.co/datasets/fengyanzi/8KDehaze_mini
The first dataset of extremely large image dehazing
from CVPR2025 paper "Tokenize Image Patches: Global Context Fusion for Effective Haze Removal in Large Images"
Paper link: https://openaccess.thecvf.com/content/CVPR2025/html/Chen_Tokenize_Image_Patches_Global_Context_Fusion_for_Effective_Haze_Removal_CVPR_2025_paper.html
DehazeXL code: https://github.com/CastleChen339/DehazeXL… See the full description on the dataset page: https://huggingface.co/datasets/CastleChen339/8KDehaze.UGround-V1-8k
UGround-WebHybrid-8K
This dataset is a curated 8K-sample subset from the original UGround-V1-Data (Web-Hybrid), as mentioned in our paper. It serves as part of the training corpus for GUI grounding tasks, focusing on diverse web interface screenshots across resolutions and aspect ratios.
Paper and Code
Paper: ZonUI-3B: A Lightweight Vision-Language Model for Cross-Resolution GUI Grounding
Code: https://github.com/zonghanHZH/ZonUI-3B
Dataset Details
Source:… See the full description on the dataset page: https://huggingface.co/datasets/zonghanHZH/UGround-V1-8k.UltraData-SFT-2605-no-think-8k-32k
UltraData-SFT-2605 · no_think · 8k–32k
A length-filtered subset of the no_think split of
openbmb/UltraData-SFT-2605,
containing conversations whose token length falls in the 8k–32k range.
This is the medium-length tier intended for standard long-context SFT.
Two companion tiers were produced from the same source:
Dataset
Length range
Records
this repo — fxmeng/UltraData-SFT-2605-no-think-8k-32k
8k–32k tokens
623,421
fxmeng/UltraData-SFT-2605-no-think-32k-200k… See the full description on the dataset page: https://huggingface.co/datasets/fxmeng/UltraData-SFT-2605-no-think-8k-32k.TCGA-UniformTumor-8K
Dataset Card for TCGA-UniformTumor-8K
What is TCGA-UniformTumor-8K?
TCGA-UniformTumor-8K dataset is a region-level pan-cancer subtyping resource comprising 25,495 ROIs of 8,192 × 8,192 pixels. These regions were extracted from 9,662 H&E-stained FFPE diagnostic histopathology WSIs sourced from TCGA. The tumor regions were manually annotated by two expert pathologists, with slide exclusion due to poor staining, poor focus, lacking cancerous regions and incorrect… See the full description on the dataset page: https://huggingface.co/datasets/MahmoodLab/TCGA-UniformTumor-8K.Arabic_Flicker_8kTaur_CoT_Analysis_Project___microsoft__Phi-3-small-8k-instruct
