datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
meowcat-predictions
MeowCat cell-type predictions on TCGA-LUAD and CPTAC-CCRCC
Per-pixel cell-type predictions generated by MeowCat on
H&E whole-slide images from two public cohorts:
Cohort
Tissue
Samples
h5ad payload
TCGA-LUAD
Lung adenocarcinoma
531
~60 GB
CPTAC-CCRCC
Clear-cell renal cell carcinoma
831
~93 GB
File layout
composition.parquet # long format: sample × cell_type → count, fraction
metadata.parquet # sample_id, cohort, patient_id, n_pixels… See the full description on the dataset page: https://huggingface.co/datasets/liranmao/meowcat-predictions.trace-rx-eval-predictions
TRACE-RX Evaluation Predictions
Per-image detector scores from an independent evaluation of the two TechJam 2026 TRACE-RX
detectors, run 30 Aug – 1 Sep 2026.
No images here. Every file contains scores, labels, asset ids and transform names only — this is
derived evaluation metadata, not a redistribution of any source imagery. The underlying corpora
(Joshyxwa/data_draft, Joshyxwa/techjam2026, techjam-aigc/wildfake-eval-subset) keep their own
terms, and data_draft's WildFake rows… See the full description on the dataset page: https://huggingface.co/datasets/joelleoqiyi/trace-rx-eval-predictions.galaxy-vit-gz-desi-dirichlet-predictions
Galaxy-ViT — GZ DESI Dirichlet predictions
Per-galaxy Dirichlet-Multinomial concentration parameters (α) for the
10-question / 34-answer Galaxy Zoo DESI decision tree, predicted by a
Zoobot ConvNeXt-nano encoder finetuned with a Dirichlet-Multinomial
head on the DR8 subset of the
mwalmsley/gz_desi_wds
labeled split.
Dataset summary
Rows
61,440
Columns
36 (key, dr8_id, alpha_0 … alpha_33)
Format
Apache Parquet
File size
~16 MB
Source images
DECaLS DR8… See the full description on the dataset page: https://huggingface.co/datasets/roth1414/galaxy-vit-gz-desi-dirichlet-predictions.autotrain-data-image-attribute-prediction
AutoTrain Dataset for project: image-attribute-prediction
Dataset Description
This dataset has been automatically processed by AutoTrain for project image-attribute-prediction.
Languages
The BCP-47 code for the dataset's language is unk.
Dataset Structure
Data Instances
A sample from this dataset looks as follows:
[
{
"image": "<261x300 RGB PIL image>",
"target": 0
},
{
"image": "<300x300 RGB PIL image>",
"target": 0… See the full description on the dataset page: https://huggingface.co/datasets/snjv90/autotrain-data-image-attribute-prediction.
