axion
Datasets
All datasets matching “axion”imagenet-r
ImageNet-R
This repo is made to facilitate the evaluation of various pretraining models. It's constructed from the source file provided by official implementation.
Usage
from datasets import load_dataset
dataset = load_dataset('axiong/imagenet-r')
Dataset Summary
ImageNet-R(endition) contains art, cartoons, deviantart, graffiti, embroidery, graphics, origami, paintings, patterns, plastic objects, plush objects, sculptures, sketches, tattoos, toys, and video… See the full description on the dataset page: https://huggingface.co/datasets/axiong/imagenet-r.pmc_oaFoundation models trained on large-scale dataset gain a recent surge in CV and NLP. In contrast, development in biomedical domain lags far behind due to data scarcity.
To address this issue, we build and release PMC-OA, a biomedical dataset with 1.6M image-caption pairs collected from PubMedCentral's OpenAccess subset, which is 8 times larger than before.
PMC-OA covers diverse modalities or diseases, with majority of the image-caption samples aligned at finer-grained level, i.e., subfigure and subcaption.
While pretraining a CLIP-style model on PMC-OA, our model named PMC-CLIP achieves state-of-the-art results on various downstream tasks,
including image-text retrieval on ROCO, MedMNIST image classification, Medical VQA, i.e. +8.1% R@10 on image-text retrieval, +3.9% accuracy on image classification.pmc_llama_instructionsThis repo provides part of the dataset used for PMC-LLaMA-13B's instruction tuning.
Data
Size
Link
ChatDoctor
100K
https://www.yunxiangli.top/ChatDoctor/
MedQA
10.2K
https://huggingface.co/datasets/GBaker/MedQA-USMLE-4-options
MedMCQA
183K
https://huggingface.co/datasets/medmcqa
PubmedQA
211K
https://huggingface.co/datasets/pubmed_qa
LiveQA
635
https://huggingface.co/datasets/truehealth/liveqa
MedicationQA
690
https://huggingface.co/datasets/truehealth/medicationqa
UMLS… See the full description on the dataset page: https://huggingface.co/datasets/axiong/pmc_llama_instructions.AxionLimitBench
AxionLimitBench
A benchmark for extracting experimental exclusion limits from particle-physics papers.
Given a paper that bounds an axion, dark-photon or other light-boson coupling, a system
must decide whether the paper reports a new measured limit, identify the coupling, and
return the excluded-region boundary as a curve in a canonical (mass, coupling) plane.
Each answer is graded against the curve the maintainer of the
AxionLimits compilation committed for that paper.
292… See the full description on the dataset page: https://huggingface.co/datasets/FaroutYLq/AxionLimitBench.pmc_oa_demoFoundation models trained on large-scale dataset gain a recent surge in CV and NLP. In contrast, development in biomedical domain lags far behind due to data scarcity.
To address this issue, we build and release PMC-OA, a biomedical dataset with 1.6M image-caption pairs collected from PubMedCentral's OpenAccess subset, which is 8 times larger than before.
PMC-OA covers diverse modalities or diseases, with majority of the image-caption samples aligned at finer-grained level, i.e., subfigure and subcaption.
While pretraining a CLIP-style model on PMC-OA, our model named PMC-CLIP achieves state-of-the-art results on various downstream tasks,
including image-text retrieval on ROCO, MedMNIST image classification, Medical VQA, i.e. +8.1% R@10 on image-text retrieval, +3.9% accuracy on image classification.ThinkSet-PTBR
🧠 ThinkSet-PTBR
A synthetic dataset for training small language models to simulate structured reasoning in Portuguese.
🚀 Overview
ThinkSet-PTBR is a synthetic dataset designed to train and evaluate reasoning behavior in small-scale language models.
It focuses on structured reasoning patterns rather than raw knowledge, making it especially suitable for:
Small models (1M–50M parameters)
CPU-based training environments
Research on reasoning simulation
💡… See the full description on the dataset page: https://huggingface.co/datasets/AxionLab-Co/ThinkSet-PTBR.
