datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
COIN-collection
COIN collection dataset
This is the dataset for our paper "Predicting the Encoding Error of SIRENs". It consists of 300,000 small SIREN networks trained to encode images from the MSCOCO dataset.
We will publish a loading script for this dataset soon, but until then, see the following instructions:
How to Use
First, download this repository using:
huggingface-cli repo download predict-SIREN-PSNR/COIN-collection --repo_type datasets
There are two types of files in this… See the full description on the dataset page: https://huggingface.co/datasets/predict-SIREN-PSNR/COIN-collection.siren-screening
SIREN Screening Dataset
~60,000 articles with synthetic relevance queries for training biomedical document screeners. Each article has multiple queries at three relevance levels: Relevant, Partial, Irrelevant.
Why this dataset?
Systematic reviews require screening thousands of articles against inclusion criteria (e.g., "RCTs in adults with diabetes, published after 2015"). Existing retrieval models (MedCPT, PubMedBERT)… See the full description on the dataset page: https://huggingface.co/datasets/Praise2112/siren-screening.sirena-0.1-qa
