CoolFace
10 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01AlekseyKorshuk /product-photography-v1-tiny-prompts-tasks-collage-filteredimage1K<n<10K1 likes2.6k downloads3y agoHugging Face02CollagenHelixLabs /cdsm_benchmarking_data CDSM Collagen Structure Benchmark — Data Structures and scores for a benchmark comparing a deterministic collagen triple-helix builder (CDSM) against four co-folding models — Boltz-2, Chai-1, Protenix and AlphaFold3, the last in both with-MSA (af3_msa) and no-MSA (af3_nomsa) conditions — on 80 experimentally resolved collagen triple helices from the RCSB PDB. Code: https://github.com/bm-howard/cdsm_benchmarking Layout Prefix Contents Size experimental/… See the full description on the dataset page: https://huggingface.co/datasets/CollagenHelixLabs/cdsm_benchmarking_data.tabular10K<n<100K0 likes636 downloads21d agoHugging Face03m-Just /VisCoT_VStar_CollageThis is part of the training data for vSearcher introduced in "InSight-o3: Empowering Multimodal Foundation Models with Generalized Visual Search". The data comprise collages made from a subset of images from VisualCoT and the training data of V*. Each entry of this dataset contains a collage (with a randomly placed "core" image within it) and a QA for the core image. The other images are filler images sampled from the same image pool as the core images. Every image (both core and filler) is… See the full description on the dataset page: https://huggingface.co/datasets/m-Just/VisCoT_VStar_Collage.imagequestion-answering10K<n<100K3 likes345 downloads8mo agoHugging Face04CollageBio /oligogym OligoGym Datasets A mirror of the 12 curated benchmark datasets shipped with OligoGym (Roche), repackaged as HuggingFace configs for convenient loading. Covers oligonucleotide activity, toxicity, and immunomodulation across ASO, siRNA, and shRNA modalities. This is not a new dataset. Row counts, column sets, and values are identical to the files in OligoGym's oligogym/resources/pkg_dataset/. All credit for curation belongs to the OligoGym authors — please cite them (see… See the full description on the dataset page: https://huggingface.co/datasets/CollageBio/oligogym.tabulartabular-regression100K<n<1M0 likes164 downloads2mo agoHugging Face05jjobear /collage-layout-dataset Collage Layout Synthetic Dataset Synthetic photo-collage layouts for layout-quality analysis & correction, built on a six-ingredient framework (Format, Photos, Visual Weight, Hierarchy, Readability, Harmony). Corrector-not-generator: every collage carries a naive (v1_center) and a corrected (fit) placement, so a model can learn the correction. Faces are synthetically replaced (privacy-safe). How to load from datasets import load_dataset ds =… See the full description on the dataset page: https://huggingface.co/datasets/jjobear/collage-layout-dataset.imageimage-to-imagen<1K0 likes53 downloads2mo agoHugging Face06lamm-mit /collagen-cdsm_benchmarking_datagated CDSM Collagen Structure Benchmark — Data Structures and scores for a benchmark comparing a deterministic collagen triple-helix builder (CDSM) against four co-folding models — Boltz-2, Chai-1, Protenix and AlphaFold3, the last in both with-MSA (af3_msa) and no-MSA (af3_nomsa) conditions — on 80 experimentally resolved collagen triple helices from the RCSB PDB. Code: https://github.com/bm-howard/cdsm_benchmarking Layout Prefix Contents Size experimental/… See the full description on the dataset page: https://huggingface.co/datasets/lamm-mit/collagen-cdsm_benchmarking_data.tabular10K<n<100K0 likes37 downloads20d agoHugging Face07totally-not-an-llm /collage-40k Collage 40k Dataset Collage 40k is a dataset of approximately 40,000 user/assistant conversations generated by GPT-3.5 and GPT-4. This dataset was created by extracting filtered subsets from other datasets. Dataset Details Total Conversations: 39,819 GPTeacher general-instruct: 13,567 ShareGPT: 16,409 OpenOrca (step by step): 9,843 The dataset has undergone filtering to remove censorship, refusals, alignment, and low-quality conversations. The data is provided in… See the full description on the dataset page: https://huggingface.co/datasets/totally-not-an-llm/collage-40k.text10K<n<100K0 likes21 downloads3y agoHugging Face08CollageBio /aso-atlas Citation This dataset is derived from the ASO Atlas resource introduced in Hill et al. (2025). If you use this dataset, please cite: @article{hill2025accurately, title={Accurately modelling RNase H-mediated antisense oligonucleotide efficacy}, author={Hill, Barney and Jaques, Maisie R and Nair, Remya R and Whiffin, Nicola and Wood, Matthew JA and Sanders, Stephan J and Oliver, Peter L and Hill, Alyssa C and Rinaldi, Carlo}, journal={bioRxiv}, pages={2025--10}, year={2025}… See the full description on the dataset page: https://huggingface.co/datasets/CollageBio/aso-atlas.tabular100K<n<1M0 likes13 downloads5mo agoHugging Face09AlekseyKorshuk /product-photography-v1-tiny-prompts-tasks-collagegatedimage1K<n<10K0 likes5 downloads3y agoHugging Face10AlekseyKorshuk /product-photography-v1-tiny-prompts-tasks-collage-filtered-annotatedgatedimage1K<n<10K1 likes3 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.