CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01deepcoder2024 /cochrane-screening-sft Cochrane Screening SFT Supervised fine-tuning (SFT) chat dataset for Cochrane-style title and abstract screening. Each example is a chat conversation that asks a model to predict a screening decision (include / exclude / uncertain) and a short justification (reason). Code: ljwa2323/cochrane-screening-slm Dataset summary Split / config Records Role train 416,799 LoRA SFT training validation 46,311 Training-time validation (10% stratified holdout from… See the full description on the dataset page: https://huggingface.co/datasets/deepcoder2024/cochrane-screening-sft.texttext-classification100K<n<1M0 likes193 downloads1mo agoHugging Face02allenai /cochrane_sparse_maxThis is a copy of the Cochrane dataset, except the input source documents of its validation split have been replaced by a sparse retriever. The retrieval pipeline used: query: The target field of each example corpus: The union of all documents in the train, validation and test splits. A document is the concatenation of the title and abstract. retriever: BM25 via PyTerrier with default settings top-k strategy: "max", i.e. the number of documents retrieved, k, is set as the maximum number of… See the full description on the dataset page: https://huggingface.co/datasets/allenai/cochrane_sparse_max.textsummarization1K<n<10K0 likes58 downloads4y agoHugging Face03allenai /cochrane_dense_meanThis is a copy of the Cochrane dataset, except the input source documents of its train, validation and test splits have been replaced by a dense retriever. The retrieval pipeline used: query: The target field of each example corpus: The union of all documents in the train, validation and test splits. A document is the concatenation of the title and abstract. retriever: facebook/contriever-msmarco via PyTerrier with default settings top-k strategy: "max", i.e. the number of documents retrieved… See the full description on the dataset page: https://huggingface.co/datasets/allenai/cochrane_dense_mean.textsummarization1K<n<10K0 likes54 downloads4y agoHugging Face04ben-yu /cochrane_combinedtabular1K<n<10K0 likes50 downloads4y agoHugging Face05clinicalnlplab /CochranePLS_testtext1K<n<10K0 likes39 downloads3y agoHugging Face06clinicalnlplab /CochranePLS_1shot_testtext1K<n<10K3 likes34 downloads3y agoHugging Face07allenai /cochrane_dense_maxThis is a copy of the Cochrane dataset, except the input source documents of its validation split have been replaced by a dense retriever. The retrieval pipeline used: query: The target field of each example corpus: The union of all documents in the train, validation and test splits. A document is the concatenation of the title and abstract. retriever: facebook/contriever-msmarco via PyTerrier with default settings top-k strategy: "max", i.e. the number of documents retrieved, k, is set as… See the full description on the dataset page: https://huggingface.co/datasets/allenai/cochrane_dense_max.textsummarization1K<n<10K1 likes32 downloads4y agoHugging Face08allenai /cochrane_dense_oracleThis is a copy of the Cochrane dataset, except the input source documents of the train, validation, and test splits have been replaced by a dense retriever. query: The target field of each example corpus: The union of all documents in the train, validation and test splits. A document is the concatenation of the title and abstract. retriever: facebook/contriever-msmarco via PyTerrier with default settings top-k strategy: "oracle", i.e. the number of documents retrieved, k, is set as the… See the full description on the dataset page: https://huggingface.co/datasets/allenai/cochrane_dense_oracle.textsummarization1K<n<10K0 likes28 downloads4y agoHugging Face09allenai /cochrane_sparse_oracleThis is a copy of the Cochrane dataset, except the input source documents of its validation split have been replaced by a sparse retriever. The retrieval pipeline used: query: The target field of each example corpus: The union of all documents in the train, validation and test splits. A document is the concatenation of the title and abstract. retriever: BM25 via PyTerrier with default settings top-k strategy: "oracle", i.e. the number of documents retrieved, k, is set as the original number… See the full description on the dataset page: https://huggingface.co/datasets/allenai/cochrane_sparse_oracle.textsummarization1K<n<10K0 likes23 downloads4y agoHugging Face10clinicalnlplab /CochranePLS_traintext1K<n<10K0 likes22 downloads3y agoHugging Face11allenai /cochrane_sparse_meanThis is a copy of the Cochrane dataset, except the input source documents of its validation split have been replaced by a sparse retriever. The retrieval pipeline used: query: The target field of each example corpus: The union of all documents in the train, validation and test splits. A document is the concatenation of the title and abstract. retriever: BM25 via PyTerrier with default settings top-k strategy: "mean", i.e. the number of documents retrieved, k, is set as the mean number of… See the full description on the dataset page: https://huggingface.co/datasets/allenai/cochrane_sparse_mean.textsummarization1K<n<10K0 likes19 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.