CoolFace
22 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01starmpcc /Asclepius-Synthetic-Clinical-Notes Asclepius: Synthetic Clincal Notes & Instruction Dataset Dataset Summary This dataset is official dataset for Asclepius (arxiv) This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs. We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5 Then, we generate instruction-answer pairs for 157k synthetic discharge summaries Supported Tasks This dataset covers below 8 tasks Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes.textquestion-answering100K<n<1M117 likes967 downloads2y agoHugging Face02ASCCCCCCCC /amazon_zhthis is a datasets about amazon reviews text100K<n<1M2 likes260 downloads5y agoHugging Face03ASCCCCCCCC /amazon_zh_simpletext10K<n<100K1 likes159 downloads5y agoHugging Face04filwsyl /ascend Dataset Card for ASCEND Dataset Summary ASCEND (A Spontaneous Chinese-English Dataset) introduces a high-quality resource of spontaneous multi-turn conversational dialogue Chinese-English code-switching corpus collected in Hong Kong. ASCEND consists of 10.62 hours of spontaneous speech with a total of ~12.3K utterances. The corpus is split into 3 sets: training, validation, and test with a ratio of 8:1:1 while maintaining a balanced gender proportion on each set.… See the full description on the dataset page: https://huggingface.co/datasets/filwsyl/ascend.audioautomatic-speech-recognition1K<n<10K1 likes48 downloads4y agoHugging Face05ASCCCCCCCC /billtext10K<n<100K0 likes37 downloads4y agoHugging Face06gate369 /alpaca-star-asciisame as the original alpaca star, this one however encourages to include a mental image. it will out put a flow chart or ascii image for each prompt textn<1K5 likes37 downloads2y agoHugging Face07SaProtHub /Dataset-AsCas12f Description This dataset contains signle site mutation of protein AsCas12f amino acid sequence and the correspond mutation effect score from a deep mutation scanning experiment. Protein Format: SA sequence (AF2) Splits traing: 6306 valid: 801 test: 834 Related paper The dataset is from An AsCas12f-based compact genome-editing tool derived by deep mutational scanning and structural analysis. Label Label means fitness score of each mutant… See the full description on the dataset page: https://huggingface.co/datasets/SaProtHub/Dataset-AsCas12f.text1K<n<10K1 likes28 downloads2y agoHugging Face08Aunsiels /Ascent-GenTtextquestion-answering10M<n<100M1 likes27 downloads3y agoHugging Face09Sanjay1905 /ascii_art_dataset_for_llmstext100K<n<1M1 likes26 downloads1y agoHugging Face10Xavier1234 /Asclepius-Synthetic-Clinical-Notes Asclepius: Synthetic Clincal Notes & Instruction Dataset Dataset Summary This dataset is official dataset for Asclepius (arxiv) This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs. We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5 Then, we generate instruction-answer pairs for 157k synthetic discharge summaries Supported Tasks This dataset covers below 8 tasks Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/Xavier1234/Asclepius-Synthetic-Clinical-Notes.textquestion-answering100K<n<1M1 likes22 downloads5mo agoHugging Face11NAMAA-Space /ASCAT-Arabic-Scientific-Translation ASCAT: Arabic Scientific Corpus for Advanced Translation ASCAT (Arabic Scientific Corpus for Advanced Translation) is a high-quality English–Arabic parallel corpus of full scientific abstracts designed for rigorous evaluation and training of domain-specific machine translation (MT) systems. Unlike existing Arabic–English corpora that rely on short sentences or narrow domains, ASCAT targets long-form scientific abstracts validated through a multi-engine translation and expert… See the full description on the dataset page: https://huggingface.co/datasets/NAMAA-Space/ASCAT-Arabic-Scientific-Translation.tabulartranslationn<1K1 likes16 downloads6mo agoHugging Face12ASCCCCCCCC /mix_infotext10K<n<100K0 likes14 downloads4y agoHugging Face13LampsteR /Asclepius-Synthetic-Clinical-Notes Asclepius: Synthetic Clincal Notes & Instruction Dataset Dataset Summary This dataset is official dataset for Asclepius (arxiv) This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs. We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5 Then, we generate instruction-answer pairs for 157k synthetic discharge summaries Supported Tasks This dataset covers below 8 tasks Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/LampsteR/Asclepius-Synthetic-Clinical-Notes.textquestion-answering100K<n<1M0 likes14 downloads10mo agoHugging Face14haizelabs /ASCII-Benchtextn<1K0 likes13 downloads1y agoHugging Face15as-cle-bert /DebateLLMstextn<1K4 likes10 downloads2y agoHugging Face16FrancophonIA /Ascenseurs [!NOTE] Dataset origin: https://termini.gov.lv/kolekcijas/130 tabulartranslationn<1K0 likes10 downloads1y agoHugging Face17HusseinEssam /ascii-arttext1K<n<10K1 likes9 downloads1y agoHugging Face18nmixx-fin /twice_korfin-asc_sent_cls FinascSent-CLS-ko Sentiment classification of text in financial reports (POSITIVE / NEUTRAL / NEGATIVE). Dataset Details Dataset Description A task that classifies the sentiment of SRC (POSITIVE / NEUTRAL / NEGATIVE) based on KLUE-TC and Naver Finance analysis reports, considering different aspects. Dataset Creation Source Data amphora/korfin-asc (Original Source : KLUE-TC and analyst reports from Naver Finance) texttext-classification1K<n<10K1 likes7 downloads2y agoHugging Face19a-scarlett /upiterbarg-lintseq-reproductiontabular10K<n<100K0 likes7 downloads1y agoHugging Face20as-cle-bert /VirBiCla-training Dataset Card for VirBiCla-training VirBiCla is a ML-based viral DNA detector designed for long-read sequencing metagenomics. This dataset is a support dataset for training the base ML model. Dataset Details Dataset Sources [optional] Repository: GitHub repository for VirBiCla Uses This dataset is intended as support for training the base VirBiCla model Dataset Structure Dataset is a CSV file composed of 60.003 record sequences (coming… See the full description on the dataset page: https://huggingface.co/datasets/as-cle-bert/VirBiCla-training.tabular10K<n<100K1 likes6 downloads3y agoHugging Face21haizelabs /ASCII-Bench-Litetextn<1K0 likes5 downloads1y agoHugging Face22as-cle-bert /architecture_vs_normal_image_promptstext1K<n<10K2 likes4 downloads2y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.