datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Asclepius-Synthetic-Clinical-Notes
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is official dataset for Asclepius (arxiv)
This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs.
We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5
Then, we generate instruction-answer pairs for 157k synthetic discharge summaries
Supported Tasks
This dataset covers below 8 tasks
Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/starmpcc/Asclepius-Synthetic-Clinical-Notes.amazon_zhthis is a datasets about amazon reviews
amazon_zh_simpleascend
Dataset Card for ASCEND
Dataset Summary
ASCEND (A Spontaneous Chinese-English Dataset) introduces a high-quality resource of spontaneous multi-turn conversational dialogue Chinese-English code-switching corpus collected in Hong Kong. ASCEND consists of 10.62 hours of spontaneous speech with a total of ~12.3K utterances. The corpus is split into 3 sets: training, validation, and test with a ratio of 8:1:1 while maintaining a balanced gender proportion on each set.… See the full description on the dataset page: https://huggingface.co/datasets/filwsyl/ascend.billalpaca-star-asciisame as the original alpaca star, this one however encourages to include a mental image.
it will out put a flow chart or ascii image for each prompt
Dataset-AsCas12f
Description
This dataset contains signle site mutation of protein AsCas12f amino acid sequence and the correspond mutation effect score from a deep mutation scanning experiment.
Protein Format: SA sequence (AF2)
Splits
traing: 6306
valid: 801
test: 834
Related paper
The dataset is from An AsCas12f-based compact genome-editing tool derived by deep mutational scanning and structural analysis.
Label
Label means fitness score of each mutant… See the full description on the dataset page: https://huggingface.co/datasets/SaProtHub/Dataset-AsCas12f.Ascent-GenTascii_art_dataset_for_llmsAsclepius-Synthetic-Clinical-Notes
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is official dataset for Asclepius (arxiv)
This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs.
We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5
Then, we generate instruction-answer pairs for 157k synthetic discharge summaries
Supported Tasks
This dataset covers below 8 tasks
Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/Xavier1234/Asclepius-Synthetic-Clinical-Notes.ASCAT-Arabic-Scientific-Translation
ASCAT: Arabic Scientific Corpus for Advanced Translation
ASCAT (Arabic Scientific Corpus for Advanced Translation) is a high-quality English–Arabic parallel corpus of full scientific abstracts designed for rigorous evaluation and training of domain-specific machine translation (MT) systems.
Unlike existing Arabic–English corpora that rely on short sentences or narrow domains, ASCAT targets long-form scientific abstracts validated through a multi-engine translation and expert… See the full description on the dataset page: https://huggingface.co/datasets/NAMAA-Space/ASCAT-Arabic-Scientific-Translation.mix_infoAsclepius-Synthetic-Clinical-Notes
Asclepius: Synthetic Clincal Notes & Instruction Dataset
Dataset Summary
This dataset is official dataset for Asclepius (arxiv)
This dataset is composed with Clinical Note - Question - Answer format to build a clinical LLMs.
We first synthesized synthetic notes from PMC-Patients case reports with GPT-3.5
Then, we generate instruction-answer pairs for 157k synthetic discharge summaries
Supported Tasks
This dataset covers below 8 tasks
Named Entity… See the full description on the dataset page: https://huggingface.co/datasets/LampsteR/Asclepius-Synthetic-Clinical-Notes.ASCII-BenchDebateLLMsAscenseurs
[!NOTE]
Dataset origin: https://termini.gov.lv/kolekcijas/130
ascii-arttwice_korfin-asc_sent_cls
FinascSent-CLS-ko
Sentiment classification of text in financial reports (POSITIVE / NEUTRAL / NEGATIVE).
Dataset Details
Dataset Description
A task that classifies the sentiment of SRC (POSITIVE / NEUTRAL / NEGATIVE) based on KLUE-TC and Naver Finance analysis reports, considering different aspects.
Dataset Creation
Source Data
amphora/korfin-asc (Original Source : KLUE-TC and analyst reports from Naver Finance)
upiterbarg-lintseq-reproductionVirBiCla-training
Dataset Card for VirBiCla-training
VirBiCla is a ML-based viral DNA detector designed for long-read sequencing metagenomics.
This dataset is a support dataset for training the base ML model.
Dataset Details
Dataset Sources [optional]
Repository: GitHub repository for VirBiCla
Uses
This dataset is intended as support for training the base VirBiCla model
Dataset Structure
Dataset is a CSV file composed of 60.003 record sequences (coming… See the full description on the dataset page: https://huggingface.co/datasets/as-cle-bert/VirBiCla-training.ASCII-Bench-Litearchitecture_vs_normal_image_prompts
