datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
vi-dataset-for-pretrain
Dataset Card for "vi-dataset-for-pretrain"
This is a combination of multiple Vietnamese dataset for pretraining CLMs such as GPT, GPT2, etc.
The dataset consists of:
vietgpt/covid_19_news_vi
hieunguyen1053/binhvq-news-corpus
oscar (unshuffled_deduplicated_vi)
vietgpt/wikipedia_vi
Dataset info
Splits
N.o examples
Size
Train
23,891,116
77.36 GB
Validation
1,257,428
4.06 GB
Total
25,148,544
81.43 GB
vi-dataset-for-pretrain
Dataset Card for "vi-dataset-for-pretrain"
This is a combination of multiple Vietnamese dataset for pretraining CLMs such as GPT, GPT2, etc.
The dataset consists of:
vietgpt/covid_19_news_vi
hieunguyen1053/binhvq-news-corpus
oscar (unshuffled_deduplicated_vi)
vietgpt/wikipedia_vi
Dataset info
Splits
N.o examples
Size
Train
23,891,116
77.36 GB
Validation
1,257,428
4.06 GB
Total
25,148,544
81.43 GB
signature-to-mechanisms
Signature-to-Mechanisms (S2M)
Signature-to-Mechanisms (S2M) provides standardized tasks designed to enable the assessment of mechanistic reasoning in AI agents. Each task supplies the elements necessary for evaluation, including experimental context, molecular signatures, and task prompts, so that agents can be tested on their ability to reconstruct mechanistic explanations reported in peer-reviewed biological studies.
S2M formalizes a core challenge in computational biology:… See the full description on the dataset page: https://huggingface.co/datasets/vida-nyu/signature-to-mechanisms.
