datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
anything-v3.0-glazed
Dataset Card for Anything v3.0 Glazed Samples
Dataset Description
Dataset Summary
This dataset contains image samples originally generated by Linaqruf/anything-v3.0
and subsequently processed by Glaze tool.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/hanamizuki-ai/anything-v3.0-glazed.sphere_cohere_embed-english-v3.0wikipedia_0715_clean_cohere_embed-english-v3.0lm-eval-results-shyamieee-Padma-SLM-7b-v3.0-private
Dataset Card for Evaluation run of shyamieee/Padma-SLM-7b-v3.0
Dataset automatically created during the evaluation run of model shyamieee/Padma-SLM-7b-v3.0
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shyamieee-Padma-SLM-7b-v3.0-private.miracl-th-corpus-multilingual-v3.0stackexchange_cohere_embed-english-v3.0Hercules-v3.0
Hercules-v3.0
Dataset Name: Hercules-v3.0
Version: 3.0
Release Date: 2024-2-14
Number of Examples: 1,637,895
Domains: Math, Science, Biology, Physics, Instruction Following, Conversation, Computer Science, Roleplay, and more
Languages: Mostly English, but others can be detected.
Task Types: Question Answering, Conversational Modeling, Instruction Following, Code Generation, Roleplay
Data Source Description
Hercules-v3.0 is an extensive and diverse dataset that… See the full description on the dataset page: https://huggingface.co/datasets/Locutusque/Hercules-v3.0.hyperion-v3.0Hyperion-3.0 has significantly improved performance over its predecessors.
"I found that having more code datasets than general purpose datasets ironically decreases performance in both coding and general tasks."
Data sources:
OpenOrca/SlimOrca
cognitivecomputations/dolphin (300k examples)
microsoft/orca-math-word-problems-200k (60k examples)
glaiveai/glaive-code-assistant
Vezora/Tested-22k-Python-Alpaca
Unnatural Instructions
BI55/MedText
LDJnr/Pure-Dove
Various domain-specific datasets by… See the full description on the dataset page: https://huggingface.co/datasets/Locutusque/hyperion-v3.0.lm-eval-results-shyamieee-B3E3-SLM-7b-v3.0-private
Dataset Card for Evaluation run of shyamieee/B3E3-SLM-7b-v3.0
Dataset automatically created during the evaluation run of model shyamieee/B3E3-SLM-7b-v3.0
The dataset is composed of 62 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/nyu-dice-lab/lm-eval-results-shyamieee-B3E3-SLM-7b-v3.0-private.han-instruct-dataset-v3.0
Dataset Card for Han Instruct Dataset v3.0
The newest dataset version is https://huggingface.co/datasets/pythainlp/han-instruction-dataset.
🪿 Han (ห่าน or goose) Instruct Dataset is a Thai instruction dataset by PyThaiNLP. This dataset collects all Thai instruct datasets that were made by humans and our old model. The dataset can be used to train Instruction Following models like ChatGPT or others.
Many questions are collect from Reference desk at Thai wikipedia.
Data sources:… See the full description on the dataset page: https://huggingface.co/datasets/pythainlp/han-instruct-dataset-v3.0.twitter_customer_support_weaviate_export_200000_cohere-embed-multilingual-light-v3.0deltakv_qwen_train_v3.0_num40000_seqlen8192Tengentoppa-sft-v3.0alice_pick_HS_ROS-bag_0916_EEF_RPY_v3.0context-conditioned-molecule-transfer-v3.0-bioavailability-ma-vote-mean-intern
Bioavailability_Ma context-conditioned molecule transfer v3_0
This release is independent of the molecule-only V1.x family. Its node is a parent molecule under one exact selected external-condition signature.
train rows: 23,100
validation_ranking rows: 2,852
validation_ranking_random_balanced rows: 2,852
test_ranking rows: 22,682
test_ranking_random_balanced rows: 22,682
Validation: Morgan top-75 parents, then every context row
Test: Morgan top-75 parents, then every context… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/context-conditioned-molecule-transfer-v3.0-bioavailability-ma-vote-mean-intern.context-conditioned-molecule-transfer-v3.0-skin-reaction-vote-mean-intern
Skin_Reaction context-conditioned molecule transfer v3_0
This release is independent of the molecule-only V1.x family. Its node is a parent molecule under one exact selected external-condition signature.
train rows: 23,496
validation_ranking rows: 2,965
validation_ranking_random_balanced rows: 2,965
test_ranking rows: 18,697
test_ranking_random_balanced rows: 18,697
Validation: Morgan top-75 parents, then every context row
Test: Morgan top-75 parents, then every context row… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/context-conditioned-molecule-transfer-v3.0-skin-reaction-vote-mean-intern.ChunkedTextReasoning_v3.0createml-ome-src-v3.00
Orthogonal Model of Emotions
Abbreviated: OME
Author
C.J. Pitchford
Creation Date
Originally created 2016, first version published September, 2017, at Medium.
Version
v3
Base Model
Working with BERT as base
toxic-full-uncensored-v3.0Tengentoppa-sft-reasoning-v3.0molecule-property-transfer-v3.0-bioavailability-ma-vote-mean-intern
Bioavailability_Ma molecule property transfer v3_0
Molecule-level transfer prompts derived from processed Starling gold labels. Molecule A exposes its property value; Molecule B remains hidden.
Label mode: vote_mean
Training rows: 19,692
Validation ranking rows: 2,475
Test ranking rows: 15,675
Balanced random validation rows: 2,475
Balanced random test rows: 15,675
Evaluation geometry: vote_mean in native units
KNN regression MAE: absolute query error after averaging the top… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/molecule-property-transfer-v3.0-bioavailability-ma-vote-mean-intern.Diacritized-Case-Ending-Errors-TTT-V3.0AIME-2021-2025-Paraphrased-Gemini-2.5-Flash-v3.0Triangle104__Chatty-Harry_V3.0-details
Dataset Card for Evaluation run of Triangle104/Chatty-Harry_V3.0
Dataset automatically created during the evaluation run of model Triangle104/Chatty-Harry_V3.0
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Triangle104__Chatty-Harry_V3.0-details.duncan-gamabunta-v3.0wikimia24_hard_64_seed2-Paraphrased-Gemini-2.5-Flash-v3.0context-conditioned-molecule-transfer-v3.0-bbb-martins-vote-mean-intern
BBB_Martins context-conditioned molecule transfer v3_0
This release is independent of the molecule-only V1.x family. Its node is a parent molecule under one exact selected external-condition signature.
train rows: 35,808
validation_ranking rows: 4,483
validation_ranking_random_balanced rows: 4,483
test_ranking rows: 30,451
test_ranking_random_balanced rows: 30,451
Validation: Morgan top-75 parents, then every context row
Test: Morgan top-75 parents, then every context row… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/context-conditioned-molecule-transfer-v3.0-bbb-martins-vote-mean-intern.molecule-property-transfer-v3.0-bbb-martins-vote-mean-intern
BBB_Martins molecule property transfer v3_0
Molecule-level transfer prompts derived from processed Starling gold labels. Molecule A exposes its property value; Molecule B remains hidden.
Label mode: vote_mean
Training rows: 34,512
Validation ranking rows: 4,425
Test ranking rows: 27,450
Balanced random validation rows: 4,425
Balanced random test rows: 27,450
Evaluation geometry: vote_mean in native units
KNN regression MAE: absolute query error after averaging the top five… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/molecule-property-transfer-v3.0-bbb-martins-vote-mean-intern.molecule-property-transfer-v3.0-skin-reaction-vote-mean-intern
Skin_Reaction molecule property transfer v3_0
Molecule-level transfer prompts derived from processed Starling gold labels. Molecule A exposes its property value; Molecule B remains hidden.
Label mode: vote_mean
Training rows: 23,124
Validation ranking rows: 2,925
Test ranking rows: 18,375
Balanced random validation rows: 2,925
Balanced random test rows: 18,375
Evaluation geometry: vote_mean in native units
KNN regression MAE: absolute query error after averaging the top five… See the full description on the dataset page: https://huggingface.co/datasets/jiosephlee/molecule-property-transfer-v3.0-skin-reaction-vote-mean-intern.classif-v3.0
