datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
encodec_24khz-opt-125m-pretrained-ft-librispeech_asr
Dataset Card for "encodec_24khz-opt-125m-pretrained-ft-librispeech_asr"
More Information needed
encodec_24khz-opt-125m-pretrained-ft-librispeech_asr-train.clean.100-features
Dataset Card for "encodec_24khz-opt-125m-pretrained-ft-librispeech_asr-train.clean.100-features"
More Information needed
OxfordPets_test_facebook_opt_125m_Visclues_ns_3669
Dataset Card for "OxfordPets_test_facebook_opt_125m_Visclues_ns_3669"
More Information needed
encodec_24khz-opt-125m-pretrained-ft-librispeech_asr_dummy-validation-features
Dataset Card for "encodec_24khz-opt-125m-pretrained-ft-librispeech_asr_dummy-validation-features"
More Information needed
pretokenized__HuggingFaceFW_fineweb-edu__EleutherAI__gpt-neo-125mOxfordPets_test_facebook_opt_125m_Attributes_ns_3669
Dataset Card for "OxfordPets_test_facebook_opt_125m_Attributes_ns_3669"
More Information needed
Caltech101_not_background_test_facebook_opt_125m_Attributes_ns_5647
Dataset Card for "Caltech101_not_background_test_facebook_opt_125m_Attributes_ns_5647"
More Information needed
EleutherAI__gpt-neo-125m-details
Dataset Card for Evaluation run of EleutherAI/gpt-neo-125m
Dataset automatically created during the evaluation run of model EleutherAI/gpt-neo-125m
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/EleutherAI__gpt-neo-125m-details.dolly-15k-clustered-fulltext-less-sweep-gpt-neo-125M_p0.0-20250109dolly-15k-clustered-fulltext-less-sweep-gpt-neo-125M_p0.4-20250109dolly-15k-clustered-fulltext-less-sweep-gpt-neo-125M_p0.6-20250109OxfordPets_test_facebook_opt_125m_Attributes_Caption_ns_3669
Dataset Card for "OxfordPets_test_facebook_opt_125m_Attributes_Caption_ns_3669"
More Information needed
pythia_125M_inference_trainslm-125m-raft-dataset
slm-125m RAFT dataset
23,830 RAFT examples derived from the QA set. Each: question + several ~200-token document
chunks -> answer. 70% positive (golden chunk + 3 distractors), 30% negative (no golden ->
"That is not stated in the context."). Chat schema (messages + meta). Files: raft_train.jsonl (23,354), raft_val.jsonl (476).
Caltech101_not_background_test_facebook_opt_125m_Visclues_ns_5647
Dataset Card for "Caltech101_not_background_test_facebook_opt_125m_Visclues_ns_5647"
More Information needed
dolly-15k-clustered-fulltext-less-sweep-gpt-neo-125M_p1.0-20250109Caltech101_not_background_test_facebook_opt_125m_Attributes_Caption_ns_5647
Dataset Card for "Caltech101_not_background_test_facebook_opt_125m_Attributes_Caption_ns_5647"
More Information needed
dolly-15k-clustered-fulltext-less-sweep-gpt-neo-125M_p0.8-2025010927-11-MobileLLM-125Mag_news-mia_ag_news_client9slm-125m-qa-dataset
slm-125m QA dataset (SFT)
24,713 reviewed grounded-QA pairs (LLM-judged, kept >=4) over US case law, SEC filings,
and educational web text. Columns: source, source_file, type, question, answer, context,
llm_judge_score, llm_judge_verdict, llm_judge_reason. Used to fine-tune Sudhanshu1985/slm-125m-sft.
Types: lookup / reasoning / unanswerable (refusals).
pythia_125M_inference_testpythia_synthetic_125M_inference_testencodec_24khz-opt-125m-pretrained-ft-librispeech_asr-validation.clean-features
Dataset Card for "encodec_24khz-opt-125m-pretrained-ft-librispeech_asr-validation.clean-features"
More Information needed
27-11-MobileLLM-125Mag_news-mia_ag_news_client6pythia_synthetic_125M_inference_test_humanencodec_24khz-opt-125m-pretrained-ft-librispeech_asr-test.clean-features
Dataset Card for "encodec_24khz-opt-125m-pretrained-ft-librispeech_asr-test.clean-features"
More Information needed
27-11-MobileLLM-125Mxsum-mia_xsum_client7encodec_24khz-opt-125m-lm_pretraining_ls960_1qt-librispeech_asr-test.clean-features
Dataset Card for "encodec_24khz-opt-125m-lm_pretraining_ls960_1qt-librispeech_asr-test.clean-features"
More Information needed
27-11-MobileLLM-125Mxsum-mia_xsum_client8pythia-125M-test-gen
