datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
einstein-speech-corpus
🌌 Albert Einstein Lifetime Scientific Lectures & Pacifist Speeches Corpus (1909–1955)
Historical Eras Distribution
The Miracle Year & General Relativity (1909–1920): 60 foundational lectures & academy addresses
Nobel Prize & Global Relativity Lectures (1921–1932): 80 Princeton lectures, Nobel oration, Solvay debates, and world tours
Princeton IAS & Anti-Fascism (1933–1945): 60 Royal Albert Hall farewell, IAS seminars, and Roosevelt atomic letter
Nuclear… See the full description on the dataset page: https://huggingface.co/datasets/sanjeevafk/einstein-speech-corpus.mlc-mlsw-melspects
MLCommons Multilingual Spoken Words Mel-Spectograms
This dataset contains all English words from the dataset available at MLCommons (or also available on huggingface). These audio files have been processed into Mel spectrograms for downstream usage in DCNNs or similar processes.
Dataset description
There's a total of 6624343 samples of Mel spectograms. There are a total of 38150 different words, the cls is the index of that word in alphabetical order. With every entry… See the full description on the dataset page: https://huggingface.co/datasets/einstein8612/mlc-mlsw-melspects.einstein_answers
What would Einstein Say?
This dataset contains a set of questions and answers, mimicking Einstein's approach to answer general scientific and philosophical queries.
The data points have been generated synthetically, however the factual correctness of the data is ensured, not guaranteed whatsoever.
Weyaxi__Einstein-v4-7Beinstein_mindset_25k_dataset
Einstein Mindset Training Dataset (25k)
A high-quality synthetic dataset designed to instill Albert Einstein's distinctive thinking patterns, voice, and philosophical mindset into large language models through fine-tuning.
Overview
This dataset contains 25,000 instruction-response pairs crafted to train models to reason and respond in the style of Albert Einstein — emphasizing:
Profound curiosity and relentless questioning
The supremacy of imagination over rote… See the full description on the dataset page: https://huggingface.co/datasets/11-47/einstein_mindset_25k_dataset.Weyaxi__Einstein-v6.1-developed-by-Weyaxi-Llama3-8BWeyaxi__Einstein-v6.1-Llama3-8BWeyaxi__Einstein-v8-Llama3.2-1BWeyaxi__Einstein-v4-7B-details
Dataset Card for Evaluation run of Weyaxi/Einstein-v4-7B
Dataset automatically created during the evaluation run of model Weyaxi/Einstein-v4-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An additional… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Weyaxi__Einstein-v4-7B-details.details_Weyaxi__Einstein-v7-Qwen2-7B
Dataset Card for Evaluation run of Weyaxi/Einstein-v7-Qwen2-7B
Dataset automatically created during the evaluation run of model Weyaxi/Einstein-v7-Qwen2-7B.
The dataset is composed of 136 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/OALL/details_Weyaxi__Einstein-v7-Qwen2-7B.Weyaxi__Einstein-v7-Qwen2-7B-details
Dataset Card for Evaluation run of Weyaxi/Einstein-v7-Qwen2-7B
Dataset automatically created during the evaluation run of model Weyaxi/Einstein-v7-Qwen2-7B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.
An… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Weyaxi__Einstein-v7-Qwen2-7B-details.Weyaxi__Einstein-v6.1-Llama3-8B-details
Dataset Card for Evaluation run of Weyaxi/Einstein-v6.1-Llama3-8B
Dataset automatically created during the evaluation run of model Weyaxi/Einstein-v6.1-Llama3-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Weyaxi__Einstein-v6.1-Llama3-8B-details.Weyaxi__Einstein-v8-Llama3.2-1B-details
Dataset Card for Evaluation run of Weyaxi/Einstein-v8-Llama3.2-1B
Dataset automatically created during the evaluation run of model Weyaxi/Einstein-v8-Llama3.2-1B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Weyaxi__Einstein-v8-Llama3.2-1B-details.testingWeyaxi__Einstein-v6.1-developed-by-Weyaxi-Llama3-8B-details
Dataset Card for Evaluation run of Weyaxi/Einstein-v6.1-developed-by-Weyaxi-Llama3-8B
Dataset automatically created during the evaluation run of model Weyaxi/Einstein-v6.1-developed-by-Weyaxi-Llama3-8B
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Weyaxi__Einstein-v6.1-developed-by-Weyaxi-Llama3-8B-details.
