datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
einstein-mpc-spectra
Einstein MPC Spectra
The Einstein Monitor Proportional Counter (MPC) release contains 605 observed spectra and their 605 backgrounds, accumulated over the intervals of good MPC data overlapping SSS observations. The pairing is quasi-simultaneous, not simultaneous.
Use
from datasets import load_dataset
ds = load_dataset("astro-legacy-archive/einstein-mpc-spectra", "a053526.pha", split="train")
good = ds.filter(lambda row: row["QUALITY"] == 0)
row = good[0]… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/einstein-mpc-spectra.einstein-sss-spectra
Einstein SSS Spectra
The Einstein Solid State Spectrometer (SSS) release contains source, background and correction spectra. It covers the 0.5–4.5 keV instrument band and includes both 128-channel SSS records and eight-channel quasi-simultaneous MPC products distributed in the SSS directory.
Use
from datasets import load_dataset
ds = load_dataset("astro-legacy-archive/einstein-sss-spectra", "sa053526.pha", split="train")
good = ds.filter(lambda row:… See the full description on the dataset page: https://huggingface.co/datasets/astro-legacy-archive/einstein-sss-spectra.details_giannisan__penny5-dolphin-einstein-llama3-dare-ties-chatmldetails_Weyaxi__Einstein-openchat-7B
Dataset Card for Evaluation run of Weyaxi/Einstein-openchat-7B
Dataset automatically created during the evaluation run of model Weyaxi/Einstein-openchat-7B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Weyaxi__Einstein-openchat-7B.details_Weyaxi__Einstein-7B
Dataset Card for Evaluation run of Weyaxi/Einstein-7B
Dataset automatically created during the evaluation run of model Weyaxi/Einstein-7B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest results.… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Weyaxi__Einstein-7B.einstein-speech-corpus
🌌 Albert Einstein Lifetime Scientific Lectures & Pacifist Speeches Corpus (1909–1955)
Historical Eras Distribution
The Miracle Year & General Relativity (1909–1920): 60 foundational lectures & academy addresses
Nobel Prize & Global Relativity Lectures (1921–1932): 80 Princeton lectures, Nobel oration, Solvay debates, and world tours
Princeton IAS & Anti-Fascism (1933–1945): 60 Royal Albert Hall farewell, IAS seminars, and Roosevelt atomic letter
Nuclear… See the full description on the dataset page: https://huggingface.co/datasets/sanjeevafk/einstein-speech-corpus.details_Weyaxi__Einstein-v4-7B
Dataset Card for Evaluation run of Weyaxi/Einstein-v4-7B
Dataset automatically created during the evaluation run of model Weyaxi/Einstein-v4-7B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Weyaxi__Einstein-v4-7B.details_PulsarAI__Einstein-v3-7B
Dataset Card for Evaluation run of Weyaxi/Einstein-v3-7B
Dataset automatically created during the evaluation run of model Weyaxi/Einstein-v3-7B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_PulsarAI__Einstein-v3-7B.details_Weyaxi__Einstein-v6-7B
Dataset Card for Evaluation run of Weyaxi/Einstein-v6-7B
Dataset automatically created during the evaluation run of model Weyaxi/Einstein-v6-7B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 2 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Weyaxi__Einstein-v6-7B.details_Weyaxi__Einstein-v6.1-Llama3-8B
Dataset Card for Evaluation run of Weyaxi/Einstein-v6.1-Llama3-8B
Dataset automatically created during the evaluation run of model Weyaxi/Einstein-v6.1-Llama3-8B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Weyaxi__Einstein-v6.1-Llama3-8B.details_Weyaxi__Einstein-v4-Qwen-1.5-32BEinstein-Puzzles-Data
Einstein-Puzzles
Communication and Verification in LLM Agents towards Collaboration under Information Asymmetry (Arxiv)
Run Peng*, Ziqiao Ma*, Amy Pang, Sikai Li, Zhang Xi-Jia, Yingzhuo Yu, Cristian-Paul Bara, Joyce Chai
Dataset Details
There are four *.jsonl files under train/ folder, which corresponds to training data for models with four different communicative action spaces. The chain-of-thought reasoning traces are generated by gpt4o given the current game state and… See the full description on the dataset page: https://huggingface.co/datasets/Roihn/Einstein-Puzzles-Data.mlc-mlsw-melspects
MLCommons Multilingual Spoken Words Mel-Spectograms
This dataset contains all English words from the dataset available at MLCommons (or also available on huggingface). These audio files have been processed into Mel spectrograms for downstream usage in DCNNs or similar processes.
Dataset description
There's a total of 6624343 samples of Mel spectograms. There are a total of 38150 different words, the cls is the index of that word in alphabetical order. With every entry… See the full description on the dataset page: https://huggingface.co/datasets/einstein8612/mlc-mlsw-melspects.einstein_answers
What would Einstein Say?
This dataset contains a set of questions and answers, mimicking Einstein's approach to answer general scientific and philosophical queries.
The data points have been generated synthetically, however the factual correctness of the data is ensured, not guaranteed whatsoever.
details_Weyaxi__Einstein-bagel-7B
Dataset Card for Evaluation run of Weyaxi/Einstein-bagel-7B
Dataset automatically created during the evaluation run of model Weyaxi/Einstein-bagel-7B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Weyaxi__Einstein-bagel-7B.details_Weyaxi__Einstein-v5-v0.2-7B
Dataset Card for Evaluation run of Weyaxi/Einstein-v5-v0.2-7B
Dataset automatically created during the evaluation run of model Weyaxi/Einstein-v5-v0.2-7B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Weyaxi__Einstein-v5-v0.2-7B.details_Weyaxi__Einstein-v6.1-phi2
Dataset Card for Evaluation run of Weyaxi/Einstein-v6.1-phi2
Dataset automatically created during the evaluation run of model Weyaxi/Einstein-v6.1-phi2 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Weyaxi__Einstein-v6.1-phi2.details_mvpmaster__Einstein-4D-Marcoro14-7b-full-slerp
Dataset Card for Evaluation run of mvpmaster/Einstein-4D-Marcoro14-7b-full-slerp
Dataset automatically created during the evaluation run of model mvpmaster/Einstein-4D-Marcoro14-7b-full-slerp on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_mvpmaster__Einstein-4D-Marcoro14-7b-full-slerp.details_mvpmaster__Einstein-4d-Marcoro14-nddmpk-KrishnaHercules-7b-slerp
Dataset Card for Evaluation run of mvpmaster/Einstein-4d-Marcoro14-nddmpk-KrishnaHercules-7b-slerp
Dataset automatically created during the evaluation run of model mvpmaster/Einstein-4d-Marcoro14-nddmpk-KrishnaHercules-7b-slerp on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_mvpmaster__Einstein-4d-Marcoro14-nddmpk-KrishnaHercules-7b-slerp.Einstein_Telsa_Inventor_100kdetails_mvpmaster__Einstein-4D-MoE-2x7b-test
Dataset Card for Evaluation run of mvpmaster/Einstein-4D-MoE-2x7b-test
Dataset automatically created during the evaluation run of model mvpmaster/Einstein-4D-MoE-2x7b-test on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_mvpmaster__Einstein-4D-MoE-2x7b-test.details_TitleOS__EinsteinBagel-8BWeyaxi__Einstein-v4-7Beinstein_mindset_25k_dataset
Einstein Mindset Training Dataset (25k)
A high-quality synthetic dataset designed to instill Albert Einstein's distinctive thinking patterns, voice, and philosophical mindset into large language models through fine-tuning.
Overview
This dataset contains 25,000 instruction-response pairs crafted to train models to reason and respond in the style of Albert Einstein — emphasizing:
Profound curiosity and relentless questioning
The supremacy of imagination over rote… See the full description on the dataset page: https://huggingface.co/datasets/11-47/einstein_mindset_25k_dataset.details_Weyaxi__Einstein-v4-phi2
Dataset Card for Evaluation run of Weyaxi/Einstein-v4-phi2
Dataset automatically created during the evaluation run of model Weyaxi/Einstein-v4-phi2 on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Weyaxi__Einstein-v4-phi2.details_Weyaxi__einstein-v2-test-model
Dataset Card for Evaluation run of Weyaxi/Einstein-v2-7B
Dataset automatically created during the evaluation run of model Weyaxi/Einstein-v2-7B on the Open LLM Leaderboard.
The dataset is composed of 63 configuration, each one coresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the latest… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard-old/details_Weyaxi__einstein-v2-test-model.Einstein_Telsa_Inventor_MOE_100kWeyaxi__Einstein-v6.1-developed-by-Weyaxi-Llama3-8BEinstein_Telsa_Inventor_rationalize_100kWeyaxi__Einstein-v6.1-Llama3-8B
