mindset
Datasets
All datasets matching “mindset”Machine_Mindset_MBTI_datasetHere are the behavior datasets used for supervised fine-tuning (SFT). And they can also be used for direct preference optimization (DPO).
The exact copy can also be found in Github.
Prefix 'en' denotes the datasets of the English version.
Prefix 'zh' denotes the datasets of the Chinese version.
Dataset introduction
There are four dimension in MBTI. And there are two opposite attributes within each dimension.
To be specific:
Energe: Extraversion (E) - Introversion (I)… See the full description on the dataset page: https://huggingface.co/datasets/pandalla/Machine_Mindset_MBTI_dataset.MINDSET
MINDSET
MINDSET is the pretraining dataset for MIND, a coordinate-only location encoder distilled from static location encoder teachers and annual AlphaEarth Foundations (AEF) embeddings.
We release the embeddings at the 12.1M training coordinates. The dataset contains 12,099,072 land coordinates in WGS84. Coordinates are dense around cities and not uniformly sampled over land.
The files are in GeoParquet format and can be joined on point_id:
file
grain
rows
columns… See the full description on the dataset page: https://huggingface.co/datasets/taylor-geospatial/MINDSET.mindset_datasets_5x25k
5 × 25k Mindset Datasets (Instruction-Tuning JSONL)
These datasets are synthetic instruction-tuning conversations designed to train a model to emulate working mindsets associated with:
Benjamin Franklin
Thomas Edison
Albert Einstein
Nikola Tesla
Leonardo da Vinci
Format
Each line is a JSON object with:
id, person, category
messages: [system, user, assistant]
developer: Within Us AI
Notes
Content is written as original paraphrased training… See the full description on the dataset page: https://huggingface.co/datasets/11-47/mindset_datasets_5x25k.Five_Phases_Mindset_datasetsWelcome to our Traditional Chinese Medicine (TCM) Consultation Dataset! This dataset contains approximately one hundred thousand TCM consultation dialogue records, aiming to provide a rich resource for research and development in the field of TCM. These dialogue data cover various TCM diseases, diagnoses, and treatment methods, serving as an important reference for TCM research and clinical practice.
The dataset was created using a method that combines manual annotation with extraction from… See the full description on the dataset page: https://huggingface.co/datasets/cookey39/Five_Phases_Mindset_datasets.huberman_lab_Dr__David_Yeager_How_to_Master_Growth_Mindset_to_Improve_Performanceeinstein_mindset_25k_dataset
Einstein Mindset Training Dataset (25k)
A high-quality synthetic dataset designed to instill Albert Einstein's distinctive thinking patterns, voice, and philosophical mindset into large language models through fine-tuning.
Overview
This dataset contains 25,000 instruction-response pairs crafted to train models to reason and respond in the style of Albert Einstein — emphasizing:
Profound curiosity and relentless questioning
The supremacy of imagination over rote… See the full description on the dataset page: https://huggingface.co/datasets/11-47/einstein_mindset_25k_dataset.
