datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CrystalXRD-Bench
LLM4Mat-Bench: XRD Max-Peak HKL Identification Benchmark
Dataset Description
LLM4Mat-Bench is a multimodal benchmark for evaluating Vision-Language Models (VLMs) on crystallographic reasoning tasks. Given a theoretical X-ray Diffraction (XRD) pattern image and the corresponding crystal structure (CIF), the model must identify the Miller indices (HKL) of the crystallographic planes contributing to the highest-intensity peak.
This benchmark tests the intersection of visual… See the full description on the dataset page: https://huggingface.co/datasets/xiaodu-ali/CrystalXRD-Bench.Code-feedback-sharegpt-renamedMoD-150k
Introduction
I'm excited to share the MoD 150k subset, a selection from the broader Mixture of Data project I've been working on. This subset is crafted for those looking to fine-tune AI models for both Mixture of Experts (MoE) architectures and standard architectures, with a keen eye on accessibility for those with limited computational resources.
My Experimentation
After diving deep into MoEs and conducting various experiments, I've found this 150k subset not only… See the full description on the dataset page: https://huggingface.co/datasets/Crystalcareai/MoD-150k.Crystal-CleanedMoD
Please note this is a dataset that accompanies the model; https://huggingface.co/Crystalcareai/Qwen1.5-8x7b. The readme is the same for both, with more detail below
Hey, I'm Lucas
I'm excited to share an early release of a project that has kept me busy for the last couple of weeks. Mixtral's release propelled me into a deep dive into MoEs.
With the release of Qwen1.5, I was curious to see how it would compare to Mixtral.
Coming from a background as an acting teacher and… See the full description on the dataset page: https://huggingface.co/datasets/Crystalcareai/MoD.CodeFeedback-AlpacaArcEasy-ExplainChoiceorca-cohereru_en_Crystallography_and_Spectroscopyopenhermes_200k_unfilteredpromptengineerSynthetic-Weakaura-DatasetCrystal-tenebrisIntroducing Crystal-tenebris, part of Project Crystal to collect massive amounts of data from the dark-web and release it as public dataset. Tenebris is a reddit alternative for dark web.
it is created under Project Crystal Blue
truthyDPO-intelFrom https://huggingface.co/jondurbin - I just renamed one of the columns to make axolotl happier.
Truthy DPO
This is a dataset designed to enhance the overall truthfulness of LLMs, without sacrificing immersion when roleplaying as a human.
For example, in normal AI assistant model, the model should not try to describe what the warmth of the sun feels like, but if the system prompt indicates it's a human, it should.
Mostly targets corporeal, spacial, temporal awareness, and common… See the full description on the dataset page: https://huggingface.co/datasets/Crystalcareai/truthyDPO-intel.Teknium-Openhermes-2.5-500k-TRLSelf-Discover-MM-Instruct-AlpacaNatural-Instructions-Small-AlpacaIntel-DPO-Pairs-Norefusalsjondurbin-airoboros-3.2-unfilteredArgilla-OpenHermes-2.5-DPO-Intelsynthetic_reasoning_natural_Alpaca_CombinedMoD-AlpacaOrca-RekaSelf-Discover-MM-InstructThis dataset was synthetically generated using the Mistral Medium model for a project I am currently developing. It draws inspiration from the Self-Discover framework outlined in a paper by Google Deepmind 1. While this implementation is a basic interpretation and does not fully capture the essence of the original framework, it resulted in a robust Instruct dataset that meets the project's objectives. Further details will be shared upon the project's release. Below is the Python code utilized… See the full description on the dataset page: https://huggingface.co/datasets/Crystalcareai/Self-Discover-MM-Instruct.WATOP600slimorca-dedup-alpaca-100kAESCOpenHermes2.5-CodeFeedback-Alpacadistilabel-intel-orca-dpo-pairs_intel_formatOH2.5strict
