datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
market-brief-data-hub
Market Brief Dataset Hub (ship with V3 model)
Companion data for Market Brief Adapter V3 (Mixtral-8x7B LoRA, win rate 0.7806)Adaption AutoScientist Challenge 2026 · Market Analysis & News
Story
Artifact
Role
Model V3
Ship — grounded briefs + richer TrueNorth-enriched train path + Mixtral
This hub
Full data program: pilot → synthetic → TrueNorth → merged → Adaptive Data exports
V3 training used Adaptive Data on the merged seed (8e47668b… /… See the full description on the dataset page: https://huggingface.co/datasets/kongclaves/market-brief-data-hub.odaigen_hindi_pre_trained_spHindi Language Pre-Trained LLM Datasets Overview
Welcome to the Hindi Language Pre-Training Datasets repository! This README provides a comprehensive overview of various pre-training datasets available for Hindi, including essential details such as licenses, sources, and statistical information. These datasets are invaluable resources for training and fine-tuning large language models (LLMs) for a wide range of natural language processing (NLP) tasks.
-Data Overview and Statistics
This README… See the full description on the dataset page: https://huggingface.co/datasets/Hindi-data-hub/odaigen_hindi_pre_trained_sp.dataHub_schema_jsonlrino-huberman-data-model
