datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
us-army-fm-instructThis is a multiturn instruct tuning dataset with 2,333,924 trainable tokens, created with Augmentoolkit, covering the material in the majority of the US Army Field Manuals that are publicly available.
Unlike many previous Augmentoolkit datasets, the questions and answers here are without fluff and are more "to the point". This "sharper" data is intended to help the LLM with recalling facts.
There are three main datasets included here: "vanilla", "negative" and "long".
Vanilla data is simple… See the full description on the dataset page: https://huggingface.co/datasets/Heralax/us-army-fm-instruct.US-Army-Survival-Sharegptus-army-fm-instructThis is a multiturn instruct tuning dataset with 2,333,924 trainable tokens, created with Augmentoolkit, covering the material in the majority of the US Army Field Manuals that are publicly available.
Unlike many previous Augmentoolkit datasets, the questions and answers here are without fluff and are more "to the point". This "sharper" data is intended to help the LLM with recalling facts.
There are three main datasets included here: "vanilla", "negative" and "long".
Vanilla data is simple… See the full description on the dataset page: https://huggingface.co/datasets/BadBerad5222/us-army-fm-instruct.us-army-public-fm-chunkedus-army-fm-instruct-cloneConvert the dataset from here to the hugginge face conversational format.
The data remains the same, the keys just change
Usage
./download-convert-fm-instruct.sh
Description
(copied from original dataset)
There are three main datasets included here: "vanilla", "negative" and "long".
Vanilla data is simple, where a human user asks a question and the AI answers it.
Negative data is meant to help the AI be a bit more robust: the user asks a misinformed, flawed, or… See the full description on the dataset page: https://huggingface.co/datasets/Bobby060/us-army-fm-instruct-clone.us-army-fm-text-cleanedus-army-fm-instructThis is a multiturn instruct tuning dataset with 2,333,924 trainable tokens, created with Augmentoolkit, covering the material in the majority of the US Army Field Manuals that are publicly available.
Unlike many previous Augmentoolkit datasets, the questions and answers here are without fluff and are more "to the point". This "sharper" data is intended to help the LLM with recalling facts.
There are three main datasets included here: "vanilla", "negative" and "long".
Vanilla data is simple… See the full description on the dataset page: https://huggingface.co/datasets/SpaceLr/us-army-fm-instruct.
