datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ciggml-vicuna-v0-quantizedThese are quantized ggml binary files for vicuna 7B and 13B models. The version of vicuna for these models are v0.
These files can be used in conjunction with minigpt4 ggml models 7B and 13B in minigpt4.cpp
Recommended are the Q5_K and Q6_K implementations. If there are any issues, use Q4_1 or Q4_0.
Vicuna Model Card
Model details
Model type:
Vicuna is an open-source chatbot trained by fine-tuning LLaMA on user-shared conversations collected from ShareGPT.
It is… See the full description on the dataset page: https://huggingface.co/datasets/maknee/ggml-vicuna-v0-quantized.Supra2-100M-tokens-backup
Supra2-100M — pretraining token bins
DO NOT USE. Internal artifact of the Supra2-100M training run, published only
because public storage is free. These are raw uint16 token bins for one specific
tokenizer (supra2-tokenizer.json) and carry the licenses of their upstream sources
(FineWeb-Edu, DCLM, The Stack v3, FineMath, Cosmopedia, Wikipedia). No redistribution
rights are granted or implied. Files may be deleted at any time.
Layout: tokens_XXX.bin (4B tok each), val_000.bin… See the full description on the dataset page: https://huggingface.co/datasets/GGMLGuy/Supra2-100M-tokens-backup.ggml-japanese-gpt2Windowsの方はggml-japanese-gpt2の実行ファイルで動くと思います。
同じファイル名のggml形式のbinとSentencePiece形式のmodelをダウンロードして保存してください。
使い方は以下のようになります。
gpt-2.exe -m ggml-model-japanese-gpt2-medium-f16.bin -p "こんにちは"
ggml-model-japanese-gpt2-xsmallのファイル形式がおかしくなっているので、ダウンロードしても動きません。
あとで修正します。
minigpt4-7b-ggmlThese are quantized ggml binary files for minigpt4 7B model.
These files can be used in conjunction with vicuna v0 ggml models to get minigpt4 working.
Not all implementations were tested. If there are any issues, use f16.
minigpt4-13b-ggmlThese are quantized ggml binary files for minigpt4 13B model.
These files can be used in conjunction with vicuna v0 ggml models to get minigpt4 working.
Not all implementations were tested. If there are any issues, use f16.
niv2_explanation_targets_h2ogpt-gm-oasst1-en-2048-falcon-40b-v2-GGML
Dataset Card for "niv2_explanation_targets_h2ogpt-gm-oasst1-en-2048-falcon-40b-v2-GGML"
More Information needed
LaMini-LM-dataset-TheBloke-h2ogpt-falcon-40b-v2-GGML-eval-llama2
Dataset Card for "LaMini-LM-dataset-TheBloke-h2ogpt-falcon-40b-v2-GGML"
More Information needed
LaMini-LM-dataset-TheBloke-h2ogpt-falcon-40b-v2-GGML-eval-llama2-gpt4
Dataset Card for "LaMini-LM-dataset-TheBloke-h2ogpt-falcon-40b-v2-GGML-eval-llama2-gpt4"
More Information needed
Llama-2-7B-Chat-GGML
Dataset Card for "Llama-2-7B-Chat-GGML"
More Information needed
cot_explanation_targets_h2ogpt-gm-oasst1-en-2048-falcon-40b-v2-GGML
Dataset Card for "cot_explanation_targets_h2ogpt-gm-oasst1-en-2048-falcon-40b-v2-GGML"
More Information needed
openhathi-7b-base-q4_0.ggmlThis dataset contains the ggml version of OpenHathi model released by Sarvam AI. Link to original model.
The ggml file provided is 4 bit quantized version, it can be run on local devices such as an M1 MacBook or other hardware.
How to use?
Download llama.cpp from here
git clone https://github.com/ggerganov/llama.cpp
Note: The ggml support has been deprecated, new file format is gguf. But since this repository contains ggml file, we have to switch back to an older commit of… See the full description on the dataset page: https://huggingface.co/datasets/sumitj39/openhathi-7b-base-q4_0.ggml.t0_explanation_targets_h2ogpt-gm-oasst1-en-2048-falcon-40b-v2-GGML
Dataset Card for "t0_explanation_targets_h2ogpt-gm-oasst1-en-2048-falcon-40b-v2-GGML"
More Information needed
LaMini-LM-raw-instruction-only-dataset-TheBloke-h2ogpt-falcon-40b-v2-GGML
Dataset Card for "LaMini-LM-raw-instruction-only-dataset-TheBloke-h2ogpt-falcon-40b-v2-GGML"
More Information needed
flan2021_explanation_targets_h2ogpt-gm-oasst1-en-2048-falcon-40b-v2-GGML
Dataset Card for "flan2021_explanation_targets_h2ogpt-gm-oasst1-en-2048-falcon-40b-v2-GGML"
More Information needed
ggmlwhisper-kannada-base-ggmlGGML-Touch-Teach-LGA1whisper-kannada-tiny-ggmlGGML-Touch-Teach-LGA0krishna_dummy-ggmlgguf-ggml-pad-overflow-poc
