CoolFace
Datasetpublic

echodict/llama.cpp

version https://git-lfs.github.com/spec/v1 oid sha256:cfc44b7ba25614df70e6b65e3341cae0310163bd32fd31a6b928a542df433faf size 30786

sourceHugging Faceupdated 5mo agoView on Hugging Face
0likes773downloads
README.md15 linesDownload Raw Back to parallel
1# llama.cpp/example/parallel2 3Simplified simulation of serving incoming requests in parallel4 5## Example6 7Generate 128 client requests (`-ns 128`), simulating 8 concurrent clients (`-np 8`). The system prompt is shared (`-pps`), meaning that it is computed once at the start. The client requests consist of up to 10 junk questions (`--junk 10`) followed by the actual question.8 9```bash10llama-parallel -m model.gguf -np 8 -ns 128 --top-k 1 -pps --junk 10 -c 1638411```12 13> [!NOTE]14> It's recommended to use base models with this example. Instruction tuned models might not be able to properly follow the custom chat template specified here, so the results might not be as expected.15