Knowledge vault (Orion/PARA) migrated from the pre-Orion 484vault on 2026-10-01. Deliberately orphaned: prior history contained a plaintext password and stays local-only on branch archive/pre-boilerplate-history. Secrets and live Hermes state are gitignored.
41 lines
No EOL
2.3 KiB
Markdown
41 lines
No EOL
2.3 KiB
Markdown
|
|
is there somethig more efficient than lama.cpp
|
|
|
|
Yes. With an H200, `llama.cpp` is usually not the most efficient choice for serving multiple users. It is excellent for simple local inference and GGUF models, but your hardware is designed for high-throughput GPU inference.
|
|
|
|
### Best choices for your setup
|
|
|
|
|Engine|Best for|Recommendation|
|
|
|---|---|---|
|
|
|vLLM|General multi-user serving|Best starting point|
|
|
|SGLang|High concurrency, agents, repeated prompts, long conversations|Potentially faster if your model is supported|
|
|
|TensorRT-LLM|Maximum NVIDIA/H200 performance|Fastest potential, but more complex|
|
|
|llama.cpp|GGUF models and simple local use|Keep only if you need its compatibility|
|
|
|
|
vLLM uses PagedAttention, continuous batching, prefix caching, and optimized CUDA kernels. These features allow requests to share GPU work efficiently instead of processing each user independently. [vllm.ai](https://docs.vllm.ai/)
|
|
|
|
TensorRT-LLM is the NVIDIA-specific option. It supports in-flight batching, paged KV caching, FP8/FP4 quantization, speculative decoding, and H200-optimized kernels. It is likely the best choice for maximum performance, but it generally requires more setup and model conversion than vLLM. nvidia.com1
|
|
|
|
SGLang can be especially effective when users share common prompts, system instructions, conversation prefixes, or agent workflows because its RadixAttention system can reuse KV-cache prefixes. arXiv2
|
|
|
|
My recommendation would be:
|
|
|
|
1. Start with vLLM.
|
|
2. Test six simultaneous users.
|
|
3. Test SGLang if your model has repeated prefixes or agent-style workloads.
|
|
4. Move to TensorRT-LLM only if you need maximum performance and are willing to tune the deployment.
|
|
|
|
A basic vLLM launch might look like:
|
|
|
|
bash
|
|
|
|
```
|
|
vllm serve /path/to/model \ --host 0.0.0.0 \ --port 8000 \ --max-model-len 8192 \ --max-num-seqs 6 \ --gpu-memory-utilization 0.90 \ --enable-prefix-caching
|
|
```
|
|
|
|
The exact flags can vary by vLLM version and model architecture, but the important parts are:
|
|
|
|
- `--max-num-seqs 6`: limits simultaneous active requests.
|
|
- `--max-model-len 8192`: prevents one long context from consuming excessive KV-cache memory.
|
|
- `--gpu-memory-utilization 0.90`: leaves some VRAM for CUDA and server overhead.
|
|
- `--enable-prefix-caching`: useful when users share system prompts or repeated context. |