Orion vault — clean initial history
Knowledge vault (Orion/PARA) migrated from the pre-Orion 484vault on 2026-10-01. Deliberately orphaned: prior history contained a plaintext password and stays local-only on branch archive/pre-boilerplate-history. Secrets and live Hermes state are gitignored.
This commit is contained in:
commit
a66996ac10
233 changed files with 103810 additions and 0 deletions
41
2. Areas/Home Personal/knowledge/H200 VRAM Efficiency.md
Normal file
41
2. Areas/Home Personal/knowledge/H200 VRAM Efficiency.md
Normal file
|
|
@ -0,0 +1,41 @@
|
|||
|
||||
is there somethig more efficient than lama.cpp
|
||||
|
||||
Yes. With an H200, `llama.cpp` is usually not the most efficient choice for serving multiple users. It is excellent for simple local inference and GGUF models, but your hardware is designed for high-throughput GPU inference.
|
||||
|
||||
### Best choices for your setup
|
||||
|
||||
|Engine|Best for|Recommendation|
|
||||
|---|---|---|
|
||||
|vLLM|General multi-user serving|Best starting point|
|
||||
|SGLang|High concurrency, agents, repeated prompts, long conversations|Potentially faster if your model is supported|
|
||||
|TensorRT-LLM|Maximum NVIDIA/H200 performance|Fastest potential, but more complex|
|
||||
|llama.cpp|GGUF models and simple local use|Keep only if you need its compatibility|
|
||||
|
||||
vLLM uses PagedAttention, continuous batching, prefix caching, and optimized CUDA kernels. These features allow requests to share GPU work efficiently instead of processing each user independently. [vllm.ai](https://docs.vllm.ai/)
|
||||
|
||||
TensorRT-LLM is the NVIDIA-specific option. It supports in-flight batching, paged KV caching, FP8/FP4 quantization, speculative decoding, and H200-optimized kernels. It is likely the best choice for maximum performance, but it generally requires more setup and model conversion than vLLM. nvidia.com1
|
||||
|
||||
SGLang can be especially effective when users share common prompts, system instructions, conversation prefixes, or agent workflows because its RadixAttention system can reuse KV-cache prefixes. arXiv2
|
||||
|
||||
My recommendation would be:
|
||||
|
||||
1. Start with vLLM.
|
||||
2. Test six simultaneous users.
|
||||
3. Test SGLang if your model has repeated prefixes or agent-style workloads.
|
||||
4. Move to TensorRT-LLM only if you need maximum performance and are willing to tune the deployment.
|
||||
|
||||
A basic vLLM launch might look like:
|
||||
|
||||
bash
|
||||
|
||||
```
|
||||
vllm serve /path/to/model \ --host 0.0.0.0 \ --port 8000 \ --max-model-len 8192 \ --max-num-seqs 6 \ --gpu-memory-utilization 0.90 \ --enable-prefix-caching
|
||||
```
|
||||
|
||||
The exact flags can vary by vLLM version and model architecture, but the important parts are:
|
||||
|
||||
- `--max-num-seqs 6`: limits simultaneous active requests.
|
||||
- `--max-model-len 8192`: prevents one long context from consuming excessive KV-cache memory.
|
||||
- `--gpu-memory-utilization 0.90`: leaves some VRAM for CUDA and server overhead.
|
||||
- `--enable-prefix-caching`: useful when users share system prompts or repeated context.
|
||||
Loading…
Add table
Add a link
Reference in a new issue