A Rust/Candle CUDA inference server for MiniMax-M2.7 split GGUF weights. It provides an OpenAI-compatible API subset for a workstation with two Blackwell GPUs.
- Two NVIDIA RTX PRO 6000 Blackwell Workstation Edition GPUs (
sm_120) - CUDA 13.3 with a 595-series driver that provides CUDA 13.2 compatibility
- A Rust toolchain with Rust 2024 support
- Exactly four regular
.ggufshard files in the model directory
The model graph is fixed to 62 layers, four shards, and two GPUs. Multi-machine inference is not supported.
The server does not download or distribute model weights. The current quantization uses about 128 GB for weights and can take several minutes to load.
export MINIMAX_MODEL_DIR="<directory-containing-the-four-gguf-shards>"
CUDA_COMPUTE_CAP=120f cargo run --releaseUse --model DIR to override MINIMAX_MODEL_DIR:
CUDA_COMPUTE_CAP=120f cargo run --release -- --model /path/to/model| Mode | Layout | Use |
|---|---|---|
pipeline |
31 layers per GPU | Default, best measured overall performance |
tensor |
One worker per GPU, TP=2, head-sharded KV caches | Lower decode latency |
Select tensor mode with --parallelism tensor or MINIMAX_PARALLELISM=tensor:
CUDA_COMPUTE_CAP=120f cargo run --release -- --parallelism tensorBoth modes use the same API schemas.
Dry run checks shard metadata, the tokenizer, and GPU startup without loading weights or starting HTTP:
CUDA_COMPUTE_CAP=120f cargo run --release -- --dry-run
CUDA_COMPUTE_CAP=120f cargo run --release -- --parallelism tensor --dry-runThe server listens on 127.0.0.1:8000 by default and has no authentication or TLS. Use --host 0.0.0.0:8000 for remote access. Do not expose the server to an untrusted network.
Sampling defaults to temperature=1.0, top_p=0.95, and top_k=40. Override them with --temp/--temperature, --top-p, --top-k, or request fields. Use temperature=0 for greedy decoding. Use top_p=1 or top_k=0 to disable those filters.
GET /healthGET /v1/modelsPOST /v1/completionsPOST /v1/chat/completionsPOST /v1/perplexity
curl --fail-with-body -sS http://127.0.0.1:8000/v1/chat/completions \
-H 'content-type: application/json' \
-d '{
"messages": [{"role": "user", "content": "Hello"}],
"max_tokens": 64,
"temperature": 0
}'Supported behavior:
- Completion prompts accept one string or one token-ID array. Batches are rejected.
- Chat supports a leading
systemordevelopermessage, prior tool calls, and tool results. - SSE streams content,
reasoning_content, and OpenAI tool-call deltas. tool_choicesupportsauto,none,required, or a named function. Arguments use JSON Schema validation.- Native MiniMax XML tool calls are converted to OpenAI tool calls.
stopaccepts one string or up to four strings, plus configured EOG tokens.- Completions and chat accept
max_tokens. Chat also acceptsmax_completion_tokens. - The context limit is 196,608 tokens. Output defaults to 131,072 tokens and is clipped to the remaining context.
- Responses report KV reuse in
usage.prompt_tokens_details.cached_tokens.
POST /v1/perplexity accepts the same completion prompt formats. It clears the shared cache, applies teacher forcing, and returns token IDs with ln p(token[i] | token[:i]). The first token is context and is not scored.
Requests are serialized through one model mutex and share one longest-prefix KV cache. An unrelated request can replace the reusable prefix.
- Prompt prefill uses 512-token chunks and FlashAttention over the retained KV prefix. Set
MINIMAX_PREFILL_CHUNKfrom 1 through 2,048. - Q4_K/Q5_K expert prefill uses direct Q8_1 MMQ at 32 or more tokens. Shorter batches use the generic parallel kernel.
- Other supported GGUF expert types use the repaired F16 WMMA path at 192 or more tokens.
- Tensor mode keeps CUDA state in two long-lived rank workers. Both workers must report the same cache length for each cache operation.
- A rank failure stops both workers and makes the server unavailable. Rank 0 performs greedy selection on-device and stochastic sampling in-process.
- Each generation logs prompt, cache, prefill, decode, latency, throughput, and finish reason.
Warmed target results:
| Workload | Pipeline | Tensor |
|---|---|---|
| 5-token prompt and 100-token decode | 125.7 tokens/s | 140.2 tokens/s |
| 512-token prefill | 229 ms | 381 ms |
| 513-token prefill | 232 ms | 392 ms |
| Mixed 125-request matrix | 30.2 s | 36.4 s |
Run the default suite:
cargo fmt
cargo check
cargo testThese tests do not load weights or initialize CUDA, but compilation requires the configured CUDA toolchain.
Run tests that require CUDA or model weights:
export MINIMAX_MODEL_DIR=/path/to/model
cargo test gguf_tokenizer_round_trip -- --ignored
cargo test --test tp_collective -- --ignored --nocapture
cargo test --test tp_layer -- --ignored --nocapture
cargo test --test gqa_prefill -- --ignored --nocaptureValidate a running server:
./scripts/test-completion.sh
./scripts/test-server.py
./scripts/test-server.py --stress-cycles 25Set BASE_URL for test-completion.sh or pass --base-url to test-server.py when the server uses another address.
Compare perplexity with one server mode loaded at a time. Run this command in pipeline mode:
./scripts/validate-ppl.py corpus.txt --output pipeline-ppl.jsonRestart in tensor mode, then run:
./scripts/validate-ppl.py corpus.txt \
--reference pipeline-ppl.json \
--output tensor-ppl.jsonUse --token-ids for a JSON token-ID corpus. Optional gates are --max-relative-ppl-delta, --max-mean-abs-token-nll-delta, and --max-p99-abs-token-nll-delta.
Code outside vendor/ uses the Apache License 2.0. Copyright 2026 Stephen Murray.
Vendored crates keep their upstream licenses. Model weights use the MiniMax-M2.7 model license.