Skip to content

About

Custom inference engine to run Minimax-M2.X series of models on dual RTX Pro 6000 (sm120)

Resources

Stars

5 stars

Watchers

1 watching

Forks

Latest commit

 

History

37 Commits

Folders and files

Repository files navigation

minimax-rs

A Rust/Candle CUDA inference server for MiniMax-M2.7 split GGUF weights. It provides an OpenAI-compatible API subset for a workstation with two Blackwell GPUs.

Requirements

  • Two NVIDIA RTX PRO 6000 Blackwell Workstation Edition GPUs (sm_120)
  • CUDA 13.3 with a 595-series driver that provides CUDA 13.2 compatibility
  • A Rust toolchain with Rust 2024 support
  • Exactly four regular .gguf shard files in the model directory

The model graph is fixed to 62 layers, four shards, and two GPUs. Multi-machine inference is not supported.

The server does not download or distribute model weights. The current quantization uses about 128 GB for weights and can take several minutes to load.

Run

export MINIMAX_MODEL_DIR="<directory-containing-the-four-gguf-shards>"
CUDA_COMPUTE_CAP=120f cargo run --release

Use --model DIR to override MINIMAX_MODEL_DIR:

CUDA_COMPUTE_CAP=120f cargo run --release -- --model /path/to/model

Parallelism

Mode Layout Use
pipeline 31 layers per GPU Default, best measured overall performance
tensor One worker per GPU, TP=2, head-sharded KV caches Lower decode latency

Select tensor mode with --parallelism tensor or MINIMAX_PARALLELISM=tensor:

CUDA_COMPUTE_CAP=120f cargo run --release -- --parallelism tensor

Both modes use the same API schemas.

Dry run checks shard metadata, the tokenizer, and GPU startup without loading weights or starting HTTP:

CUDA_COMPUTE_CAP=120f cargo run --release -- --dry-run
CUDA_COMPUTE_CAP=120f cargo run --release -- --parallelism tensor --dry-run

The server listens on 127.0.0.1:8000 by default and has no authentication or TLS. Use --host 0.0.0.0:8000 for remote access. Do not expose the server to an untrusted network.

Sampling defaults to temperature=1.0, top_p=0.95, and top_k=40. Override them with --temp/--temperature, --top-p, --top-k, or request fields. Use temperature=0 for greedy decoding. Use top_p=1 or top_k=0 to disable those filters.

API

  • GET /health
  • GET /v1/models
  • POST /v1/completions
  • POST /v1/chat/completions
  • POST /v1/perplexity
curl --fail-with-body -sS http://127.0.0.1:8000/v1/chat/completions \
  -H 'content-type: application/json' \
  -d '{
    "messages": [{"role": "user", "content": "Hello"}],
    "max_tokens": 64,
    "temperature": 0
  }'

Supported behavior:

  • Completion prompts accept one string or one token-ID array. Batches are rejected.
  • Chat supports a leading system or developer message, prior tool calls, and tool results.
  • SSE streams content, reasoning_content, and OpenAI tool-call deltas.
  • tool_choice supports auto, none, required, or a named function. Arguments use JSON Schema validation.
  • Native MiniMax XML tool calls are converted to OpenAI tool calls.
  • stop accepts one string or up to four strings, plus configured EOG tokens.
  • Completions and chat accept max_tokens. Chat also accepts max_completion_tokens.
  • The context limit is 196,608 tokens. Output defaults to 131,072 tokens and is clipped to the remaining context.
  • Responses report KV reuse in usage.prompt_tokens_details.cached_tokens.

POST /v1/perplexity accepts the same completion prompt formats. It clears the shared cache, applies teacher forcing, and returns token IDs with ln p(token[i] | token[:i]). The first token is context and is not scored.

Requests are serialized through one model mutex and share one longest-prefix KV cache. An unrelated request can replace the reusable prefix.

Implementation

  • Prompt prefill uses 512-token chunks and FlashAttention over the retained KV prefix. Set MINIMAX_PREFILL_CHUNK from 1 through 2,048.
  • Q4_K/Q5_K expert prefill uses direct Q8_1 MMQ at 32 or more tokens. Shorter batches use the generic parallel kernel.
  • Other supported GGUF expert types use the repaired F16 WMMA path at 192 or more tokens.
  • Tensor mode keeps CUDA state in two long-lived rank workers. Both workers must report the same cache length for each cache operation.
  • A rank failure stops both workers and makes the server unavailable. Rank 0 performs greedy selection on-device and stochastic sampling in-process.
  • Each generation logs prompt, cache, prefill, decode, latency, throughput, and finish reason.

Warmed target results:

Workload Pipeline Tensor
5-token prompt and 100-token decode 125.7 tokens/s 140.2 tokens/s
512-token prefill 229 ms 381 ms
513-token prefill 232 ms 392 ms
Mixed 125-request matrix 30.2 s 36.4 s

Test

Run the default suite:

cargo fmt
cargo check
cargo test

These tests do not load weights or initialize CUDA, but compilation requires the configured CUDA toolchain.

Run tests that require CUDA or model weights:

export MINIMAX_MODEL_DIR=/path/to/model
cargo test gguf_tokenizer_round_trip -- --ignored
cargo test --test tp_collective -- --ignored --nocapture
cargo test --test tp_layer -- --ignored --nocapture
cargo test --test gqa_prefill -- --ignored --nocapture

Validate a running server:

./scripts/test-completion.sh
./scripts/test-server.py
./scripts/test-server.py --stress-cycles 25

Set BASE_URL for test-completion.sh or pass --base-url to test-server.py when the server uses another address.

Compare perplexity with one server mode loaded at a time. Run this command in pipeline mode:

./scripts/validate-ppl.py corpus.txt --output pipeline-ppl.json

Restart in tensor mode, then run:

./scripts/validate-ppl.py corpus.txt \
  --reference pipeline-ppl.json \
  --output tensor-ppl.json

Use --token-ids for a JSON token-ID corpus. Optional gates are --max-relative-ppl-delta, --max-mean-abs-token-nll-delta, and --max-p99-abs-token-nll-delta.

License

Code outside vendor/ uses the Apache License 2.0. Copyright 2026 Stephen Murray.

Vendored crates keep their upstream licenses. Model weights use the MiniMax-M2.7 model license.

About

Custom inference engine to run Minimax-M2.X series of models on dual RTX Pro 6000 (sm120)

Resources

Stars

5 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages