This guide explains how to configure and run TOLERANT Match in a cluster for production use.
It is written for administrators and operators.
Use this as the shortest production-oriented flow:
- Prepare at least 2 nodes (3 preferred), synchronized time (NTP), and node-to-node network access.
- Configure one shared cluster identity (
clusterName) and consistent cluster ports on all nodes. - Use one discovery method across all nodes (Kubernetes service discovery, static TCP list, or AWS discovery).
- Ensure each node has unique node identity and separate local data storage.
- Start nodes and verify all members join the same cluster and report expected
clusterSize. - Route traffic only to healthy nodes and only when project state is
SYNCHRONIZED. - Monitor synchronization and cluster metrics; investigate any
ERRORstate before full traffic.
Path selection:
- Kubernetes/Helm operators: start with sections 4.5, 4.7, 7, and 8.
- On-prem/custom operators: start with sections 4.4, 4.8, 4.9, 17.1, and 17.2.
A Match cluster runs two or more Match service nodes for:
- higher availability
- better request throughput
- controlled data synchronization between nodes
In a typical setup, clients call a load balancer, and the load balancer forwards requests to cluster nodes.
Use this baseline architecture:
- At least 2 Match nodes (3 preferred for production).
- One load balancer in front of Match nodes.
- Separate local data storage per node.
- Stable node identity and stable network naming.
- Time synchronization (NTP) on all nodes.
Important: cluster nodes must not share the same local data directory or database schema tables.
Do:
- keep Match versions identical across all cluster nodes
- keep cluster/discovery configuration consistent across all nodes
- keep local node data directories separated per node
- route production traffic only to projects in
SYNCHRONIZED
Don't:
- do not share one local data directory across nodes
- do not mix discovery methods unintentionally in the same cluster
- do not route full production traffic to
SYNCHRONIZINGorERRORproject states - do not change cluster identity/ports on only a subset of nodes
- do not keep Savepoint Settings on all nodes the same if copying the same config files between servers
Before enabling cluster mode, confirm:
- all nodes run the same Match version
- all nodes have compatible platform/runtime setup
- required node-to-node ports are open
- load balancer can reach all Match nodes
- all nodes can resolve each other by DNS/IP
- project configuration is aligned across nodes
Match uses Hazelcast configuration files for cluster discovery and communication.
Quick path selection:
Kubernetes/Helm: use chart-generatedhazelcast.yaml(recommended)On-prem/custom: use external Hazelcast config files, ormatchRuntimefallback if no Hazelcast config file exists
Place one of these files in the Match config directory:
hazelcast.yamlhazelcast.ymlhazelcast.xmlclusterconfig.yamlclusterconfig.ymlclusterconfig.xml
The first file found in this order is used.
Both YAML and XML are supported.
Use one format consistently across all nodes.
Common discovery modes:
- Kubernetes service discovery
- static TCP member list
- AWS discovery (for AWS profile deployments)
Use static member discovery when multicast or platform discovery is not suitable.
Example (hazelcast.yaml):
hazelcast:
cluster-name: match-cluster
network:
port:
port: 5701
join:
multicast:
enabled: false
tcp-ip:
enabled: true
member-list:
- 192.168.10.11:5701
- 192.168.10.12:5701
- 192.168.10.13:5701Guidelines:
- disable multicast when using TCP/IP member list
- list all expected cluster members
- keep cluster name and port consistent on all nodes
- ensure firewall rules allow member-to-member traffic
For Kubernetes, prefer chart-managed Hazelcast config and StatefulSet identity.
Example Helm values:
global:
db:
type: rocksdb
rocksdbCluster:
enabled: true
match:
environmentProfile: k8s
clusterSyncServicePort: 10100
hazelcast:
enabled: true
port: 5701
clusterName: match-cluster
joinMode: kubernetesNotes:
- use stable pod identity (StatefulSet)
- keep per-pod persistent storage
- keep readiness checks enabled and route traffic only to ready pods
Important:
- Match checks Hazelcast files in the lookup order from section 4.1 (
hazelcast.yaml,hazelcast.yml,hazelcast.xml,clusterconfig.yaml,clusterconfig.yml,clusterconfig.xml) - if a file is found, that file is used for Hazelcast cluster configuration
matchRuntimecluster attributes are used only when no Hazelcast configuration file was found
In Helm/Kubernetes deployments:
- when
global.match.hazelcast.enabled=true, the chart generates and mountshazelcast.yaml - this generated file is the effective runtime cluster configuration
- use Helm values (for example
global.match.hazelcast.joinMode,global.match.hazelcast.tcpMembers,global.match.hazelcast.aws.*, ports, and cluster name) instead of manually editing externalclusterconfig.*files
For on-prem installations, cluster mode is activated by setting cluster identity in matchRuntime.
Minimum activation:
- set
clusterNamein<matchRuntime ...>
All service instances with the same clusterName are treated as one cluster.
These attributes are used as fallback baseline cluster settings only if no Hazelcast config file is found:
clusterName: shared cluster identity across nodesclusterPort: port for cluster member communicationclusterSyncServicePort: port for synchronization data exchange serviceclusterSyncServiceEncrKey(optional): enables encrypted sync payload exchange when set
Example:
<matchRuntime clusterName="cl1" clusterPort="5701" clusterSyncServicePort="10100" />For advanced network/member settings, use Hazelcast configuration files.
You can use software or hardware load balancers.
A reverse proxy setup (for example, Apache/Nginx) is common.
Recommendations:
- route client traffic only to healthy Match nodes
- keep health checks active
- use session stickiness only if your architecture requires it
- for clustered projects, route full traffic only when project state is
SYNCHRONIZED
Configure consistent cluster identity and ports across nodes:
- cluster name (shared by all nodes in one cluster)
- node name / service identity (unique per node)
- cluster sync service port
- Hazelcast/member communication port
If values differ unintentionally between nodes, synchronization issues may occur.
For Helm-based deployments:
- keep
global.match.environmentProfile: k8s(default) - enable Hazelcast for clustered synchronization
- run multi-replica setups as StatefulSet for stable identity
- keep per-node persistent storage
For RocksDB multi-replica deployments:
- enable RocksDB cluster mode
- enable Hazelcast
- keep per-pod PVCs
For SQL deployments (Postgres/MariaDB):
- use node-specific service IDs/prefixes
- keep SQL connection settings valid for all replicas
Check cluster status from the Match admin interface/API/CLI.
Typical project cluster states:
SYNCHRONIZING: node is catching up and should not receive production trafficSYNCHRONIZED: node is ready for normal trafficERROR: synchronization/configuration problem; investigate before routing traffic
Operational rule: only route full traffic to nodes in SYNCHRONIZED.
The following example shows the readiness of a cluster with four members:
~ $ service.sh backend --endpoint health --function readiness
_____ ___ _ ___ ___ _ _ _ _____ __ __ _ _
|_ _/ _ \| | | __| _ \ /_\ | \| |_ _| | \/ |__ _| |_ __| |_
| || (_) | |__| _|| / / _ \| .` | | | | |\/| / _` | _/ _| ' \
|_| \___/|____|___|_|_\/_/ \_\_|\_| |_| |_| |_\__,_|\__\__|_||_|
Version 12.1.58634, Copyright (c) 2026 TOLERANT Software GmbH & Co KG
{
"name": "match",
"status": "UP",
"details": {
"projects": {
"name": "match",
"status": "UP",
"details": {
"matchProject-1": {
"name": "match",
"status": "UP",
"details": {
"errorMessage": "",
"cluster-information": {
"cluster-project-state": "SYNCHRONIZED",
"cluster-mode": "MASTER_SLAVE",
"cluster-size": 4,
"cluster-node-name": "tolerant-match-smoke-match-1",
"cluster-uuid": "153d789f-3dde-4cc9-a1cc-24d92a901241",
"cluster-name": "match-cluster"
},
"componentStatus": "RUNNING",
"active": true,
"configurationTime": "2026-05-22T11:34:43.033157072"
}
}
}
},
"compositeDiscoveryClient()": {
"name": "match",
"status": "UP",
"details": {
"services": {}
}
},
"readiness": {
"name": "match",
"status": "UP",
"details": {
"runtimeState": "RUNNING",
"projectCount": 1,
"allProjectsRunning": true,
"allClusterProjectsSynchronized": true,
"unsynchronizedClusterProjects": []
}
},
"service": {
"name": "match",
"status": "UP"
},
"diskSpace": {
"name": "match",
"status": "UP",
"details": {
"total": 1081101176832,
"free": 1003819864064,
"threshold": 10485760
}
},
"runtime": {
"name": "match",
"status": "UP",
"details": {
"cluster-information": {
"cluster-name": "match-cluster",
"cluster-member-state": "CONNECTED",
"cluster-node-name": "tolerant-match-smoke-match-1",
"cluster-UUID": "153d789f-3dde-4cc9-a1cc-24d92a901241"
},
"errorMessage": "",
"componentStatus": "RUNNING",
"configurationTime": "2026-05-22T11:34:38.426613761"
}
},
"gracefulShutdown": {
"name": "match",
"status": "UP",
"details": {
"activeTasks": 1
}
}
}
}
Processing successful (RC: 0, Elapsed time: 1s)The same result can be achieved by using the following http request:
wget -q -O- http://<Servername>:<Serverport>/health/readinessThe following metrics give you a quick view over the cluster state:
- cluster.member.state
- cluster.project.member-count
- cluster.project.state
State code mapping (cluster.member.state / ClusterMemberState):
0: not available / cluster disabled1:UNCONFIGURED,CONFIGURED,CONNECTED,SPLIT_BRAIN4:ERROR
State code mapping (cluster.project.state / ClusterProjectState):
0: not available / cluster disabled1:UNCONFIGURED,CONFIGURED_ACTIVE,CONFIGURED_INACTIVE2:SYNCHRONIZING,SYNCHRONIZING_WAITING3:SYNCHRONIZED4:ERROR
Example:
~ $ service.sh backend --endpoint metrics --function cluster.project.member-count
_____ ___ _ ___ ___ _ _ _ _____ __ __ _ _
|_ _/ _ \| | | __| _ \ /_\ | \| |_ _| | \/ |__ _| |_ __| |_
| || (_) | |__| _|| / / _ \| .` | | | | |\/| / _` | _/ _| ' \
|_| \___/|____|___|_|_\/_/ \_\_|\_| |_| |_| |_\__,_|\__\__|_||_|
Version 12.1.58634, Copyright (c) 2026 TOLERANT Software GmbH & Co KG
{
"name": "cluster.project.member-count",
"measurements": [{
"statistic": "VALUE",
"value": 4
}],
"availableTags": [{
"tag": "projectId",
"values": ["matchProject-1"]
}],
"description": "Cluster member count for this project (0 when project is not clustered)"
}
Processing successful (RC: 0, Elapsed time: 1s)When a node restarts or rejoins:
- it finds a suitable peer
- it synchronizes missing data
- it becomes
SYNCHRONIZED
Synchronization can use backlog/history when available, which is usually faster than full transfer.
If history is insufficient, Match can use larger recovery transfer paths, which take longer.
Keep history retention long enough to cover expected outage windows.
Practical recommendation:
- set retention so normal node restarts can recover from backlog/history
- avoid very short retention values that force frequent full re-sync
- configure retention with project attribute
maxHistoryDuration(seconds)
If you run an initial load on one node while other nodes are active:
- expect synchronization work after startup/rejoin
- monitor all affected projects until they return to
SYNCHRONIZED
Do not treat the cluster as fully available until synchronization completes.
Track at least:
- project cluster state transitions
- synchronization duration and failures
- repeated retries/timeouts
- node availability and readiness endpoints (
/health)
Use alerts for:
- nodes stuck in
SYNCHRONIZING - any
ERRORstate - recurring sync failures after restart events
Check:
- same cluster name on all nodes
- matching discovery method/config
- network/firewall access to cluster ports
- config file is present and valid (YAML/XML syntax)
Check:
- donor node health
- backlog/history retention settings
- storage and network throughput
- very large data changes since node went offline
Check:
- project config consistency across nodes
- runtime config consistency across nodes
- accidental drift after manual changes
- allow cluster ports only within trusted networks
- do not expose internal cluster sync ports to the public internet
- protect credentials/secrets via your platform secret manager
- use least-privilege firewall and security group rules
For safe production changes:
- apply config changes in staging first
- validate node join and synchronization
- roll out progressively in production
- verify all projects return to
SYNCHRONIZED
- Always verify functionality on testing environment before production
- Check whether a full initial load is required after a product update
- Verify that all nodes use the same Tolerant Match major and minor version
- Follow steps of playbook to ensure full functionality of services
- Perform UAT or Testcases to ensure response integrity
- After successful setup on testing environment move to production
-
Prepare infrastructure
- use at least two nodes (three recommended)
- ensure node-to-node connectivity for
clusterPortandclusterSyncServicePort - ensure load balancer can reach all nodes
- ensure time sync (NTP) on all nodes
-
Configure cluster identity and ports
- choose one cluster name for all nodes
- configure either:
- Hazelcast config file (
hazelcast.yaml/hazelcast.xml/clusterconfig.*), or matchRuntimefallback attributes (clusterName,clusterPort,clusterSyncServicePort)
- Hazelcast config file (
- ensure each node has unique service identity/node identity
-
Start services
- start node A
- start node B (and remaining nodes)
- wait until each node is reachable and healthy
-
Configure and enable load balancer routing
- add all nodes as targets
- route only to healthy/ready nodes
-
Verify cluster join
- check that all nodes report the same
clusterName - check expected
clusterSize - check project states move to
SYNCHRONIZED
- check that all nodes report the same
Use this checklist after setup and after operational changes.
Node health checks:
- backend process is running on every node
/healthendpoints respond successfully- required ports are open/listening
Cluster checks:
- all nodes show same cluster name
clusterSizeequals expected node count- project state is
SYNCHRONIZEDon all active nodes - no persistent cluster errors in logs
Load balancer checks:
- requests through LB are distributed to active nodes
- unhealthy nodes are not receiving production traffic
Synchronization checks:
- restart one node and verify it returns to
SYNCHRONIZED - verify other nodes stay available during that restart
- verify data changes are visible across nodes after synchronization
Optional resilience test:
- stop one node intentionally
- verify service remains available via remaining node(s)
- start node again and confirm successful rejoin/synchronization
Acceptance checklist (pass/fail):
- all nodes healthy
- cluster joined with expected size
- all projects synchronized
- LB routing healthy
- restart/rejoin synchronization successful
Use this recovery flow when one restarted node does not return cleanly:
- Remove the affected node from load balancer routing.
- Stop the affected Match service instance.
- Verify the old process is fully terminated and releases file handles.
- Start the node again and wait for health/readiness.
- Verify cluster member count and project state reaches
SYNCHRONIZED. - Re-enable load balancer routing for that node.
Use this flow when new project data must be loaded into an existing cluster.
-
Select one source node for the load.
- Prefer a healthy node that is currently
SYNCHRONIZED. - Remove the source node from load balancer routing if the load may affect response latency or project availability.
- Prefer a healthy node that is currently
-
Run the initial load on the selected source node.
- Stop the affected project on that node before loading.
- Set
--data-epochif the data is to be rolled out to all other cluster members. - Keep the other nodes running so the service remains available through synchronized nodes.
Linux shell example:
service.sh backend --endpoint operations --function stop.project --parameter "projectId=matchProject-1" matchInitialLoad.sh -delete-backlog config/matchserviceconfig.xml matchProject-1 --data-epoch 1713700000000 service.sh backend --endpoint operations --function start.project --parameter "projectId=matchProject-1"
Windows cmd example:
service.exe backend --endpoint operations --function stop.project --parameter "projectId=matchProject-1" matchInitialLoad.bat -delete-backlog config\matchserviceconfig.xml matchProject-1 --data-epoch 1713700000000 service.exe backend --endpoint operations --function start.project --parameter "projectId=matchProject-1"
-
Verify that the affected project is running again on the source node.
-
Monitor cluster synchronization.
- Expect other nodes to enter synchronization while they catch up.
- Do not route full production traffic to nodes whose affected project is
SYNCHRONIZINGorERROR.
-
Verify completion.
- all nodes show the expected
clusterSize - all affected projects return to
SYNCHRONIZED - smoke-test representative requests and response integrity
- all nodes show the expected
-
Re-enable normal load balancer routing after validation.
- This file is the customer-facing cluster manual.
- Helm deployment details remain in
helm/README.md.
These scenarios summarize expected cluster failure recovery behavior:
- One node fails, savepoints available.
- One node fails, no savepoint for restarting node.
- One node fails, no usable history.
- One node fails, no usable history and limited savepoint coverage.
- One node fails, no savepoints available.
- One node fails, no savepoint on preferred master node.
- One node fails, no savepoints on either node.
- Two nodes fail at different times, savepoints available.
- Two nodes fail at different times, no usable history.
- Cluster communication outage, savepoint available.
- Cluster communication outage, no usable history.
Operational interpretation:
- If backlog/history is sufficient, recovery is usually faster.
- If backlog/history is insufficient, full savepoint/database/index transfer paths may be required.
- A higher
maxHistoryDurationtypically reduces expensive full recovery operations.
The descriptions below use a consistent scenario format and should be read together with each scenario diagram.
Common context for all scenario images:
- two nodes (
A,B) F: represents an errorS: startSP: full savepointH: history start timeT(x): timestamp of eventx
Timing expressions such as T(S) > T(F) mean event S happened after event F.
Important note:
- importing an incremental savepoint alone from another cluster node is not possible
- incremental savepoints always depend on the latest full savepoint
- therefore, only full savepoints are exchanged directly
- for database-based synchronization, Match transfers the full savepoint and all incremental savepoints written after that full savepoint
Scenario 1: One node fails, savepoints available
Context: Node B fails and is restarted.
Timing relation: T(S) > T(SP A2) > T(F) > T(H A)
Recovery: use data from SP B1, Backlog B, and History A > T(F) / Backlog A.

Scenario 2: One node fails, no savepoint available for restarting node
Context: Node B fails and is restarted; no savepoint exists for B.
Timing relation: T(S) > T(SP A2) > T(F) > T(H A), no SP for B
Recovery: use data from Backlog B, History A > T(F), and Backlog A.

Scenario 3: One node fails, no history available
Context: Node B fails and is restarted; required history is not available.
Timing relation: T(S) > T(SP A2) > T(H A) > T(F)
Recovery: use data from SP A2 (full savepoint + increments); Backlog B is discarded.

Scenario 4: One node fails, no history and no local savepoint available
Context: Node B fails and is restarted; history is missing and local savepoint is missing.
Timing relation: T(S) > T(SP A2) > T(H A) > T(F), no SP for B
Recovery: use data from SP A2 (full savepoint + increments); Backlog B is discarded.

Scenario 5: One node fails, no usable savepoints in current path
Context: Node B fails and is restarted.
Timing relation: T(S) > T(SP B1) > T(SP A2)
Recovery: use data from SP B1, plus Backlog B and Backlog A > T(F).

Scenario 6: One node fails, no savepoint for master
Context: Node B fails and is restarted; no savepoint exists for A, savepoint exists for B.
Timing relation: no SP for A, SP for B exists
Recovery: use data from SP B1, Backlog B, and Backlog A > T(F).

Scenario 7: One node fails, no savepoints available
Context: Node B fails and is restarted; no savepoints exist for A or B.
Timing relation: no SP for A and B
Recovery: use data from Backlog B and Backlog A > T(F).

Scenario 8: Two nodes fail at different times, savepoints available
Context: Nodes A and B fail in sequence.
Timing relation: T(S A) > T(S B) > T(F A) > T(F B)
Recovery:
A uses Backlog A with Backlog B > T(F A).
B at T(S B): no action.
B at T(S A): use Backlog A > T(F B) for keys without newer updates.

Scenario 9: Two nodes fail at different times, no history available
Context: Nodes A and B fail in sequence; history coverage is insufficient.
Timing relation: T(S A) > T(S B) > T(F A) > T(H A) > T(F B)
Recovery:
A uses Backlog A with Backlog B > T(F A).
B at T(S B): no action.
B at T(S A): use SP A1 (full savepoint + increments), after A synchronization.

Scenario 10: Cluster communication fails, savepoint available
Context: Cluster communication is interrupted temporarily.
Timing relation: T(CS) > max(T(SP A1), T(SP B1))
Recovery: A uses Backlog B > T(CS) and B uses Backlog A > T(CS).

Scenario 11: Cluster communication fails, no history available
Context: Cluster communication is interrupted and history coverage is insufficient.
Timing relation: T(CR) > T(SP B1) > T(H B) > T(CS) or T(H A) > T(CS)
Recovery: broader export/import reconciliation between A and B; in key conflicts, master-node entries win.
