Oracle Voting Disks: Cluster Membership, Node Evictions, and Split-Brain Protection
If the Oracle Cluster Registry (OCR) stores the configuration of the cluster, the Voting Disks protect the cluster itself.
Every Oracle RAC administrator has heard messages such as:
- Node eviction
- CSS communication failure
- Split-brain detected
- Lost voting disk
- Node rebooted by Clusterware
In almost every one of these situations, Voting Disks play an important role.
Understanding how they work is essential for troubleshooting RAC environments because many cluster failures begin long before the database instance goes offline.
In this article, we’ll explore the purpose of Voting Disks, how Oracle uses them to maintain cluster integrity, and what happens when nodes lose communication.
What are Voting Disks?
Voting Disks are shared storage files used by Cluster Synchronization Services (CSSD) to determine which nodes belong to the cluster.
Their primary responsibilities are:
- Maintain cluster membership
- Detect node failures
- Prevent split-brain
- Participate in quorum decisions
- Coordinate node eviction
Unlike the OCR, Voting Disks do not store cluster configuration.
Why are Voting Disks Needed?
Imagine a two-node cluster.
+-----------+
| Node 1 |
+-----------+
|
|
Shared Storage
|
+-----------+
| Node 2 |
+-----------+
Everything works normally while both nodes can communicate.
But what happens if communication suddenly stops?
Node1 X----------------X Node2
Each node believes it is still alive.
If both continue writing to the database independently, the database could become corrupted.
This situation is called split-brain.
Voting Disks help Oracle decide which node is allowed to continue running.
Understanding Split-Brain
A split-brain condition occurs when:
- Nodes lose communication.
- Both believe the other has failed.
- Both attempt to access the same database.
Example:
Network Failure Node1 --------------------X Node2 --------------------X
Without protection:
Node1 continues writing.
Node2 continues writing.
Shared storage becomes inconsistent.
Oracle avoids this scenario by evicting one side of the cluster.
What is Quorum?
Oracle RAC uses the concept of quorum.
A node must maintain communication with the majority of the cluster.
Example:
Three-node cluster
Node1 Node2 Node3
If Node3 becomes isolated while Nodes1 and 2 still communicate:
- Nodes1 and 2 form the majority.
- Node3 loses quorum.
- CSSD evicts Node3.
This prevents multiple independent clusters from accessing the same storage.
Voting Disk Architecture
Modern Oracle RAC stores Voting Disks inside ASM.
Example:
+OCR Disk Group
+----------------------+
| OCR |
| Voting Files |
+----------------------+
Every node accesses the same Voting Disks.
CSSD and Voting Disks
The Cluster Synchronization Services Daemon (CSSD) continuously exchanges heartbeat information.
Heartbeats are exchanged through:
- Private interconnect
- Voting Disks
Example:
Node1
| Heartbeat
Private Network
| Heartbeat
Node2
|
Voting Disk
As long as both heartbeat mechanisms remain healthy, the cluster is stable.
Network Heartbeat vs Disk Heartbeat
Oracle uses two independent heartbeat mechanisms.
Network Heartbeat
Uses the private interconnect.
Checks whether nodes can communicate.
Disk Heartbeat
Uses Voting Disks.
Checks whether nodes still have access to shared storage.
Oracle evaluates both before making an eviction decision.
Node Eviction
Node eviction is one of the most misunderstood RAC events.
When Oracle evicts a node, it is protecting the cluster—not causing a failure.
Example:
Private Network Failure
Node1
X
Node2
CSSD determines that Node2 no longer has quorum.
Result:
Node2 Rebooted Cluster Protected
The surviving node continues operating safely.
Viewing Voting Disk Configuration
Display Voting Disk information:
Example:
crsctl query css votedisk
## STATE File Universal Id File Name
ONLINE 8d7c3b4a... +OCR
Located 1 voting disk(s).
This is one of the most commonly used commands during RAC troubleshooting.
Check Cluster Health
Verify CSS:
crsctl check css
Example:
CRS-4529: Cluster Synchronization Services is online
Verify Cluster:
crsctl check cluster
Check ASM Disk Groups
Because Voting Disks usually reside in ASM:
asmcmd lsdg
Verify that the +OCR disk group is mounted.
Viewing Cluster Resources
crsctl stat res -t
Confirm that:
- CSSD
- ASM
- Database
- Listeners
remain ONLINE.
Common Voting Disk Problems
Voting Disk Unavailable
Symptoms:
- CSS startup failure.
- CRS startup failure.
- Node reboot.
Verify:
crsctl query css votedisk
Check ASM availability.
Private Interconnect Failure
Symptoms:
- Node eviction.
- CSS errors.
- Cluster instability.
Verify:
oifcfg getif
Inspect network latency and packet loss.
ASM Disk Group Offline
Symptoms:
- Voting Disk inaccessible.
Verify:
asmcmd lsdg
Storage Failure
Possible causes:
- SAN outage.
- Multipath issues.
- Fibre Channel failure.
- Storage controller failure.
Always investigate the storage layer before restarting Clusterware.
Reading CSS Logs
Most Voting Disk problems appear first in the CSS logs.
Typical locations:
Grid_Home/log/<node>/cssd/
Example:
cd $GRID_HOME/log/racnode1/cssd
Useful files include:
ocssd.log
Search for keywords such as:
- voting disk
- eviction
- heartbeat
- misscount
- reboot
- fatal
These messages often provide the first indication of the root cause.
Real Production Scenario
Consider a two-node RAC cluster.
Node1 and Node2 are operating normally.
Suddenly, a network switch fails.
Node1 X----------------X Node2
Node1 still has access to:
- Shared storage
- Voting Disk
Node2 loses access to the private interconnect.
CSSD determines that Node2 cannot safely remain in the cluster.
Oracle immediately reboots Node2.
At first point, it appears that Oracle caused an outage.
In reality, Oracle prevented a much more serious issue: simultaneous writes from two independent nodes.
Useful Voting Disk Commands
Display Voting Disks:
crsctl query css votedisk
Check CSS:
crsctl check css
Check Cluster:
crsctl check cluster
Display resources:
crsctl stat res -t
Display disk groups:
asmcmd lsdg
These commands should become part of every RAC administrator’s troubleshooting routine.
Best Practices
- Store Voting Disks in a redundant ASM disk group.
- Use reliable shared storage.
- Configure redundant private interconnects.
- Monitor network latency continuously.
- Regularly review CSS logs for warning messages.
- Investigate node evictions rather than simply restarting the affected node.
- Keep Grid Infrastructure patched with current Release Updates.
DBA Tip
When a node is unexpectedly rebooted in a RAC environment, many administrators immediately focus on the alert log of the database instance.
In most cases, that’s not where the story begins.
Start with the Clusterware logs:
- Check the CSS logs.
- Verify Voting Disk availability.
- Confirm the health of the private interconnect.
- Review ASM disk group status.
- Then investigate the database.
I’ve seen many incidents where the database was functioning perfectly, but a brief network interruption or storage delay caused CSSD to evict a node to protect the cluster.
A node eviction is usually a symptom of an underlying infrastructure problem—not the problem itself. Always investigate why the eviction occurred before bringing the node back into service.
Conclusion
Voting Disks are a critical part of Oracle RAC. They help maintain cluster membership, enforce quorum, and prevent split-brain scenarios that could otherwise lead to database corruption.
Although they operate behind the scenes, understanding how CSSD uses Voting Disks, how heartbeats are exchanged, and why Oracle performs node evictions will greatly improve your ability to troubleshoot RAC clusters with confidence.


