Oracle RAC Private Interconnect: Architecture, Performance, Monitoring and Troubleshooting
In the previous part, we looked at GCS and GES and saw how heavily RAC depends on communication between instances.
That leads us to one of the most important infrastructure components in an Oracle RAC environment:
The private interconnect.
If the interconnect is slow, unstable, congested, misconfigured, or incorrectly designed, the impact can be much larger than a simple network problem.
It can affect:
- Cache Fusion
- GCS/GES
- Cluster heartbeats
- Node membership
- Database performance
- Node eviction
- RAC stability
This is why an Oracle RAC expert should be able to troubleshoot the interconnect from both the Oracle and Linux layers.
1. What Is the RAC Private Interconnect?
The private interconnect is the network used by RAC nodes for internal cluster communication.
It carries traffic such as:
- Cache Fusion
- GCS communication
- GES communication
- Cluster heartbeats
- Global resource coordination
A typical RAC architecture looks like:
Clients
|
Public Network
|
+------+------+
| |
Node 1 Node 2
| |
eth0 eth0
| |
+-------------+
Public LAN
eth1 eth1
| |
+-------------+
Private
Interconnect
The public network handles client-facing traffic.
The private network handles internal RAC traffic.
2. Why Separate the Networks?
Imagine the following design:
Application Traffic
+
RAC Interconnect
|
NIC
During a large application workload, the network interface could become saturated.
Now RAC’s internal traffic competes with application traffic.
That is a bad design.
A better architecture is:
Public Network
|
eth0
|
Node
eth1
|
Private Network
|
RAC
This gives RAC its own communication path.
3. What Travels Over the Interconnect?
The interconnect carries significantly more than database blocks.
Examples include:
Cache Fusion
Instance 1
|
| Database Block
v
Interconnect
|
v
Instance 2
GCS/GES
Global cache and lock coordination.
Cluster Heartbeats
Used by Clusterware to determine whether nodes are alive.
Therefore:
A private interconnect failure can become a cluster availability problem, not merely a performance problem.
4. Oracle’s View of the Interconnect
Use:
SELECT inst_id,
name,
ip_address,
is_public
FROM gv$cluster_interconnects;
Example:
INST_ID NAME IP_ADDRESS IS_PUBLIC ------- ---- --------------- --------- 1 eth1 192.168.20.11 NO 2 eth1 192.168.20.12 NO
This tells you which interface Oracle is using for the RAC interconnect.
5. Check Oracle’s Network Configuration
Use:
oifcfg getif
Example:
eth0 10.10.10.0 global public
eth1 192.168.20.0 global cluster_interconnect
This is one of the first commands I run when troubleshooting RAC networking.
6. Adding an Interconnect
The oifcfg utility is used to configure Oracle network interfaces.
Example:
oifcfg iflist
This displays interfaces that Oracle can potentially use.
Then:
oifcfg getif
shows the interfaces currently configured.
If configuration is required, use Oracle-supported oifcfg commands appropriate for your environment.
Do not manually modify internal Clusterware configuration files.
7. Public vs Private Interface
Example:
eth0 10.10.10.0/24 PUBLIC eth1 192.168.20.0/24 PRIVATE
The roles are different.
Public
Used for:
- Client connections
- SCAN
- VIP
- Listener traffic
Private
Used for:
- Cache Fusion
- GCS
- GES
- Cluster communication
8. How Fast Should the Interconnect Be?
There is no single bandwidth value that is correct for every RAC environment.
The correct choice depends on:
- Number of RAC nodes
- Database workload
- Cache Fusion traffic
- Transaction volume
- Database block size
- Application design
- Consolidation
- Network architecture
For modern production RAC systems, 10 GbE should generally be considered a baseline rather than an ambitious target, while 25/40/100 GbE may be appropriate for high-throughput environments.
The important point is not simply bandwidth.
You need:
Low latency + low packet loss + sufficient bandwidth + redundancy.
9. Bandwidth Is Not Everything
Consider two networks:
Network A 100 Gb/s 2 ms latency Network B 25 Gb/s 20 μs latency
For some RAC workloads, Network B may provide better behavior despite having lower bandwidth.
RAC is highly sensitive to latency because Cache Fusion operations can occur during SQL execution.
10. Check NIC Speed
On Linux:
ethtool eth1
Example:
Speed: 25Gb/s Duplex: Full Link detected: yes
Check all interfaces:
ip link
11. Check NIC Errors
Run:
ip -s link show eth1
Look for:
RX errors TX errors RX dropped TX dropped
A healthy RAC interconnect should not have unexplained increasing error counters.
12. Check Network Statistics
Use:
sar -n DEV 1 10
This allows you to observe traffic over time.
Look for:
- High utilization
- Drops
- Errors
- Unexpected traffic patterns
13. Check Packet Loss
Basic test:
ping 192.168.20.12
Example:
64 bytes from 192.168.20.12: time=0.12 ms
You should test from every RAC node to every other RAC node.
For a three-node RAC:
Node1 → Node2 Node1 → Node3 Node2 → Node1 Node2 → Node3 Node3 → Node1 Node3 → Node2
14. Don’t Rely Only on Ping
This is important.
A successful ping does not prove that the RAC interconnect is healthy.
You can have:
Ping: OK
while experiencing:
Packet drops NIC errors Switch congestion Microbursts TCP retransmissions High application latency
Therefore, RAC network troubleshooting must combine:
- Oracle statistics
- Linux statistics
- Switch statistics
- Application performance
15. MTU and Jumbo Frames
Jumbo frames are often discussed in RAC deployments.
Traditional Ethernet:
MTU = 1500
Jumbo frames:
MTU ≈ 9000
Potential benefits include:
- Lower packet-processing overhead
- More efficient high-throughput transfers
But there is an important rule:
Do not enable jumbo frames on only part of the communication path.
All relevant devices must support the configured MTU consistently.
16. Testing MTU
If your environment uses MTU 9000, test it end-to-end.
For example:
ping -M do -s 8972 192.168.20.12
The exact payload depends on IP/ICMP headers.
If fragmentation occurs or packets fail, investigate the entire path:
Node ↓ NIC ↓ Switch ↓ Switch ↓ NIC ↓ Node
17. Jumbo Frames Are Not Automatically Better
This is an important expert-level point.
Don’t configure MTU 9000 simply because someone says:
“RAC performs better with jumbo frames.”
The correct question is:
Does the complete network path support it correctly and does testing demonstrate a benefit?
If the configuration is inconsistent, you can create difficult-to-diagnose connectivity problems.
18. Redundancy
A production RAC environment should avoid a single point of failure in the interconnect.
A typical design may use:
Node 1 ├── NIC1 ────── Switch A └── NIC2 ────── Switch B Node 2 ├── NIC1 ────── Switch A └── NIC2 ────── Switch B
The exact implementation depends on your Oracle and network architecture.
19. Multiple Private Networks
Modern Oracle RAC supports multiple private networks and can use them for redundancy and load distribution.
The important architectural goal is:
Failure of one path
↓
RAC communication continues
Rather than:
Private NIC failure
↓
Cluster instability
20. Network Bonding
Linux bonding or vendor-specific network teaming may be used depending on the platform and architecture.
But don’t assume that every bonding mode is appropriate for RAC.
Always validate:
- Oracle support
- Operating system support
- Switch configuration
- Driver compatibility
- Failover behavior
A technically functional network configuration isn’t necessarily a supported Oracle architecture.
21. Interconnect Failure Scenario
Consider a two-node RAC:
Node1 ======== Node2
Private
Network
The private switch suddenly fails.
The nodes can no longer exchange heartbeats.
Clusterware sees:
Node1 ←→ Node2 Communication Lost
CSSD must determine which node should remain in the cluster.
If quorum and cluster membership cannot safely be maintained, Oracle may evict a node.
22. Why Node Eviction Happens
A node eviction can look frightening:
Node2 ↓ Reboot
But Oracle is protecting the database.
The alternative could be:
Node1 → Database Node2 → Database Both believe they are active
This creates the possibility of split-brain.
Therefore:
Node eviction is often the protection mechanism, not the root cause.
23. Diagnosing a Node Eviction
When a node unexpectedly reboots, don’t start with the database alert log.
Start with Clusterware.
Check:
crsctl check cluster
Then:
crsctl query css votedisk
Then inspect CSS logs:
$GRID_HOME/log/<hostname>/cssd/
Look for:
heartbeat misscount voting disk eviction network
24. Check CSSD Logs
For example:
cd $GRID_HOME/log/$(hostname)/cssd
Then:
grep -iE "evict|heartbeat|voting|network|misscount" ocssd.log
This can quickly reveal whether the event was related to:
- Interconnect
- Voting Disk
- Storage
- Node communication
25. RAC Interconnect and Cache Fusion
Now connect the concepts from Part 12 and Part 13.
Application
|
v
Instance 1
|
| Cache Fusion
v
Private Interconnect
|
v
Instance 2
If the interconnect becomes slow:
SQL ↓ GC request ↓ Interconnect latency ↓ Longer response time
The database may still be healthy.
The infrastructure beneath it may be the bottleneck.
26. Identifying GC Performance Problems
Start with:
SELECT inst_id,
event,
total_waits,
time_waited_micro
FROM gv$system_event
WHERE event LIKE 'gc%'
ORDER BY time_waited_micro DESC;
Then examine:
SELECT inst_id,
name,
value
FROM gv$sysstat
WHERE name LIKE 'gc%'
ORDER BY inst_id, name;
If both GC activity and network latency increase at the same time, the interconnect becomes a strong candidate.
27. A Practical Correlation Model
A good RAC investigation correlates:
SQL latency
↓
RAC wait events
↓
GC statistics
↓
Interconnect traffic
↓
NIC statistics
↓
Switch statistics
For example:
Application response time ↑
↓
gc current request ↑
↓
Interconnect traffic ↑
↓
NIC utilization 95%
↓
Switch interface congestion
Now you have an evidence-based diagnosis.
28. Useful SQL Health Check
You can create a basic RAC interconnect report:
SELECT inst_id,
name,
ip_address,
is_public
FROM gv$cluster_interconnects
ORDER BY inst_id;
Then:
SELECT inst_id,
event,
total_waits,
ROUND(time_waited_micro / 1000000, 2) AS seconds_waited
FROM gv$system_event
WHERE event LIKE 'gc%'
ORDER BY seconds_waited DESC;
And:
SELECT inst_id,
name,
value
FROM gv$sysstat
WHERE name IN (
'gc cr blocks received',
'gc cr blocks served',
'gc current blocks received',
'gc current blocks served'
)
ORDER BY inst_id, name;
29. Linux Health Check
On each node:
echo "===== INTERFACES =====" ip -br link echo "===== IP CONFIGURATION =====" ip -br addr echo "===== ROUTES =====" ip route echo "===== NIC STATUS =====" ethtool eth1 echo "===== NIC STATISTICS =====" ip -s link show eth1
Then:
sar -n DEV 1 5
This gives you a good first-level infrastructure snapshot.
30. What About TCP?
Oracle RAC interconnect communication has historically used protocols such as UDP, with Oracle supporting different communication mechanisms depending on platform and release.
The important DBA lesson is:
Don’t assume RAC interconnect performance is equivalent to normal application TCP traffic.
RAC’s internal communication is controlled by Oracle Clusterware and RAC internals.
Therefore, don’t troubleshoot it solely as an ordinary application network.
31. A Production Troubleshooting Example
Imagine a four-node RAC:
ORCL1 ORCL2 ORCL3 ORCL4
Users report:
“The database is randomly slow.”
AWR shows:
gc current request gc cr request
increasing.
CPU:
40%
Storage latency:
Normal
ASM:
Healthy
Interconnect:
25 GbE
But Linux shows increasing:
RX dropped TX dropped
Switch monitoring reveals:
Private VLAN interface 95% utilization
The database isn’t CPU-bound.
The database isn’t storage-bound.
The bottleneck is the RAC interconnect.
After resolving the network congestion, GC wait times return to normal.
32. What an Expert Should Never Do
Don’t immediately change:
RAC timeout parameters CSS parameters GC parameters _hidden parameters
just because a node was evicted or GC waits increased.
First determine:
What failed? Why did it fail? Can the problem be reproduced? Which layer owns the failure?
Only then consider configuration changes.
33. RAC Interconnect Checklist
For a production RAC health check:
Oracle Layer
oifcfg getif
SELECT * FROM gv$cluster_interconnects;
Linux Layer
ip -s link
ethtool <interface>
sar -n DEV 1 10
Network Layer
Check:
- Switch errors
- Interface utilization
- Packet drops
- MTU
- Duplex
- LACP/bonding
- VLAN configuration
- Redundancy
Database Layer
Check:
gcwaits- Cache Fusion statistics
- AWR
- ASH
- SQL workload
34. RAC Interconnect Design — Expert View
A strong production architecture should aim for:
+-------------+
| Application |
+------+------+
|
Public Network
|
+----------------+----------------+
| |
RAC Node 1 RAC Node 2
| |
Public NIC Public NIC
| |
Private NIC Private NIC
| |
+---------- Private Network -----+
For larger environments:
Redundant Private Fabric
/ \
Switch A Switch B
/ \ / \
Node1 Node2 Node1 Node2
The exact topology should be validated against Oracle’s supported architecture and the organization’s network standards.
35. DBA Expert Tip
When troubleshooting RAC performance, don’t ask only:
“Is the interconnect up?”
Ask:
“Is the interconnect fast, stable, uncongested, redundant, and correctly configured?”
There is a huge difference between:
Interface = UP
and:
RAC Interconnect = Healthy
A network can be technically online while still causing significant RAC performance problems.
Final Takeaway
The RAC private interconnect is one of the most important components of the entire Oracle RAC architecture.
It connects:
GCS GES CSSD Cache Fusion RAC Instances
A problem at this layer can manifest as:
- Slow SQL
- High
gcwaits - Excessive Cache Fusion
- Node eviction
- Cluster instability
- Application disconnects
The best RAC administrators therefore treat the interconnect as a first-class database performance component, not simply a networking detail.
In RAC, the network between the instances is part of the database engine.
Practical Exercise
Before moving to the next part, run these commands on each RAC node:
oifcfg getif
crsctl query css votedisk
ip -s link
ethtool <private_interface>
Then run:
SELECT inst_id,
name,
ip_address,
is_public
FROM gv$cluster_interconnects
ORDER BY inst_id;
and:
SELECT inst_id,
event,
total_waits,
ROUND(time_waited_micro / 1000000,2) seconds_waited
FROM gv$system_event
WHERE event LIKE 'gc%'
ORDER BY seconds_waited DESC;
You now have the foundation for building a real RAC interconnect health-check report rather than simply checking whether the interface is online.


