Oracle RAC 19c Masterclass – Part 14

Oracle RAC Private Interconnect: Architecture, Performance, Monitoring and Troubleshooting

In the previous part, we looked at GCS and GES and saw how heavily RAC depends on communication between instances.

That leads us to one of the most important infrastructure components in an Oracle RAC environment:

The private interconnect.

If the interconnect is slow, unstable, congested, misconfigured, or incorrectly designed, the impact can be much larger than a simple network problem.

It can affect:

  • Cache Fusion
  • GCS/GES
  • Cluster heartbeats
  • Node membership
  • Database performance
  • Node eviction
  • RAC stability

This is why an Oracle RAC expert should be able to troubleshoot the interconnect from both the Oracle and Linux layers.

1. What Is the RAC Private Interconnect?

The private interconnect is the network used by RAC nodes for internal cluster communication.

It carries traffic such as:

  • Cache Fusion
  • GCS communication
  • GES communication
  • Cluster heartbeats
  • Global resource coordination

A typical RAC architecture looks like:

                    Clients
                       |
                 Public Network
                       |
                +------+------+
                |             |
             Node 1         Node 2
                |             |
              eth0           eth0
                |             |
                +-------------+
                  Public LAN

                eth1           eth1
                  |             |
                  +-------------+
                  Private
                Interconnect

The public network handles client-facing traffic.

The private network handles internal RAC traffic.

2. Why Separate the Networks?

Imagine the following design:

Application Traffic
        +
RAC Interconnect
        |
      NIC

During a large application workload, the network interface could become saturated.

Now RAC’s internal traffic competes with application traffic.

That is a bad design.

A better architecture is:

              Public Network
                    |
                  eth0
                    |
                  Node

                  eth1
                    |
            Private Network
                    |
                 RAC

This gives RAC its own communication path.

3. What Travels Over the Interconnect?

The interconnect carries significantly more than database blocks.

Examples include:

Cache Fusion
Instance 1
    |
    | Database Block
    v
Interconnect
    |
    v
Instance 2
GCS/GES

Global cache and lock coordination.

Cluster Heartbeats

Used by Clusterware to determine whether nodes are alive.

Therefore:

A private interconnect failure can become a cluster availability problem, not merely a performance problem.

4. Oracle’s View of the Interconnect

Use:

SELECT inst_id,
       name,
       ip_address,
       is_public
FROM gv$cluster_interconnects;

Example:

INST_ID  NAME  IP_ADDRESS      IS_PUBLIC
-------  ----  --------------- ---------
1        eth1  192.168.20.11   NO
2        eth1  192.168.20.12   NO

This tells you which interface Oracle is using for the RAC interconnect.

5. Check Oracle’s Network Configuration

Use:

oifcfg getif

Example:

eth0  10.10.10.0      global  public
eth1  192.168.20.0    global  cluster_interconnect

This is one of the first commands I run when troubleshooting RAC networking.

6. Adding an Interconnect

The oifcfg utility is used to configure Oracle network interfaces.

Example:

oifcfg iflist

This displays interfaces that Oracle can potentially use.

Then:

oifcfg getif

shows the interfaces currently configured.

If configuration is required, use Oracle-supported oifcfg commands appropriate for your environment.

Do not manually modify internal Clusterware configuration files.

7. Public vs Private Interface

Example:

eth0
10.10.10.0/24
PUBLIC

eth1
192.168.20.0/24
PRIVATE

The roles are different.

Public

Used for:

  • Client connections
  • SCAN
  • VIP
  • Listener traffic
Private

Used for:

  • Cache Fusion
  • GCS
  • GES
  • Cluster communication

8. How Fast Should the Interconnect Be?

There is no single bandwidth value that is correct for every RAC environment.

The correct choice depends on:

  • Number of RAC nodes
  • Database workload
  • Cache Fusion traffic
  • Transaction volume
  • Database block size
  • Application design
  • Consolidation
  • Network architecture

For modern production RAC systems, 10 GbE should generally be considered a baseline rather than an ambitious target, while 25/40/100 GbE may be appropriate for high-throughput environments.

The important point is not simply bandwidth.

You need:

Low latency + low packet loss + sufficient bandwidth + redundancy.

9. Bandwidth Is Not Everything

Consider two networks:

Network A
100 Gb/s
2 ms latency

Network B
25 Gb/s
20 μs latency

For some RAC workloads, Network B may provide better behavior despite having lower bandwidth.

RAC is highly sensitive to latency because Cache Fusion operations can occur during SQL execution.

10. Check NIC Speed

On Linux:

ethtool eth1

Example:

Speed: 25Gb/s
Duplex: Full
Link detected: yes

Check all interfaces:

ip link

11. Check NIC Errors

Run:

ip -s link show eth1

Look for:

RX errors
TX errors
RX dropped
TX dropped

A healthy RAC interconnect should not have unexplained increasing error counters.

12. Check Network Statistics

Use:

sar -n DEV 1 10

This allows you to observe traffic over time.

Look for:

  • High utilization
  • Drops
  • Errors
  • Unexpected traffic patterns

13. Check Packet Loss

Basic test:

ping 192.168.20.12

Example:

64 bytes from 192.168.20.12: time=0.12 ms

You should test from every RAC node to every other RAC node.

For a three-node RAC:

Node1 → Node2
Node1 → Node3
Node2 → Node1
Node2 → Node3
Node3 → Node1
Node3 → Node2

14. Don’t Rely Only on Ping

This is important.

A successful ping does not prove that the RAC interconnect is healthy.

You can have:

Ping:
OK

while experiencing:

Packet drops
NIC errors
Switch congestion
Microbursts
TCP retransmissions
High application latency

Therefore, RAC network troubleshooting must combine:

  • Oracle statistics
  • Linux statistics
  • Switch statistics
  • Application performance

15. MTU and Jumbo Frames

Jumbo frames are often discussed in RAC deployments.

Traditional Ethernet:

MTU = 1500

Jumbo frames:

MTU ≈ 9000

Potential benefits include:

  • Lower packet-processing overhead
  • More efficient high-throughput transfers

But there is an important rule:

Do not enable jumbo frames on only part of the communication path.

All relevant devices must support the configured MTU consistently.

16. Testing MTU

If your environment uses MTU 9000, test it end-to-end.

For example:

ping -M do -s 8972 192.168.20.12

The exact payload depends on IP/ICMP headers.

If fragmentation occurs or packets fail, investigate the entire path:

Node
 ↓
NIC
 ↓
Switch
 ↓
Switch
 ↓
NIC
 ↓
Node

17. Jumbo Frames Are Not Automatically Better

This is an important expert-level point.

Don’t configure MTU 9000 simply because someone says:

“RAC performs better with jumbo frames.”

The correct question is:

Does the complete network path support it correctly and does testing demonstrate a benefit?

If the configuration is inconsistent, you can create difficult-to-diagnose connectivity problems.

18. Redundancy

A production RAC environment should avoid a single point of failure in the interconnect.

A typical design may use:

Node 1
 ├── NIC1 ────── Switch A
 └── NIC2 ────── Switch B

Node 2
 ├── NIC1 ────── Switch A
 └── NIC2 ────── Switch B

The exact implementation depends on your Oracle and network architecture.

19. Multiple Private Networks

Modern Oracle RAC supports multiple private networks and can use them for redundancy and load distribution.

The important architectural goal is:

Failure of one path
        ↓
RAC communication continues

Rather than:

Private NIC failure
        ↓
Cluster instability

20. Network Bonding

Linux bonding or vendor-specific network teaming may be used depending on the platform and architecture.

But don’t assume that every bonding mode is appropriate for RAC.

Always validate:

  • Oracle support
  • Operating system support
  • Switch configuration
  • Driver compatibility
  • Failover behavior

A technically functional network configuration isn’t necessarily a supported Oracle architecture.

21. Interconnect Failure Scenario

Consider a two-node RAC:

Node1 ======== Node2
       Private
       Network

The private switch suddenly fails.

The nodes can no longer exchange heartbeats.

Clusterware sees:

Node1 ←→ Node2

Communication Lost

CSSD must determine which node should remain in the cluster.

If quorum and cluster membership cannot safely be maintained, Oracle may evict a node.

22. Why Node Eviction Happens

A node eviction can look frightening:

Node2
   ↓
Reboot

But Oracle is protecting the database.

The alternative could be:

Node1 → Database
Node2 → Database

Both believe they are active

This creates the possibility of split-brain.

Therefore:

Node eviction is often the protection mechanism, not the root cause.

23. Diagnosing a Node Eviction

When a node unexpectedly reboots, don’t start with the database alert log.

Start with Clusterware.

Check:

crsctl check cluster

Then:

crsctl query css votedisk

Then inspect CSS logs:

$GRID_HOME/log/<hostname>/cssd/

Look for:

heartbeat
misscount
voting disk
eviction
network

24. Check CSSD Logs

For example:

cd $GRID_HOME/log/$(hostname)/cssd

Then:

grep -iE "evict|heartbeat|voting|network|misscount" ocssd.log

This can quickly reveal whether the event was related to:

  • Interconnect
  • Voting Disk
  • Storage
  • Node communication

25. RAC Interconnect and Cache Fusion

Now connect the concepts from Part 12 and Part 13.

Application
     |
     v
Instance 1
     |
     | Cache Fusion
     v
Private Interconnect
     |
     v
Instance 2

If the interconnect becomes slow:

SQL
 ↓
GC request
 ↓
Interconnect latency
 ↓
Longer response time

The database may still be healthy.

The infrastructure beneath it may be the bottleneck.

26. Identifying GC Performance Problems

Start with:

SELECT inst_id,
       event,
       total_waits,
       time_waited_micro
FROM gv$system_event
WHERE event LIKE 'gc%'
ORDER BY time_waited_micro DESC;

Then examine:

SELECT inst_id,
       name,
       value
FROM gv$sysstat
WHERE name LIKE 'gc%'
ORDER BY inst_id, name;

If both GC activity and network latency increase at the same time, the interconnect becomes a strong candidate.

27. A Practical Correlation Model

A good RAC investigation correlates:

SQL latency
     ↓
RAC wait events
     ↓
GC statistics
     ↓
Interconnect traffic
     ↓
NIC statistics
     ↓
Switch statistics

For example:

Application response time ↑

        ↓

gc current request ↑

        ↓

Interconnect traffic ↑

        ↓

NIC utilization 95%

        ↓

Switch interface congestion

Now you have an evidence-based diagnosis.

28. Useful SQL Health Check

You can create a basic RAC interconnect report:

SELECT inst_id,
       name,
       ip_address,
       is_public
FROM gv$cluster_interconnects
ORDER BY inst_id;

Then:

SELECT inst_id,
       event,
       total_waits,
       ROUND(time_waited_micro / 1000000, 2) AS seconds_waited
FROM gv$system_event
WHERE event LIKE 'gc%'
ORDER BY seconds_waited DESC;

And:

SELECT inst_id,
       name,
       value
FROM gv$sysstat
WHERE name IN (
    'gc cr blocks received',
    'gc cr blocks served',
    'gc current blocks received',
    'gc current blocks served'
)
ORDER BY inst_id, name;

29. Linux Health Check

On each node:

echo "===== INTERFACES ====="
ip -br link

echo "===== IP CONFIGURATION ====="
ip -br addr

echo "===== ROUTES ====="
ip route

echo "===== NIC STATUS ====="
ethtool eth1

echo "===== NIC STATISTICS ====="
ip -s link show eth1

Then:

sar -n DEV 1 5

This gives you a good first-level infrastructure snapshot.

30. What About TCP?

Oracle RAC interconnect communication has historically used protocols such as UDP, with Oracle supporting different communication mechanisms depending on platform and release.

The important DBA lesson is:

Don’t assume RAC interconnect performance is equivalent to normal application TCP traffic.

RAC’s internal communication is controlled by Oracle Clusterware and RAC internals.

Therefore, don’t troubleshoot it solely as an ordinary application network.

31. A Production Troubleshooting Example

Imagine a four-node RAC:

ORCL1
ORCL2
ORCL3
ORCL4

Users report:

“The database is randomly slow.”

AWR shows:

gc current request
gc cr request

increasing.

CPU:

40%

Storage latency:

Normal

ASM:

Healthy

Interconnect:

25 GbE

But Linux shows increasing:

RX dropped
TX dropped

Switch monitoring reveals:

Private VLAN interface
95% utilization

The database isn’t CPU-bound.

The database isn’t storage-bound.

The bottleneck is the RAC interconnect.

After resolving the network congestion, GC wait times return to normal.

32. What an Expert Should Never Do

Don’t immediately change:

RAC timeout parameters
CSS parameters
GC parameters
_hidden parameters

just because a node was evicted or GC waits increased.

First determine:

What failed?
Why did it fail?
Can the problem be reproduced?
Which layer owns the failure?

Only then consider configuration changes.

33. RAC Interconnect Checklist

For a production RAC health check:

Oracle Layer
oifcfg getif

SELECT * FROM gv$cluster_interconnects;
Linux Layer
ip -s link

ethtool <interface>

sar -n DEV 1 10
Network Layer

Check:

  • Switch errors
  • Interface utilization
  • Packet drops
  • MTU
  • Duplex
  • LACP/bonding
  • VLAN configuration
  • Redundancy

Database Layer

Check:

  • gc waits
  • Cache Fusion statistics
  • AWR
  • ASH
  • SQL workload

34. RAC Interconnect Design — Expert View

A strong production architecture should aim for:

                    +-------------+
                    | Application |
                    +------+------+
                           |
                     Public Network
                           |
          +----------------+----------------+
          |                                 |
       RAC Node 1                         RAC Node 2
          |                                 |
       Public NIC                        Public NIC
          |                                 |
       Private NIC                      Private NIC
          |                                 |
          +---------- Private Network -----+

For larger environments:

             Redundant Private Fabric
                 /             \
             Switch A         Switch B
              /   \             /   \
           Node1 Node2       Node1 Node2

The exact topology should be validated against Oracle’s supported architecture and the organization’s network standards.

35. DBA Expert Tip

When troubleshooting RAC performance, don’t ask only:

“Is the interconnect up?”

Ask:

“Is the interconnect fast, stable, uncongested, redundant, and correctly configured?”

There is a huge difference between:

Interface = UP

and:

RAC Interconnect = Healthy

A network can be technically online while still causing significant RAC performance problems.

Final Takeaway

The RAC private interconnect is one of the most important components of the entire Oracle RAC architecture.

It connects:

GCS
GES
CSSD
Cache Fusion
RAC Instances

A problem at this layer can manifest as:

  • Slow SQL
  • High gc waits
  • Excessive Cache Fusion
  • Node eviction
  • Cluster instability
  • Application disconnects

The best RAC administrators therefore treat the interconnect as a first-class database performance component, not simply a networking detail.

In RAC, the network between the instances is part of the database engine.

Practical Exercise

Before moving to the next part, run these commands on each RAC node:

oifcfg getif

crsctl query css votedisk

ip -s link

ethtool <private_interface>

Then run:

SELECT inst_id,
       name,
       ip_address,
       is_public
FROM gv$cluster_interconnects
ORDER BY inst_id;

and:

SELECT inst_id,
       event,
       total_waits,
       ROUND(time_waited_micro / 1000000,2) seconds_waited
FROM gv$system_event
WHERE event LIKE 'gc%'
ORDER BY seconds_waited DESC;

You now have the foundation for building a real RAC interconnect health-check report rather than simply checking whether the interface is online.

Bookmark the permalink.
Loading Facebook Comments ...

Leave a Reply