Oracle RAC 19c Masterclass — Part 13

GCS and GES: Understanding the Global Cache and Global Lock Architecture

In Part 12, we looked at Cache Fusion and how Oracle RAC transfers blocks between instances through the private interconnect.

Now we need to go one level deeper.

When Instance 1 owns a block and Instance 2 needs it, Oracle needs more than just a fast network. It needs a mechanism to coordinate:

  • Who owns the block?
  • Who can modify it?
  • Who can read it?
  • Which instance currently masters the resource?
  • Can another instance request the block?
  • What happens when multiple instances request the same resource?

This is where Global Cache Service (GCS) and Global Enqueue Service (GES) become important.

A good RAC DBA should understand these two components because many difficult RAC performance problems eventually come down to global cache coordination, lock contention, or excessive inter-instance communication.

1. GCS vs GES

The simplest way to remember the difference is:

GCS manages cached database blocks. GES manages global locks and resources.

ComponentMain responsibility
GCSCoordinates database blocks across instances
GESCoordinates global locks and enqueues
Cache FusionTransfers blocks between instances
Private InterconnectCarries RAC communication

A simplified architecture:

                 Oracle RAC
                     |
        +------------+------------+
        |                         |
    Instance 1                Instance 2
        |                         |
       GCS                       GCS
        |                         |
        +------ Cache Fusion -----+
        |                         |
       GES                       GES
        |                         |
        +---- Global Resources ---+

2. What Does GCS Actually Manage?

Suppose a block belongs to table:

CUSTOMERS

Instance 1 currently has the latest version in its buffer cache.

Instance 2 wants to update the same block.

Oracle needs to coordinate the request.

Conceptually:

Instance 2
    |
    | "I need this block"
    v
GCS
    |
    | "Instance 1 currently owns it"
    v
Instance 1
    |
    | Transfer / downgrade
    v
Instance 2

GCS coordinates the ownership and state of the block.

3. What Does GES Manage?

GES is responsible for global resources that require coordination across RAC instances.

Examples include:

  • Enqueues
  • Locks
  • Resource ownership
  • Global synchronization

Consider two transactions trying to modify the same logical resource:

Instance 1
    |
    | Lock request
    v
   GES
    ^
    | Lock request
    |
Instance 2

GES ensures that the global resource is coordinated correctly.

4. Why Do We Need Both?

Consider this scenario:

Instance 1
    |
    | Current block
    |
    +--------------------+
                         |
                    Instance 2

GCS answers:

“Where is the required database block and what state is it in?”

GES answers:

“Who has the required global resource or lock?”

Both mechanisms work together to maintain consistency across RAC instances.

5. Resource Mastering

One of the most important concepts in RAC internals is resource mastering.

Oracle assigns responsibility for managing particular global resources to RAC instances.

For example:

Resource A → Instance 1

Resource B → Instance 2

Resource C → Instance 1

The instance responsible for a resource is called its master.

The master maintains information about the resource and coordinates requests from other instances.

6. Why Does Resource Mastering Matter?

Suppose Instance 2 repeatedly accesses resources mastered by Instance 1.

Instance 2
   |
   | requests
   v
Instance 1
   |
   | response
   v
Instance 2

If this happens thousands of times per second, the private interconnect becomes heavily involved.

This can create:

  • Increased RAC messaging
  • Higher CPU overhead
  • Higher gc waits
  • Increased latency

This is one reason workload placement matters in RAC.

7. Dynamic Resource Remastering

Oracle can dynamically change resource mastering.

If workload patterns change, Oracle may determine that certain resources would be better mastered by another instance.

Conceptually:

Before:

Resource X → Instance 1

After workload changes:

Resource X → Instance 2

This process is called dynamic remastering.

It can reduce unnecessary cross-instance communication.

8. Static vs Dynamic Workload Patterns

Imagine a reporting application:

REPORTING_SERVICE
        |
        v
     ORCL2

and an OLTP application:

OLTP_SERVICE
        |
        v
     ORCL1

This type of workload separation can reduce unnecessary RAC communication.

Instead of:

OLTP → ORCL1
OLTP → ORCL2
OLTP → ORCL1
OLTP → ORCL2

we aim for:

OLTP → ORCL1
REPORTING → ORCL2

This doesn’t eliminate Cache Fusion, but it can significantly reduce unnecessary cross-instance activity for certain workloads.

9. GCS Wait Events

When sessions need blocks from another instance, RAC-related wait events may appear.

For example:

gc cr request

This generally indicates a session is waiting for a consistent-read block.

Another important event:

gc current request

This indicates a request for the current version of a block.

10. Investigating GC Waits

Start with:

SELECT inst_id,
       event,
       total_waits,
       time_waited_micro
FROM gv$system_event
WHERE event LIKE 'gc%'
ORDER BY time_waited_micro DESC;

This gives you a cluster-wide view.

Don’t automatically conclude that high gc waits mean a network problem.

The application workload may simply be generating a large amount of cross-instance block access.

11. Current Blocks vs CR Blocks

This distinction is important.

CR Block

A consistent-read version of a block.

Typically associated with queries requiring read consistency.

Current Block

The current version of the block.

Typically more important for modifications.

Therefore:

gc cr

and:

gc current

represent different types of global cache activity.

12. Measuring GC Statistics

You can inspect RAC statistics using:

SELECT inst_id,
       name,
       value
FROM gv$sysstat
WHERE name LIKE 'gc%'
ORDER BY inst_id, name;

Useful statistics include:

gc cr blocks received
gc cr blocks served
gc current blocks received
gc current blocks served

These counters help identify the volume of global cache activity.

13. Calculate the Workload by Instance

A simple query:

SELECT inst_id,
       name,
       value
FROM gv$sysstat
WHERE name IN (
    'gc cr blocks received',
    'gc current blocks received',
    'gc cr blocks served',
    'gc current blocks served'
)
ORDER BY inst_id, name;

Run it periodically rather than looking at a single snapshot.

The trend is much more useful than one number.

14. Finding RAC-Related SQL

If a particular SQL statement is generating excessive global cache activity, start by looking at SQL statistics.

For example:

SELECT inst_id,
       sql_id,
       executions,
       buffer_gets,
       cpu_time,
       elapsed_time
FROM gv$sql
ORDER BY elapsed_time DESC
FETCH FIRST 20 ROWS ONLY;

Then correlate the SQL with:

  • Application workload
  • Instance
  • Service
  • RAC wait events

15. Find Sessions Waiting on RAC Events

SELECT inst_id,
       sid,
       serial#,
       username,
       event,
       sql_id,
       seconds_in_wait
FROM gv$session
WHERE wait_class <> 'Idle'
AND event LIKE 'gc%'
ORDER BY seconds_in_wait DESC;

This is particularly useful during an active incident.

16. A More Practical RAC Investigation

Suppose you discover:

ORCL1:
gc current request → high

ORCL2:
gc current request → high

Don’t immediately change RAC parameters.

Ask:

Question 1

Which application is generating the workload?

Question 2

Which services are involved?

Question 3

Are sessions accessing the same objects from both instances?

Question 4

Are there hot blocks?

Question 5

Is the private interconnect healthy?

Question 6

Is the workload correctly distributed?

This is how a senior DBA approaches RAC performance.

17. Hot Block Example

Suppose an application repeatedly updates:

UPDATE orders
SET status = 'PROCESSED'
WHERE order_id = :id;

If many sessions from different RAC instances modify rows that reside in the same frequently accessed blocks, Oracle may have to transfer current blocks repeatedly.

Conceptually:

ORCL1
  |
  | Block 500
  |
  v
ORCL2
  |
  | Block 500
  |
  v
ORCL1

The block keeps moving.

This is commonly called block pinging.

18. Why Block Pinging Is Expensive

Every cross-instance block transfer requires:

  • Interconnect communication
  • Global cache coordination
  • CPU processing
  • Synchronization

The more frequently blocks move between instances, the greater the overhead.

Therefore:

The goal isn’t to eliminate Cache Fusion. The goal is to avoid unnecessary Cache Fusion.

19. Example: Bad Workload Distribution

Imagine:

Application
     |
     +----------------+
     |                |
   ORCL1            ORCL2
     |                |
     +-------+--------+
             |
       Same hot table

Both instances continuously modify the same blocks.

This can result in:

High gc current
High interconnect traffic
Higher latency

20. Better Workload Design

Use services to logically separate workloads:

                    Applications
                         |
              +----------+----------+
              |                     |
          OLTP_SERVICE        REPORTING_SERVICE
              |                     |
            ORCL1                 ORCL2

Now the workloads are more predictable.

This doesn’t guarantee that blocks will never cross instances, but it can substantially reduce unnecessary global cache traffic.

21. Private Interconnect Matters

Because GCS and GES communicate heavily over the private interconnect, network performance is critical.

Check the configured interfaces:

oifcfg getif

Example:

eth0  192.168.10.0  global public
eth1  192.168.20.0  global cluster_interconnect

The private interface should be:

  • Dedicated
  • Low latency
  • Reliable
  • Properly sized
  • Redundant where appropriate

22. Check Network Interfaces

Linux:

ip addr

Check errors:

ip -s link

Check interface statistics:

sar -n DEV 1 10

Look for:

  • Packet errors
  • Drops
  • Retransmissions
  • Saturation

23. Check RAC Interconnect Configuration

Oracle:

SELECT inst_id,
       name,
       ip_address,
       is_public
FROM gv$cluster_interconnects;

This helps identify the interconnect interface used by each instance.

24. RAC Interconnect Latency

A healthy RAC interconnect should have extremely low latency.

For example:

ping <private-ip>

For deeper network testing, use tools appropriate to your infrastructure and security policies.

Don’t evaluate RAC interconnect performance solely from a normal ping. Application-level latency and packet behavior can be different.

25. GES Monitoring

GES information can be investigated using:

SELECT *
FROM gv$ges_statistics;

You can also inspect global resources:

SELECT inst_id,
       resource_name,
       state
FROM gv$ges_resource;

These views become especially useful when investigating global lock contention.

26. RAC Lock Contention

Consider:

Transaction A
Instance 1
     |
     | Holds resource
     v

Transaction B
Instance 2
     |
     | Requests same resource
     v

Wait

GES coordinates the global resource.

The problem may appear as a lock wait, but the root cause can involve application design.

27. Don’t Confuse RAC Waits With Database Bugs

Seeing:

gc current request

doesn’t automatically mean Oracle RAC is malfunctioning.

It may simply indicate:

Application
      ↓
Cross-instance access
      ↓
Cache Fusion
      ↓
Normal RAC operation

The question is whether the amount of global activity is appropriate for the workload.

28. Production Example

Imagine a four-node RAC:

ORCL1
ORCL2
ORCL3
ORCL4

The application distributes connections equally across all four instances.

Performance is acceptable with 10,000 transactions per minute.

After business growth, the workload increases to 50,000 transactions per minute.

Suddenly:

gc current request

increases significantly.

CPU remains moderate.

Storage latency is normal.

Network utilization increases.

The investigation reveals that the application is heavily updating the same set of tables from all four instances.

The solution isn’t necessarily:

Add CPU
Add memory
Add RAC nodes

Instead:

  • Review service placement.
  • Identify hot objects.
  • Review application transaction design.
  • Reduce cross-instance modifications.
  • Consider partitioning where appropriate.
  • Validate interconnect capacity.

This is a classic example of why RAC tuning must include the application architecture.

29. A Practical RAC Health-Check Script

You can start building your own RAC diagnostic toolkit with:

SELECT inst_id,
       event,
       total_waits,
       time_waited_micro
FROM gv$system_event
WHERE event LIKE 'gc%'
ORDER BY time_waited_micro DESC;

Then:

SELECT inst_id,
       name,
       value
FROM gv$sysstat
WHERE name LIKE 'gc%'
ORDER BY inst_id, name;

And:

SELECT inst_id,
       sid,
       serial#,
       username,
       event,
       sql_id,
       seconds_in_wait
FROM gv$session
WHERE wait_class <> 'Idle'
AND event LIKE 'gc%'
ORDER BY seconds_in_wait DESC;

Finally:

SELECT inst_id,
       name,
       ip_address,
       is_public
FROM gv$cluster_interconnects;

These four queries provide a useful starting point for a RAC Cache Fusion investigation.

30. What a Senior DBA Should Look For

When analyzing GCS/GES issues, don’t just look at one wait event.

Build a relationship between:

SQL
 ↓
Session
 ↓
Service
 ↓
Instance
 ↓
Object
 ↓
Cache Fusion
 ↓
Interconnect

For example:

SQL_ID
  ↓
High gc current request
  ↓
Service SALES
  ↓
ORCL1 + ORCL2
  ↓
HOT_TABLE
  ↓
High block transfers

Now you have a potential root-cause path.

That’s much more valuable than simply saying:

“The RAC has high GC waits.”

Best Practices

1. Design services around workloads

Separate OLTP, reporting, batch, and other workloads where appropriate.

2. Keep the private interconnect healthy

Monitor:

  • Latency
  • Packet loss
  • Errors
  • Capacity

3. Monitor GC waits

Look for trends rather than isolated values.

4. Investigate hot blocks

High gc current activity may indicate application-level contention.

5. Don’t tune RAC parameters blindly

Parameters such as RAC-related timeouts and resource-management settings should not be changed simply because a wait event appears in AWR.

First establish the root cause.

6. Correlate database and infrastructure metrics

RAC performance requires visibility across:

Database
+
Clusterware
+
ASM
+
Network
+
Storage
+
Application

DBA Tip

When you see high RAC waits, ask one question before changing anything:

“Why does this block need to move between instances?”

That question often takes you directly to the real problem.

The answer may be:

  • Normal RAC activity
  • Poor service placement
  • Hot blocks
  • Application design
  • Excessive cross-instance updates
  • Network latency
  • Resource contention

The wait event is only the symptom.

An expert RAC DBA doesn’t tune the gc wait. The expert finds out why the application is generating the gc wait.

Conclusion

GCS and GES are fundamental components of Oracle RAC.

GCS coordinates cached database blocks, while GES manages global locks and resources. Together, they allow multiple Oracle instances to work against the same database while maintaining consistency.

For production RAC environments, understanding these components is essential for:

  • Performance tuning
  • AWR analysis
  • Cache Fusion troubleshooting
  • Hot-block analysis
  • Interconnect troubleshooting
  • Workload design

The most important lesson is that RAC performance is not simply about CPU, memory, or storage. The way workloads are distributed across instances can have a major impact on Cache Fusion traffic and application response time.

Bookmark the permalink.
Loading Facebook Comments ...

Leave a Reply