Oracle RAC 19c Masterclass – Part 10

Oracle RAC Startup Sequence: From Server Boot to an Open Database

One of the biggest differences between a junior and a senior Oracle RAC DBA is knowing how Oracle starts.

When a RAC node boots, dozens of Oracle processes start in a specific order. Every component depends on the previous one.

If one component fails, everything above it also fails.

For example:

  • If ASM cannot start, the OCR cannot be accessed.
  • If the OCR cannot be read, Cluster Ready Services (CRSD) cannot start.
  • If CRSD is unavailable, databases and listeners remain offline.

Understanding the startup sequence allows you to troubleshoot problems logically instead of restarting services randomly.

In this article, we’ll follow the complete startup process from the Linux operating system to a fully operational RAC database.

The Complete Startup Sequence

The simplified startup flow looks like this:

Linux Boot
     │
     ▼
Oracle High Availability Services (OHASD)
     │
     ▼
Oracle Local Registry (OLR)
     │
     ▼
Grid Plug and Play (GPnP)
     │
     ▼
Oracle ASM
     │
     ▼
Oracle Cluster Registry (OCR)
     │
     ▼
Cluster Synchronization Services (CSSD)
     │
     ▼
Cluster Ready Services (CRSD)
     │
     ▼
Event Manager (EVMD)
     │
     ▼
Network Resources
(VIP, SCAN, Listeners)
     │
     ▼
ASM Disk Groups
     │
     ▼
Database Instances
     │
     ▼
Database Services
     │
     ▼
Applications

Each stage depends on the successful completion of the previous one.

Step 1 – Linux Boot

Everything starts with the operating system.

After Linux boots:

  • Network interfaces are initialized.
  • Storage devices are discovered.
  • Multipath services start.
  • Time synchronization begins.
  • System services become available.

Verify the server is healthy:

uptime
systemctl status

Check available memory:

free -h

Check storage:

df -h

If the operating system is unhealthy, Oracle RAC cannot start correctly.

Step 2 – Oracle High Availability Services (OHASD)

Once Linux is ready, Oracle starts OHASD.

OHASD is responsible for bootstrapping the Grid Infrastructure stack.

Verify:

crsctl check has

Example:

CRS-4638:
Oracle High Availability Services is online

If OHAS is not running, nothing else in Grid Infrastructure will start.

Step 3 – Read the Oracle Local Registry (OLR)

OHASD reads the local registry.

The OLR contains:

  • Local startup information
  • GPnP configuration
  • Local Clusterware resources

Verify:

ocrcheck -local

Expected output:

Oracle Local Registry integrity check succeeded.

Step 4 – GPnP Initialization

Next, Oracle loads the Grid Plug and Play profile.

GPnP provides:

  • Cluster identity
  • ASM discovery string
  • Network configuration
  • Cluster profile

Display the profile:

gpnptool get

Without GPnP, Oracle cannot discover ASM.

Step 5 – ASM Starts

Now Oracle can locate ASM.

ASM provides access to:

  • OCR
  • Voting Disks
  • Database files
  • Recovery files

Verify:

srvctl status asm

or

ps -ef | grep asm_pmon

Example:

+ASM1 is running

Step 6 – OCR Becomes Available

Once ASM is running, Oracle can access the OCR.

The OCR contains:

  • Cluster configuration
  • Database resources
  • Services
  • Listeners
  • VIPs

Verify:

ocrcheck

Example:

Cluster registry integrity check succeeded.

Step 7 – CSSD Starts

CSSD (Cluster Synchronization Services) starts next.

Responsibilities:

  • Cluster membership
  • Voting Disks
  • Heartbeats
  • Node eviction
  • Split-brain protection

Verify:

crsctl check css

Example:

CRS-4529:
Cluster Synchronization Services is online

Step 8 – CRSD Starts

CRSD is now able to manage cluster resources.

It reads the OCR and determines:

  • Which databases to start
  • Which listeners to start
  • Which VIPs belong to the node
  • Service placement

Verify:

crsctl check crs

Example:

CRS-4537:
Cluster Ready Services is online

Step 9 – EVMD Starts

The Event Manager Daemon (EVMD) handles cluster events.

Examples:

  • Resource failures
  • Node joins
  • Node leaves
  • Service relocation

Verify:

crsctl check evmd

Step 10 – Network Resources Start

CRSD starts network resources.

These include:

  • VIP
  • Local Listener
  • SCAN Listener
  • Network resources

Verify:

srvctl status listener
srvctl status scan_listener
srvctl status vip

Step 11 – ASM Disk Groups Mount

ASM mounts the required disk groups.

Example:

asmcmd lsdg

Example output:

OCR
DATA
FRA

Verify each required disk group is mounted before database startup.

Step 12 – Database Instances Start

CRSD starts the database instances.

Verify:

srvctl status database -d ORCL

Example:

Instance ORCL1 is running
Instance ORCL2 is running

Step 13 – Database Opens

Verify:

SELECT instance_name,
       status
FROM v$instance;

Example:

INSTANCE_NAME   STATUS
ORCL1 OPEN

Step 14 – Services Start

Oracle starts database services.

Verify:

srvctl status service -d ORCL

Example:

Service SALES is running
Service REPORTING is running

Applications can now connect.

Visual Startup Timeline

Linux
│
OHASD
│
OLR
│
GPnP
│
ASM
│
OCR
│
CSSD
│
CRSD
│
EVMD
│
VIP
Listeners
SCAN
│
ASM Disk Groups
│
Database
│
Services
│
Users

This sequence is worth memorizing—it explains many RAC startup failures.

Common Startup Failures

OHAS Does Not Start

Verify:

crsctl check has

Check:

  • Operating system logs
  • OLR
  • Grid Infrastructure binaries

ASM Does Not Start

Verify:

srvctl status asm

Possible causes:

  • Missing disks
  • Incorrect ASM_DISKSTRING
  • Permission issues
  • Storage failure

OCR Cannot Be Read

Verify:

ocrcheck

Possible causes:

  • ASM offline
  • OCR corruption
  • Storage unavailable

CSSD Offline

Verify:

crsctl check css

Possible causes:

  • Voting Disk issue
  • Private interconnect failure
  • Storage latency

Database Does Not Start

Verify:

srvctl status database -d ORCL

Then review:

  • Alert log
  • Listener status
  • ASM disk groups
  • Database services

Useful Troubleshooting Commands

Check the entire stack:

crsctl check cluster

View all Clusterware resources:

crsctl stat res -t

Check OHAS:

crsctl check has

Check ASM:

srvctl status asm

Check listeners:

srvctl status listener

Check database:

srvctl status database -d ORCL

Check services:

srvctl status service -d ORCL

These commands provide a quick overview of the RAC environment and are often the first ones to run during an incident.

Production Scenario

A customer reported that both database instances were down after a planned server reboot.

The database alert log showed no obvious errors, so the initial assumption was a database issue.

A quick review of the startup sequence revealed:

  • Linux: OK
  • OHAS: Running
  • OLR: Healthy
  • GPnP: Loaded
  • ASM: Not running

Further investigation showed that the SAN team had not presented the ASM disks after the maintenance window.

Because ASM could not start:

  • The OCR was unavailable.
  • CRSD could not initialize.
  • Database instances never started.

The issue was resolved after restoring storage connectivity and restarting the Grid Infrastructure stack.

The lesson was straightforward: the database was never the problem—it simply couldn’t start because an earlier dependency had failed.

Best Practices

  • Learn the startup order and troubleshoot in the same sequence.
  • Use srvctl to start and stop Clusterware-managed resources.
  • Verify ASM before checking the database.
  • Monitor OCR and Voting Disk health regularly.
  • Keep operating system, storage, and network teams informed during maintenance windows.
  • Review Clusterware logs before restarting services.

DBA Tip

One habit that has saved me countless hours is never skipping layers during troubleshooting.

If someone says, “The database won’t start,” resist the urge to open the alert log immediately.

Instead, ask:

  1. Is Linux healthy?
  2. Is OHAS running?
  3. Is the OLR accessible?
  4. Has GPnP initialized?
  5. Is ASM running?
  6. Is the OCR accessible?
  7. Is CSSD online?
  8. Is CRSD online?
  9. Are listeners and services running?

By checking dependencies in order, you’ll usually identify the root cause much faster than jumping directly to the database.

Oracle RAC is like a chain. When one link breaks, every component above it is affected. Find the first broken link—not the last symptom.

Conclusion

The Oracle RAC startup process is a carefully orchestrated sequence of components, each depending on the successful initialization of the previous one. Understanding this sequence transforms troubleshooting from guesswork into a structured process.

Whether you’re dealing with ASM startup failures, OCR corruption, CSSD issues, or databases that refuse to open, following the startup chain will help you isolate the problem quickly and confidently.

Bookmark the permalink.
Loading Facebook Comments ...

Leave a Reply