Collect Diagnostic Data
Collect and manage diagnostic data, configure automatic and on-demand collections, monitor system and operating system metrics, and manage Oracle Database and Oracle Grid Infrastructure logs using Oracle Autonomous Health Framework (AHF).
Managing and Configuring Oracle Trace File Analyzer
This section introduces management and configuration of Oracle Trace File Analyzer, covering its daemon, diagnostic collections, and repository administration.
Read More: Managing and Configuring Oracle Trace File Analyzer
Querying Oracle Trace File Analyzer Status and Configuration
Use the tfactl print command to inspect TFA status and configuration, including automatic collection, file trimming, repository size, trace verbosity, automatic purging, collection retention age, and alert-log scanning thresholds.
Read More: Querying Oracle Trace File Analyzer Status and Configuration
Managing the Oracle Trace File Analyzer Daemon
TFA starts automatically with the operating system through platform-specific init or service mechanisms; administrators can manually start or stop it with tfactl start and tfactl stop, and control automatic restart behavior with tfactl enable and tfactl disable.
Read More: Managing the Oracle Trace File Analyzer Daemon
Managing the Repository
TFA stores diagnostic collections in a repository governed by configurable size and location limits. Automatic purging, enabled by default, removes eligible collections from largest to smallest when disk space or repository thresholds are reached, while administrators can configure retention age, purge behavior, location, and maximum size or manually inspect and purge collections with repository and collection commands.
Read More: Managing the Repository
Managing Collections
Collection management includes configuring directories and access policies, controlling file trimming and core-file size limits, explicitly including cores in collections, and temporarily blacking out automatic collections for selected targets or events, including resources being provisioned.
Read More: Managing Collections
Configuring the Host
Host configuration requires root or sudo access and supports viewing, adding, removing, and synchronizing cluster hosts and their authentication certificates. New hosts are integrated with tfactl syncnodes, while existing synchronized certificates can be used with tfactl host add; default certificates may be replaced with self-signed or CA-signed certificates.
Read More: Configuring the Host
Configuring the Ports
Cluster TFA daemons communicate securely through ports 5000–5005 by default. Administrators can assign a primary port or up to five sequential ports with tfactl set port, then restart TFA on every node to apply the changes.
Read More: Configuring the Ports
Configuring SSL and SSL Certificates
TFA secures cluster communication with TLS protocols, restricts obsolete protocols, and supports configurable cipher suites plus self-signed or CA-signed certificates. Administrators can inspect and restrict protocols, create or import certificates and keystores, protect keystore permissions, apply settings with tfactl set sslconfig, and restart TFA; the default cipher suite can also be replaced with a JRE 1.8-supported suite.
Read More: Configuring SSL and SSL Certificates
Configuring Email Notification Details
Configure notification addresses globally or per operating-system owner and define SMTP host, authentication, security, sender, recipient, and debugging parameters with tfactl set smtp. Use tfactl print smtp to review settings and tfactl sendmail to validate delivery before relying on automatic collection notifications.
Read More: Configuring Email Notification Details
Managing the Index
TFA uses Lucene indexes for diagnostic and telemetry metadata. Trace-file metadata indexing is disabled by default beginning with AHF 25.2, while telemetry index corruption can be handled in recreate mode for faster recovery with possible data loss or restore mode, which preserves data through backups and redo at the cost of additional resources and recovery time.
Read More: Managing the Index
Using Automatic Diagnostic Collections
TFA monitors database and Clusterware logs for significant errors and events, consolidates trimmed diagnostics from cluster nodes, stores the results in its repository, and can notify recipients or upload collections for support. Automatic collection is enabled by default, uses a delay and event coalescing window to avoid duplicate collections, supports Cluster Health Advisor events, offers masking or sanitization of sensitive data, and provides flood-control settings to limit repeated collections.
Read More: Using Automatic Diagnostic Collections
Collecting Diagnostics and Analyzing Logs On-Demand
The tfactl interface provides consistent access to Oracle diagnostic tools and supports on-demand collections when an issue is identified. TFA gathers relevant data for a specified period, trims files for diagnosis, and packages the results on the node where the command is executed.
Read More: Collecting Diagnostics and Analyzing Logs On-Demand
Viewing System and Cluster Summary
The tfactl summary command provides a real-time overview of system and cluster status, with additional syntax guidance available through tfactl summary -help.
Read More: Viewing System and Cluster Summary
Investigating Logs for Errors
Use tfactl analyze to search cluster-wide logs for recent errors over a specified hours-or-days interval or to find occurrences of a particular error, such as ORA-00600, within a defined period.
Read More: Investigating Logs for Errors
Analyzing Logs Using the Oracle Database Support Tools
When the support-tools bundle is installed, TFA provides a unified interface to health checks, operating-system monitoring, database performance analysis, log searching, process inspection, configuration summaries, event reporting, log maintenance, and related tools across Linux, UNIX, and Windows. Administrators can verify installed tools with tfactl toolstatus and run them from command-line or interactive shell mode.
Read More: Analyzing Logs Using the Oracle Database Support Tools
Searching Oracle Trace File Analyzer Metadata
TFA metadata searches use tfactl search with JSON criteria to filter indexed events by content, database, and time range. The command can list available datatypes or return all matching events, with platform-specific quoting syntax for Linux, UNIX, AIX, Solaris, and Windows.
Read More: Searching Oracle Trace File Analyzer Metadata
Oracle Trace File Analyzer Service Request Data Collections
Service Request Data Collections provide predefined, problem-specific diagnostic profiles that gather and package the appropriate logs, reports, traces, and system information across local or cluster scopes. Administrators can run SRDCs interactively or silently, constrain collection time and database scope, tag or name output, and optionally upload results to a Service Request using configured upload settings.
Read More: Oracle Trace File Analyzer Service Request Data Collections
Diagnostic Upload
AHF provides generic upload configuration and execution for TFA, ORAchk, and EXAchk collections through HTTP, SQLNET, and SFTP endpoints. Using ahfctl, administrators can set, retrieve, validate, and remove named configurations that synchronize across cluster nodes, while tfactl or related tools can upload files during or after collection; the feature supports multiple users when installed as root but is not supported on Windows.
Read More: Diagnostic Upload
Performing Custom Collections
Oracle Trace File Analyzer (TFA) supports targeted diagnostic collections by customizing the time range, event, nodes, components, directories, file sizes, collection names, and packaging behavior. Use tfactl diagcollect with options such as -last, -from, -to, -for, -node, component selectors, -collectdir, -tag, -z, -nocopy, -notrim, -silent, and -cores; event-driven collections may invoke an associated SRDC and require database dba privileges. IPS packages can also be collected and managed, while large files can be limited through maxfilecollectionsize.
Read More: Performing Custom Collections
Limit the Maximum Amount of Memory Used by Oracle Trace File Analyzer
On Linux systems with a full, root-installed Autonomous Health Framework, TFA memory usage can be constrained using ahfctl setresourcelimit. Limits apply to automatic collections, on-demand collections, and analysis operations, with supported values from 150 MB to 2 GB or 25% of system memory, whichever is lower. Use the kmem resource for system memory or swmem for combined system and swap memory; limits are enabled by default at the maximum value.
Read More: Limit the Maximum Amount of Memory Used by Oracle Trace File Analyzer
Limit Oracle Trace File Analyzer CPU Usage
On Linux, TFA CPU consumption can be limited with ahfctl setresourcelimit -value. Values range from 0.5 CPU to 4 CPUs or 75% of available CPUs, whichever is lower, and the default limit is the maximum permitted value. For example, setting the value to 0.5 limits TFA to approximately half of one CPU.
Read More: Limit Oracle Trace File Analyzer CPU Usage
Automatic TFA Self-Diagnostics on diagcollect Failure for Faster Issue Resolution
When a root user’s manual tfactl diagcollect operation fails, TFA automatically runs tfactl diagnosetfa -local -profile collection and generates a local self-diagnostic bundle named in the hostname_ahf_autodiagnostic.zip format. The fallback covers service interruptions, stalled collectors, fatal collection errors, interruptions, and repository-related aborts, and displays the bundle location and recommended next steps. Remote-node failures and non-root failures require manually running the diagnostic command as root on the affected or local node and uploading the resulting bundles to the Service Request.
Read More: Automatic TFA Self-Diagnostics on diagcollect Failure for Faster Issue Resolution
Proactively Detecting and Diagnosing Performance Issues for Oracle RAC
Oracle Cluster Health Advisor provides early warnings, root-cause diagnoses, and corrective actions for emerging performance and availability issues affecting Oracle RAC databases and cluster nodes. It compares observed metrics with expected values from trained normal-operation models, detects significant anomalies, and reports targeted recommendations through Enterprise Manager Cloud Control. This release improves CHA detection on Exadata and adds support for Oracle Solaris RAC deployments.
Read More: Proactively Detecting and Diagnosing Performance Issues for Oracle RAC
Oracle Cluster Health Advisor Architecture
Oracle Cluster Health Advisor operates as the highly available ochad cluster resource on each node. Each daemon monitors its host operating system and optionally RAC database instances, receiving operating-system metrics from Cluster Health Monitor and database metrics through memory-mapped files without requiring database connections. The Health Prognostics Engine evaluates these metrics against selected models multiple times per minute for both nodes and monitored database instances.
Read More: Oracle Cluster Health Advisor Architecture
Removing Grid Infrastructure Management Repository
Because GIMR is desupported in Oracle AI Database 26ai, existing configurations should be removed after checking GIMR and Oracle Fleet Patching and Provisioning status with srvctl. The procedure involves preparing a grid-owned directory, obtaining and extracting the deletion script, optionally exporting CHA models and FPP metadata, and executing reposScript.sh -mode="Delete" as the grid user. FPP users should use its self-upgrade process where applicable, because deleting GIMR without upgrading and reconfiguring FPP stops FPP functionality.
Read More: Removing Grid Infrastructure Management Repository
Monitoring the Oracle Real Application Clusters Environment with Oracle Cluster Health Advisor
CHA is provisioned and enabled by default with RAC or RAC One Node Grid Infrastructure installations and begins node monitoring automatically when database instances are detected. Grid users can enable database monitoring with chactl monitor database -db, stop it with chactl unmonitor database -db, and review node and database monitoring status with chactl status, optionally using -verbose. CHA monitors each RAC instance independently but does not support single-instance databases.
Read More: Monitoring the Oracle Real Application Clusters Environment with Oracle Cluster Health Advisor
Using Cluster Health Advisor for Health Diagnosis
CHA autonomously detects and clears problems, while Grid users query stored diagnoses with chactl query diagnosis, specifying a database and time range in YYYY-MM-DD HH24:MI:SS format. Results identify detected problems, descriptions, causes, impacts, and targeted corrective actions, and can be saved as HTML with -htmlfile. Diagnoses can expose issues such as control-file I/O latency, excessive database CPU usage, and frequent log switches.
Read More: Using Cluster Health Advisor for Health Diagnosis
Calibrating an Oracle Cluster Health Advisor Model for a Cluster Deployment
CHA default models are designed to avoid false warnings, but chactl calibrate can improve sensitivity for a specific normal workload. Oracle recommends at least six hours of representative data with consistent time ranges across the cluster and databases; use query calibration to verify samples and filter abnormal data with KPI sets when necessary. After calibration, save the model under a user-defined name and activate it with chactl monitor cluster -model.
Read More: Calibrating an Oracle Cluster Health Advisor Model for a Cluster Deployment
Viewing the Details for an Oracle Cluster Health Advisor Model
Use chactl query model -name to inspect a CHA model’s target type, software version, operating-system platform, calibration target, calibration date, time ranges, and KPI settings. CHA models can also be renamed, imported, exported, or deleted as part of model administration.
Read More: Viewing the Details for an Oracle Cluster Health Advisor Model
Managing the Oracle Cluster Health Advisor Repository
When GIMR is configured, the CHA repository retains historical node and database problems, metric evidence, and models. Its default capacity supports 16 targets for 72 hours, with retention decreasing as targets increase; warnings occur below 72 hours and monitoring stops below 24 hours. Use chactl query repository to inspect capacity, chactl set maxretention to define the maximum retention period, and chactl resize repository -entities to expand support for additional targets.
Read More: Managing the Oracle Cluster Health Advisor Repository
Viewing the Status of Cluster Health Advisor
Use SRVCTL to check CHA service status and configuration across active hub and leaf nodes. srvctl status cha reports where the service is running, while srvctl config cha reports where it is enabled. A target is monitored only when it is running and its host node’s CHA service is active.
Read More: Viewing the Status of Cluster Health Advisor
Enhanced Cluster Health Advisor Support for Oracle Pluggable Databases
CHA expands support to as many as 4,000 PDBs, up from 256, improving suitability for large Autonomous Database deployments. Its enhanced detection and root-cause analysis considers events such as database reconfiguration, enabling more effective identification and preventive recommendations for issues including instance evictions.
Read More: Enhanced Cluster Health Advisor Support for Oracle Pluggable Databases
New Profile to Include Cluster Health Advisor Data in Oracle Orachk and Oracle Exachk Reports
AHF 25.9 adds the cha profile to Orachk and Exachk for collecting and reporting Cluster Health Advisor data. Running either tool with -profile cha produces a report containing only the CHA section, while -includeprofile cha adds that section to the standard report; the CHA profile is not enabled by default.
Read More: New Profile to Include Cluster Health Advisor Data in Oracle Orachk and Oracle Exachk Reports
Collecting Operating System Resources Metrics
Cluster Health Monitor and System Health Monitor are lightweight, high-availability daemons that collect operating-system metrics every five seconds with low overhead. They aggregate CPU, device, process, network, NFS, protocol, filesystem, and critical-resource data into Nodeview snapshots, perform inline analysis and summarization, and provide Clusterware-aware tagging and analysis that is more consistent and useful than manually combining standard OS utilities.
Read More: Collecting Operating System Resources Metrics
Comparing CHM and SHM - Understanding their fundamental differences
CHM, provided by Grid Infrastructure as osysmond, monitors RAC cluster nodes and stores Nodeview metrics in the GI repository, while SHM, provided by AHF as ahf-sysmon, supports single-instance and non-GI systems. CHM is managed by the GI high-availability stack and supports Linux, Solaris, AIX, zLinux, ARM64, and Windows; SHM is managed by AHF, supports Linux only, and stores metrics in the AHF repository. Both automatically archive and purge historical data and include their metrics in tfactl diagcollect.
Read More: Comparing CHM and SHM - Understanding their fundamental differences
Additional Details About System Health Monitor
SHM is enabled by default in AHF on Linux single-instance and non-GI systems, collecting real-time process, memory, network, I/O, and disk metrics for diagnostics and Insights analysis. Administrators can check status with ahfctl statusahf, enable or disable it through the enhanced OS metrics property, inspect its process and JSON data, and verify resource cgroup placement. TFA collections include SHM data, which can be validated by collecting and inspecting archives for SHM-related files.
Read More: Additional Details About System Health Monitor
Collecting Cluster Health Monitor Data
Cluster Health Monitor data can be collected from any cluster node, and Oracle recommends using tfactl diagcollect whenever an Oracle Clusterware error occurs so that relevant diagnostic information is gathered for investigation.
Read More: Collecting Cluster Health Monitor Data
Operating System Metrics Collected by Cluster Health Monitor and System Health Monitor
CHM and SHM organize operating-system observations into Nodeview metric sets covering CPUs, devices, processes, network interfaces, NFS, protocols, filesystems, system resources, and Oracle process aggregates. Metrics include utilization, throughput, queueing, latency, errors, memory and swap activity, file descriptors, process states, network packets, filesystem capacity, and aggregated CPU, memory, thread, and descriptor usage for database, ASM, Clusterware, and other process groups.
Read More: Operating System Metrics Collected by Cluster Health Monitor and System Health Monitor
Detecting Component Failures and Self-Healing Autonomously
CHM’s CHMDiag daemon receives component events through the CRFE API, validates and schedules corrective actions, monitors execution, and terminates actions that exceed configured time limits. It records event details, action results, and daemon logs under the CHMDIAG base directory, while oclumon chmdiag description, query, and collect provide event documentation, reports, and data collection for troubleshooting and autonomous self-healing.
Read More: Detecting Component Failures and Self-Healing Autonomously
CHM Inline Analysis
CHM Inline Analysis preserves summarized operating-system diagnostics instead of large raw CHM or SHM files. TFA processes hourly compressed data, stores the smaller analyzed results in the AHF repository, and runs as root under resource constraints. This reduces repository requirements to about 100 MB and extends retention from hours to months, improving the likelihood that useful OS data is available in Service Requests.
Read More: CHM Inline Analysis
Integrating System Health Monitor into AHF for Standalone Non-Root Installations
AHF integrates Linux-only SHM into standalone non-root Oracle Restart and single-instance installations, allowing non-root users to collect OS metrics and generate focused Insights reports. In the initial release, users manage the daemon with sysmonctl start and sysmonctl stop; SHM reads /proc, may collect less information than root installations, and should be validated under enforcing SELinux configurations.
Read More: Integrating System Health Monitor into AHF for Standalone Non-Root Installations
Creating an AHF Insights Report for Operating System Issues - Non-Root AHF Installation
To create an OS Insights report from SHM data in a non-root AHF installation, run the CHM driver from $AHF_HOME/chm/bin with the AHF Python interpreter and specify an output directory and start and end times. The driver processes, analyzes, and reports SHM data into Insights ZIP archives; verify source files under the AHF SHM data directory and view the report by opening web/index.html after extraction.
Read More: Creating an AHF Insights Report for Operating System Issues - Non-Root AHF Installation
Diagnostic Signature - HugePagesNotUtilized
AHF 25.11 introduces the HugePagesNotUtilized diagnostic signature, which detects when HugePages are configured but none are in use because total and free HugePages are equal. The signature produces a specific alert in the CHM analysis section of Orachk reports, helping identify memory configuration problems more quickly than generic utilization warnings.
Read More: Diagnostic Signature - HugePagesNotUtilized
Monitoring System Metrics for Cluster Nodes
Oracle recommends Oracle Enterprise Manager for routine Oracle Clusterware monitoring, complemented by Cluster Health Monitor for full-stack operating-system and cluster observation. Both are enabled by default for Oracle clusters, and administrators should also review Clusterware resource activity logs to monitor managed resources and investigate operational issues.
Read More: Monitoring System Metrics for Cluster Nodes
Monitoring Oracle Clusterware with Oracle Enterprise Manager
Oracle Enterprise Manager provides Cluster Database Home, Interconnects, and Cluster Database Performance views for monitoring RAC databases and Clusterware. These views report node, VIP, node application, alert-log, registry, voting-file, service, interconnect, throughput, error, load, host, cache-fusion, active-session, and database-throughput information. Historical views and Top Activity analysis help identify performance causes, resource needs, SQL or schema tuning opportunities, and cluster wait problems.
Read More: Monitoring Oracle Clusterware with Oracle Enterprise Manager
Monitoring Oracle Clusterware with Cluster Health Monitor
The OCLUMON command-line tool queries CHM repositories for node metrics over selected periods and supports administrative tasks. Commands can change debug levels, display CHM versions, dump Nodeview information, and manage the metrics datafile size.
Read More: Monitoring Oracle Clusterware with Cluster Health Monitor
Managing Oracle Database and Oracle Grid Infrastructure Logs
TFA provides commands for managing Oracle Database and Grid Infrastructure diagnostic data and for monitoring disk-usage snapshots, helping control diagnostic storage and maintain usable log repositories.
Read More: Managing Oracle Database and Oracle Grid Infrastructure Logs
Managing Automatic Diagnostic Repository Log and Trace Files
The tfactl managelogs command manages ADR files in alert, incident, trace, core-dump, health-monitor, diagnostic, and log directories. Administrators can restrict operations by file age, preview purge results with -dryrun, purge database or GI files with -purge, and inspect diagnostic-destination usage with -show usage; appropriate operating-system privileges are required.
Read More: Managing Automatic Diagnostic Repository Log and Trace Files
Managing Disk Usage Snapshots
TFA automatically records disk-usage snapshots in its repository, using a default 60-minute interval. Administrators can change the interval with tfactl set diskUsageMonInterval=minutes and enable or disable monitoring with tfactl set diskUsageMon=ON|OFF.
Read More: Managing Disk Usage Snapshots
Purging Oracle Database and Oracle Grid Infrastructure Logs
Automatic TFA log purging is enabled by default on Domain Service Clusters and disabled elsewhere; when enabled, it periodically removes logs older than the default 30-day age. Use manageLogsAutoPurge to toggle purging, manageLogsAutoPurgePolicyAge to set the retention age, and manageLogsAutoPurgeInterval to define the purge frequency.
Read More: Purging Oracle Database and Oracle Grid Infrastructure Logs
Securing Access to Diagnostic Collections
TFA restricts diagnostic commands to authorized users, allowing selected non-root operations for the Grid Infrastructure home owner and database home owners when TFA is root-installed on Linux or UNIX. Administrators can list, add, remove, remove all, or reset users through tfactl access commands, with optional local-node scope. Deleted operating-system usernames should also be removed from the AHF access list to prevent future users with the same name from inheriting old privileges.
Read More: Securing Access to Diagnostic Collections
Database Monitoring Using Database User Credentials
AHF supports database monitoring with a configured username and password instead of requiring SYSDBA access, although the default / as SYSDBA connection remains preferred because common-user access can limit diagnostic collection. Configure an AHF database user in the CDB for multitenant databases or the database for nonmultitenant environments, preferably grant the DBA role, or provide the required dictionary, directory, advisor, session, container, and catalog privileges. For multitenant databases, enable access to data across all containers with CONTAINER_DATA=ALL, then have the Oracle software owner store the credentials in the AHF Wallet using ahf security add-credentials; the database unique name must be used to uniquely identify the database.
Read More: Database Monitoring Using Database User Credentials