Solace Insights Metrics and Checks

Solace Insights Advanced Monitoring shows metrics and checks (status information) on the various default dashboards provided by Solace. These metrics and checks provide you with in-depth information (statistics and states) collected from the event brokers in your estate. You can also use these metrics to build custom visualizations or customize existing dashboards. The information from the metrics and checks is useful to not only manage all aspects of your event brokers, but give you better insight to manage your event-driven architecture (EDA).

Some metrics and checks listed on this page are only available with Solace Remote Monitoring Service (RMMS). These metrics and checks are available for Solace Software Event Brokers, Solace Appliance Event Brokers, or non-standard deployments of Solace Cloud. These metrics and checks are not used in typical Solace Cloud deployments.

Customers using Disaster Recovery (DR) should be aware that Solace Insights does not collect certain metrics for the backup site, including Queues, Client Username, Topic Endpoints, and REST Delivery Point (RDP) metrics. Collecting these metrics for both the primary and backup sites would cause Insights to generate duplicate alerts. Insights continues to collect other metrics from the back up site to ensure the health of your DR configuration, including Queue metrics for backup sites with ACK disabled.

Unless otherwise noted in the table or metric description, all metrics and checks are available for all event broker types: Solace Cloud event broker, software event broker, and appliance event broker. Metrics and checks that are specific to a particular event broker type are indicated as such.

Here's a brief summary of the metrics and status information available:

Service Checks: These metrics characterize the status of a service for monitoring within Datadog. For more information about service checks, see the Datadog documentation. These are the service checks that are available with Insights:

Collected and Derived Metrics: There are two types of metrics as follows:

  • Collected — These are basic metrics collected directly from the event broker services. They typically represent measurements, statistics, and operational numbers provided by the event broker services.
  • Derived — These metrics are calculated or determined by some means. The metric is usually a combination of information sources that are used to calculate a value or status.

The collected and derived metrics are organized into the following alphabetical groupings:

  • Insights Metrics: A-B — Appliance Environmental Sensor Statistics, Bridge Statistics, Broker Resource Utilization Statistics
  • Insights Metrics: C — Cache Instance Statistics, Certificate Expiry Monitoring Statistics, Client Statistics, Client Username Statistics, Config-Sync Statistics
  • Insights Metrics: D — Disk Detailed Statistics, Disk Statistics, DMR Cluster Statistics, DMR Link Statistics, DNS Statistics
  • Insights Metrics: H-L — Hardware HBA Link Statistics, Health Statistics, Interface Statistics, Kafka Bridge Statistics, Limit Usage
  • Insights Metrics: Ma-Mem (Message Spool) — Maximum Guaranteed Message Size Statistics, Memory Statistics, Message Spool Statistics
  • Insights Metrics: Message VPN — Message VPN Detailed Statistics, Message VPN Message Spool Statistics, Message VPN Replication Statistics, Multi-Node Routing Statistics
  • Insights Metrics: P-R — Partitioned Queue Statistics, Queue Rate Statistics, Queue Statistics, Redundancy Statistics, Replication Statistics, REST Delivery Point Statistics
  • Insights Metrics: S-T — Server Certificate Statistics, Storage Statistics, Subscription Statistics, Time Server Statistics, Topic Endpoint Statistics

Filtering Metrics

Each Insights metric has a variety of tags associated with it. The tags provide information about where the metric's data was sourced. For example, the env tag allows you to determine the environment the data in the metric was derived from. You can filter the Insights metrics list using the tags to get data sourced from a specific object, for example, from a specific environment or event broker.

To filter the metrics list, follow these steps:

  1. In Datadog, go to MonitorsSummary to access the list of metrics.

  2. In the Tag field, either:

    • select a tag from the list of options

    • enter the tag you want to filter the list by. For example, enter env:MyFirstEnvironment to filter the metric list to show only metrics for the environment named MyFirstEnvironment.

Service Checks

These metrics characterize the status of a service for monitoring within Datadog. For more information about service checks, see the Datadog documentation.

Service Check Name Polling
Frequency
in Seconds
Description

Bridge

bridge.status

10

Whether the inbound and outbound Message VPN bridges are healthy and the associated queues are bound.

Cache Instance

cache_instance.status

10

Whether the cache instance is operational.

Config-Sync

config_sync.up_or_disabled

60

The Config-Sync status. A valid status is Shutdown or Up.

Disk

disk.status

10

The internal disk status. A valid status is disabled or up.

DNS

system.dns_status

60

The Domain Name System (DNS) reachability status.

DMR

dmr.cluster.status

10

Whether the Dynamic Message Routing (DMR) cluster is operational.

DMR link

dmr.link.status

10

Whether the DMR links and redundancy are operational.

Environment

system.sensor_status

60

This metric applies to appliance event brokers only.

System temperature and power redundancy status.

Hardware HBA link (For Appliances Only)

hba_link.status

60

This metric applies to appliance event brokers only.

The Host_Bus Adapter (HBA) link_state status.

Insights Agent Heartbeat (For Software and Appliance Event Brokers with Insights Forwarding)

datadog.agent.running 60

This metric is forwarded to third-party observability platforms when forwarding metrics and logs of software and appliance event brokers, along with your other event broker metrics and logs.

A value of 1 indicates Insights Agent is running

Includes ha_role and maas_ha_role attributes to distinguish between the primary, backup, and monitoring nodes in HA deployments

Interface

interface.status

30

The network-interface status. A valid status is disabled or up.

MNR

mnr.neighbor.status

10

Whether the Multi_Node Routing (MNR) neighbor status is OK.

Queue

queue.unbound_and_enabled_with_messages

30

Whether an enabled queue has messages, but no bound clients. If no clients are consuming from the enabled queue, the queue could fill.

Redundancy

derived_metrics.redundancy.up_or_disabled

10

The redundancy status. A valid status is Shutdown or Up.

derived_metrics.redundancy.adb_links_status

10

Whether the ADB Hello is Up.

derived_metrics.redundancy.is_service_active

10

Whether one event broker in the High-Availability (HA) group is active.

derived_metrics.message_spool_service_status

10

Whether the message spool disk on the active event broker is in the AD_ACTIVE state.

derived_metrics.message_spool_standby_status

10

Whether the message spool disk on the inactive node is in the AD_Standby state.

System POST

system.post_status

60

Whether the Power-On Self-Test (POST) is successful.

Time Server

ntp.in_sync

60

Whether the event broker clock is synchronized with the configured time server.

Topic Endpoint

topic_endpoint.unbound_and_enabled_with_messages

30

Whether an enabled topic endpoint has messages, but no bound clients. If no clients are consuming from the enabled topic endpoint, the topic endpoint could fill.

VPN Replication

vpn.replication_sync_eligible

10

Whether the replication service entered a degraded state.

Collected and Derived Metrics

The list of supported metrics and details that can be summarized as collected and derived metrics.

  • Collected: These are basic metrics collected directly from the event broker services. They typically represent measurements, statistics, and operational numbers provided by the event broker services.
  • Derived: These metrics are calculated or determined by some means. The metric is usually a combination of information sources that are used to calculate a value or status.

Common Metric Tags

These common tags are associated with every metric listed in the table below and are useful for filtering and aggregating metrics:

  • ha_role: The identity of the software event broker in a high-availability (HA) event broker for which the metric is measuring. The value can backup, monitoring, or primary.

  • peb_type: The type of event broker being monitored. The value can be cloud, software, or appliance. This tag is used to filter metrics and dashboards by event broker type.

  • service_id: The unique identifier of the event broker service, such as s1h0z84b1ar.

  • service_name: The name of the event broker service, such as My-First-Service.

  • service_type: The type of event broker service. This can be developer, standalone or enterprise.

  • service_class: The service class the metric, such as developer or enterprise-kilo.

Depending on the metric, these tags are also useful for aggregating and filtering the metrics:

  • vpn_name: The name of the Message VPN.

  • remote_name: The name of the remote Message VPN name in a bridge.

  • topic_endpoint: The name of the topic endpoint.

  • bridge_name: The name of the bridge.

  • cache_cluster_name: The name of the cache cluster.

  • cache_instance_name: The name of the cache instance.

  • cache_name: The name of the cache.

  • queue_name:The name of the queue, such as my-first-queue.

  • rdp_name: The name of the REST Delivery Point.

For detailed metric reference information, see the following pages:

  • Insights Metrics: A-B — Appliance Environmental Sensor Statistics, Bridge Statistics, Broker Resource Utilization Statistics

  • Insights Metrics: C — Cache Instance Statistics, Certificate Expiry Monitoring Statistics, Client Statistics, Client Username Statistics, Config-Sync Statistics

  • Insights Metrics: D — Disk Detailed Statistics, Disk Statistics, DMR Cluster Statistics, DMR Link Statistics, DNS Statistics

  • Insights Metrics: H-L — Hardware HBA Link Statistics, Health Statistics, Interface Statistics, Kafka Bridge Statistics, Limit Usage

  • Insights Metrics: Ma-Mem (Message Spool) — Maximum Guaranteed Message Size Statistics, Memory Statistics, Message Spool Statistics

  • Insights Metrics: Message VPN — Message VPN Detailed Statistics, Message VPN Message Spool Statistics, Message VPN Replication Statistics, Multi-Node Routing Statistics

  • Insights Metrics: P-R — Partitioned Queue Statistics, Queue Rate Statistics, Queue Statistics, Redundancy Statistics, Replication Statistics, REST Delivery Point Statistics

  • Insights Metrics: S-T — Server Certificate Statistics, Storage Statistics, Subscription Statistics, Time Server Statistics, Topic Endpoint Statistics