> For the complete documentation index, see [llms.txt](https://docs.expel.io/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.expel.io/connect-your-technology/security-devices/security-device-health.md).

# Security Device Health

Health checks serve as the first line of defense for ensuring data ingestion into Workbench. In most cases, you are responsible for resolving a failed health check that results in an unhealthy device, but there are [a few exceptions](#health-statuses-expel-must-resolve).

## About Health Checks

### Types of Health Checks

There are three types of health checks.

<details>

<summary>API Health Checks</summary>

API health checks apply to devices with “direct” integrations to Workbench via an API connection. These checks focus on ensuring there is a viable connection between your integration(s) and Workbench so that events can be ingested.

</details>

<details>

<summary>Via SIEM Health Checks</summary>

Via SIEM health checks apply to devices that are connected to Workbench via a SIEM (because a direct API connection is not available). There are two checks available:

1. The **configuration check** looks for the correct credentials and query substitution fields to verify the query syntax.
2. The **parent SIEM check** looks at the health of the parent SIEM device and updates the status of any child devices accordingly.

</details>

<details>

<summary>Device Health Monitoring (DHM) Health Checks</summary>

DHM health checks are done by Workbench to identify possible ingestion anomalies. DHM leverages statistical forecasting.

**DHM Philosophy**

When we are monitoring the health of your device, we are not just looking for gaps in connectivity (these are your API and via SIEM health checks) but also evaluating the overall pattern of data ingestion behavior. Our DHM health checks apply time-series forecasting models (Prophet and ARIMA) to determine what “healthy” or “normal” looks like for each device.

When a possible ingestion anomaly is identified, we triage the anomaly and work to determine if the model produced a false positive or a true positive. We will reach out if we identify a true positive.

</details>

{% hint style="info" %}
Health checks are not available for collector devices, and this will be reflected in the status shown for those devices.
{% endhint %}

### Health Check Failures

If any health check fails, the status of the device (and any via SIEM child devices) immediately changes to unhealthy. You will also automatically be alerted to a change in device status via your [default Workbench notifications](/workbench-setup/notifications/about-notifications.md#default-notifications) (to a channel like Slack, Teams, Webhooks, etc.) if you have [set up the platform integration](/workbench-setup/notifications/about-notifications.md#supported-platforms) for your organization. [Email notifications](/workbench-setup/notifications/manage-email-notifications.md) can be enabled as well.

Because device health is part of your default Workbench notifications, it is not necessary to manually check Workbench for a device health status change after you have enabled organization notifications. You will receive an alert from Ruxie that looks like this:

<div align="left"><figure><img src="/files/ZqITHoUQJcAqiBnWoBow" alt="Image of a Ruxie alert showing the problem and next steps."><figcaption></figcaption></figure></div>

{% hint style="info" %}
For via SIEM devices, the child device will become unhealthy if the parent device becomes unhealthy (and it will return to a healthy status only when the parent device is healthy). For example, if a Zscaler (via SIEM) device is connected to Workbench through a Splunk device, and the Splunk device becomes unhealthy, then the Zscaler (via SIEM) device will also become unhealthy.
{% endhint %}

### Recommended Preventive Device Health Tasks

While Workbench can detect when data ingestion stops, it cannot determine the root cause. To prevent unnecessary interruptions to data ingestion and the resulting unhealthy device status, we recommend you:

* Monitor for planned maintenance or configuration changes that might affect data ingestion
* Inform Expel of any scheduled downtime

## View Device Health

There are multiple ways to view the health of your device. The [Security Devices](#user-content-fn-1)[^1] page is one of them, but you can also use the [Alert Analysis Dashboard](/workbench-reference/dashboards/alert-analysis-dashboard.md).&#x20;

For a high-level snapshot of a device's connectivity status, use the **Status** column to check on the connection. If a problem exists, you will see a message here along with a timestamp for when the device became unhealthy.

* Most devices should indicate a Healthy API connection soon after the new device configuration has been saved. If this does not happen for a new device, your first step should be to check your device configuration for errors.
* See the [Reference](#reference) section for a full list of unhealthy statuses.&#x20;

{% hint style="info" %}
You may occasionally see a status of "No alerts mode" which indicates that polling has been intentionally disabled or is not supported, and we are not ingesting data at this time. [Contact Support](/support/how-to-reach-us.md) if you need help with this status or if it appears unexpectedly on a device.
{% endhint %}

<figure><img src="/files/c9TC2I8yx3WmcauLk9wq" alt="Image of Security Devices page focusing on the status column."><figcaption></figcaption></figure>

## Fix an Unhealthy Device

If you see an unhealthy message in the status column or receive a notification about an unhealthy device from Ruxie, you will usually need to resolve the issue yourself (there are [a few exceptions](#health-statuses-expel-must-resolve)). For via SIEM devices, you may need to fix the parent device in order to restore the child device to a healthy status.

<details>

<summary>Fix a newly added device that shows as unhealthy after onboarding.</summary>

1. Check your device configuration for errors.
2. If you are unable to find the issue within your configuration, or if your device becomes healthy but does not begin polling within 15 minutes and receiving data within 30 minutes, [contact Support](/support/how-to-reach-us.md) for help.

</details>

<details>

<summary>Fix an existing device that was previously healthy.</summary>

* Follow the instructions given to you by Ruxie, or
* Go into the device details (see [View Security Device Details](/connect-your-technology/security-devices/manage-security-devices.md#view-security-device-details) for more information about this screen and how to access it) and look for a matching "unhealthy" banner, then select the small arrow to display the steps to fix.&#x20;

{% hint style="warning" %}
For via SIEM devices, you may need to also check the parent device.
{% endhint %}

</details>

{% hint style="info" %}
See the [Reference](#reference) for a full list of health statuses, error messages, and the general actions required to resolve each one.&#x20;
{% endhint %}

<figure><img src="/files/XabkxQoobCl9FlThEhBW" alt="Example device details page for GitLab, showing an unhealthy device and the steps to fix."><figcaption></figcaption></figure>

## Reference

### Health Statuses You Must Resolve

The following charts show some specific definitions for unhealthy statuses that must be resolved by you, and the required action(s) to take.

<details open>

<summary>API (Direct) Connections</summary>

<table><thead><tr><th width="206.82421875">Status Message</th><th width="172.59765625">Meaning</th><th>Your Required Action</th></tr></thead><tbody><tr><td>Connection refused: Check that device is running and can be reached.</td><td>The security device refused our connection attempt. </td><td><ul><li><a href="/pages/yVrI2EJ0ezE94fTKOCfn#view-security-device-details">View the security device details</a> and validate the configuration.</li><li>Verify that you can log into the device.</li><li>Verify that your <a href="/pages/yCfOYtrsOevLRWO6FP6f">firewalls</a> allow traffic to the device.</li></ul></td></tr><tr><td>Failed to connect: Check device hostname is correct and Assembler has correct DNS server.</td><td>Workbench cannot connect due to a DNS lookup failure. </td><td><ul><li><a href="/pages/yVrI2EJ0ezE94fTKOCfn#view-security-device-details">View the security device details</a> and verify the device hostname is correct. </li><li><a href="/pages/Ex7GwaLMG9xNq3SXV3TZ">Check the Assembler</a> and verify it has the correct DNS server. </li></ul></td></tr><tr><td>Device error: &#x3C;error message>. If necessary, contact Expel for further assistance</td><td>The security device is reporting an internal error. </td><td><ul><li><a href="/pages/yVrI2EJ0ezE94fTKOCfn#view-security-device-details">View the security device details</a> and verify that it is polling and receiving data.</li><li>Follow any troubleshooting procedures from the device vendor. </li></ul></td></tr><tr><td>Failed to connect to device API: Check that server URL is correct. If necessary, contact Expel for further assistance.</td><td>We received an unexpected response when we tried to access the security device, and cannot connect.</td><td><ul><li><a href="/pages/yVrI2EJ0ezE94fTKOCfn#view-security-device-details">View the security device details</a> and verify that the server URL is correct.</li><li>Verify that the server URL is using the correct port. </li></ul></td></tr><tr><td>Device has invalid credentials.</td><td>We are unable to access the integration with the credentials supplied. </td><td><a href="/pages/yVrI2EJ0ezE94fTKOCfn#edit-a-security-device">Edit the security device</a> and update the credentials.</td></tr><tr><td>License expired.</td><td>The security device is reporting an expired vendor license and is preventing us from using it.</td><td>Renew the license with the security product vendor, or delete the device from Workbench if you are no longer using it. </td></tr><tr><td>Failed to connect: Check device address is correct.</td><td>Workbench cannot connect to the integration from the assembler due to a network routing problem. </td><td><a href="/pages/yVrI2EJ0ezE94fTKOCfn#view-security-device-details">View the security device details</a> and validate the configuration.</td></tr><tr><td>Device credentials don't have permission to connect.</td><td>The supplied credentials are valid, but Workbench does not have permission to perform the necessary actions. </td><td>Review the setup guide to ensure the correct permissions are assigned to the Expel user account.  </td></tr><tr><td>Timeout problem: If necessary, contact Expel for further assistance.</td><td>Our attempt to retrieve data from the security device timed out. </td><td><ul><li><a href="/pages/yVrI2EJ0ezE94fTKOCfn#view-security-device-details">View the security device details</a> and validate the configuration.</li><li>Verify that your <a href="/pages/yCfOYtrsOevLRWO6FP6f">firewalls</a> allow traffic to the device.</li></ul></td></tr><tr><td>Unsupported version:&#x3C;error message>.</td><td>We detected that this security device is running a version we do not support.</td><td>Modify the version or service tier for the device. </td></tr></tbody></table>

</details>

<details>

<summary>Via SIEM Connections</summary>

<table><thead><tr><th width="209.98046875">Status Message</th><th width="176.0625">Meaning</th><th>Your Required Action(s)</th></tr></thead><tbody><tr><td>&#x3C;unique error message from device></td><td>The SIEM’s query syntax is improperly defined due incorrect credentials or query substitution fields.</td><td>Validate the correct credentials and query substitution fields are used.</td></tr><tr><td>Parent SIEM is not active.</td><td>The SIEM through which this child (via SIEM) security device is connected to Workbench is unhealthy.</td><td>Go to the parent SIEM and resolve the unhealthy status message. This action will restore the child device to a healthy status. </td></tr></tbody></table>

</details>

### Health Statuses Expel Must Resolve

The following statuses must be resolved by Expel.

<table><thead><tr><th width="131.07421875">Connection Type</th><th width="148.62890625">Status Message</th><th width="142.3984375">Meaning</th><th>Expel's Required Action(s)</th></tr></thead><tbody><tr><td>API</td><td>Device is rate limited: Please increase query rate limit for the device.</td><td>You have hit your rate limit. This may require further investigation.</td><td>Expel must investigate the cause of the rate limit, and work to resolve any issues or contact you directly to discuss options for increase.</td></tr><tr><td>API</td><td>Unknown error: Expel will investigate and contact you if action is needed.</td><td>An unknown error has occurred.</td><td>Expel must investigate the root cause and work to resolve the error.</td></tr><tr><td>Via SIEM</td><td>Unknown error: Expel will investigate and contact you if action is needed.</td><td>One or more of the detection queries for this (via SIEM) device is not working. </td><td>Expel must investigate and resolve the issue. </td></tr><tr><td>N/A</td><td>No alerts mode</td><td>Polling has been intentionally disabled or is not supported, and we are not ingesting data at this time.</td><td>This is an intentional status set by Expel. <a href="/spaces/BwLjtVgtfcQ9hAVBO4rZ/pages/LThc2RqOxBKU56Qt3TMy">Contact Support</a> if you need help with this status or if it appears unexpectedly on a device.</td></tr></tbody></table>

### AWS Errors

AWS users may also see the following errors as a result of health checks, and these errors must be resolved by you. The error itself will show in the Device Details within "Steps to Fix."

| Error (located in Device Details)                                                                                | Status Message                  | Meaning                                                                                                 | Your Required Action(s)                                                                                                                                                                                                                                                                      |
| ---------------------------------------------------------------------------------------------------------------- | ------------------------------- | ------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| <p>STS:AssumeRole</p><p>ClientError: An error occurred (AccessDenied) when calling the AssumeRole operation.</p> | Device has invalid credentials. | Expel does not have the permission to use AssumeRole.                                                   | Verify that the correct organization GUID is applied to the ExternalId in the [IAM policy](https://docs.aws.amazon.com/IAM/latest/UserGuide/access_policies_manage-attach-detach.html). Refer to the onboarding guide for assistance.                                                        |
| AWS\_FILE\_DECRYPT\_FORBIDDEN                                                                                    | Device has invalid credentials. | The Expel role within your AWS environment does not have permission to decrypt the CloudTrail log file. | <ul><li>Verify the KMS:Decrypt permission is set to “Allow” in the <a href="https://docs.aws.amazon.com/kms/latest/developerguide/find-cmk-id-arn.html">KMS Key ARN</a> for your policy.</li><li>Verify the KMS:Decrypt permission is applied to the correct KMS Key ARN resource.</li></ul> |
| AWS\_FILE\_READ\_FORBIDDEN                                                                                       | Device has invalid credentials. | The Expel role within your AWS environment does not have permission to read the CloudTrail log file.    | <ul><li>Verify the get object permission is set to “Allow” in the <a href="https://docs.aws.amazon.com/IAM/latest/UserGuide/access_policies_manage-attach-detach.html">IAM policy</a>.</li><li>Verify the object permission has the correct S3 bucket resource ARN applied.</li></ul>        |
| AWS\_INVALID\_JSON                                                                                               | Device has invalid credentials. | The CloudTrail log files are not in JSON format.                                                        | Find the [log files](https://docs.aws.amazon.com/awscloudtrail/latest/userguide/cloudtrail-working-with-log-files.html) in your S3 bucket and make sure they are in JSON format.                                                                                                             |
| AWS\_LOG\_FILE\_NOT\_FOUND                                                                                       | Device has invalid credentials. | No log file was found in the S3 bucket.                                                                 | <ul><li>Make sure the notification service is connected to the correct <a href="https://docs.aws.amazon.com/AmazonS3/latest/userguide/enable-cloudtrail-logging-for-s3.html">S3 bucket</a>. </li><li>Confirm the log file was not deleted.</li></ul>                                         |
| AWS\_NON\_TAR\_FILE                                                                                              | Device has invalid credentials. | The CloudTrail log file is in the incorrect format.                                                     | Confirm the [S3 bucket](https://docs.aws.amazon.com/AmazonS3/latest/userguide/enable-cloudtrail-logging-for-s3.html) contains compressed JSON files ending with the .gz file extension.                                                                                                      |
| AWS\_NO\_RECORDS\_IN\_S3\_FILE                                                                                   | Device has invalid credentials. | No records were found in the S3 log file.                                                               | <ul><li>Check your configuration to be sure CloudTrail is writing <a href="https://docs.aws.amazon.com/awscloudtrail/latest/userguide/cloudtrail-working-with-log-files.html">logs</a> to the correct S3 bucket.</li><li>Confirm the log file was not deleted.</li></ul>                     |
| INVALID\_SQS\_MESSAGE                                                                                            | Device has invalid credentials. | The format of the SQS message is incorrect                                                              | Check your SQS dashboard and verify that it has properly formatted messages.                                                                                                                                                                                                                 |

[^1]: [Organization Settings > Security Devices](https://workbench.expel.io/settings/security-devices)


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.expel.io/connect-your-technology/security-devices/security-device-health.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
