40

2026 · Article

Hierarchical Anomaly Detection and SHAP-Based Root-Cause Attribution for Robotic Process Automation Workflows

Yanka Aleksandrova · Mihail Radev · Mila Georgieva · Desislava Koleva

Original methodological and empirical research study with proof-of-concept validation

How to cite this article

Aleksandrova, Y., Radev, M., Georgieva, M., & Koleva, D. (2026). Hierarchical Anomaly Detection and SHAP-Based Root-Cause Attribution for Robotic Process Automation Workflows. Proceedings of the International Conference on Business Excellence, 20(1), 538–552. https://doi.org/10.2478/picbe-2026-0043

Extended Summary

Introduction

Robotic Process Automation (RPA) is increasingly used to automate repetitive, rule-based, and structured business processes. The execution of RPA workflows generates detailed logs and telemetry data that can be used for process monitoring, anomaly detection, operational diagnostics, and performance improvement.

However, many conventional anomaly-detection approaches analyse workflow behaviour at a single level, most commonly at the level of the complete process run. Such aggregation may obscure the distinction between problems originating in the execution environment—such as infrastructure load, queue delays, connector-related issues, or concurrency effects—and abnormalities arising from the internal workflow logic, individual process actions, or their temporal sequence.

This distinction is particularly important in complex automated environments because identifying that an anomaly has occurred is not sufficient for effective operational management. Diagnosis also requires information about the likely source and nature of the abnormal behaviour. The study therefore addresses the need for a hierarchical approach capable of integrating information from different levels of RPA execution and providing interpretable evidence for root-cause attribution.

Aim

The study aims to develop and evaluate a hierarchical two-tier framework for anomaly detection and root-cause attribution in Robotic Process Automation workflows.

The proposed framework integrates run-level and process-level anomaly detection in order to distinguish between abnormalities associated primarily with the execution environment and those related to internal workflow behaviour.

The central research question is whether combining these two analytical levels can provide more informative and operationally useful diagnostic information than conventional single-level anomaly detection.

Materials and Methods

The proposed framework consists of two interconnected analytical tiers.

Tier 1: Run-level anomaly detection

The first tier analyses environmental and aggregate characteristics of each RPA execution. The examined variables include queue time, payload size, concurrency, connector complexity, number of workflow steps, total execution duration, and recent failure rates.

Two unsupervised machine-learning approaches are applied:

  • Isolation Forest, designed to identify observations that can be isolated from the remaining data through relatively few random partitions;
  • a feed-forward Autoencoder, which detects abnormal executions through reconstruction error.

Tier 2: Process-level anomaly detection

The second tier examines internal workflow behaviour at the level of individual actions and their temporal sequence.

Two complementary methods are used:

  • Z-score analysis to identify deviations in the duration of individual workflow actions;
  • Long Short-Term Memory Autoencoder (LSTM-AE) to detect temporal and sequential abnormalities in process execution.

The outputs from the two analytical tiers are integrated into four diagnostic states:

  • Normal — no substantial anomaly is identified;
  • Environment Stress — abnormal behaviour is predominantly detected at the run or environmental level;
  • Latent Logic Issue — process-level abnormalities are present without a corresponding substantial run-level anomaly;
  • Systemic Failure — abnormalities are detected simultaneously at both analytical levels.

To improve the interpretability of detected anomalies, the framework incorporates SHAP (SHapley Additive exPlanations). SHAP values are used to estimate the contribution of individual variables to the anomaly assessment.

The authors further introduce an External Influence Index (EII) based on SHAP values. The index is designed to estimate the relative contribution of external environmental factors versus internal workflow characteristics to an identified anomaly.

The framework is evaluated in two stages. First, validation is performed using a controlled synthetic dataset comprising 10,000 RPA runs, with predefined anomalies introduced into the data and known ground-truth labels available for evaluation. Second, the practical feasibility of the framework is examined using 325 real executions of an automated university administrative process implemented in Microsoft Power Automate. The process concerns the submission and evaluation of scientific project proposals.

Results

The numerical results presented in this section originate from the authors’ own empirical analyses and are not values extracted from external studies.

In the synthetic dataset, approximately 21.6% of the 10,000 RPA runs contained at least one injected anomaly.

At the run level, both unsupervised models demonstrated strong discrimination between anomalous and normal executions. Isolation Forest achieved an ROC-AUC of 0.9571, while the feed-forward Autoencoder achieved an ROC-AUC of 0.9025 for detection of injected anomalies.

The results also showed that run-level information alone was less effective in identifying certain operational outcomes, including SLA overruns and general execution failure. This supports the study’s central premise that aggregate execution characteristics cannot fully capture abnormalities occurring within the internal process structure.

At the process level, the LSTM Autoencoder achieved an ROC-AUC of 0.9643, while the Z-score approach achieved an ROC-AUC of 0.9472 for injected anomaly detection.

For SLA overruns, the Z-score method reached an ROC-AUC of 0.7834, compared with 0.7114 for the LSTM-based approach.

The relationship between the two process-level anomaly scores was moderate but statistically significant (r = 0.5453; p < 0.001). This indicates that the methods capture related but not identical dimensions of abnormal workflow behaviour and therefore provide complementary diagnostic information.

After integration of the two analytical tiers, the 10,000 synthetic runs were classified as follows:

  • 78.4% Normal;
  • 16.6% Latent Logic Issue;
  • 3.6% Systemic Failure;
  • 1.5% Environment Stress.

The hierarchical classification therefore extends conventional anomaly detection by providing information not only on whether abnormal behaviour is present, but also on its probable location within the automated system.

The pilot application involving 325 real RPA executions produced a broadly comparable diagnostic pattern. At Tier 1, the anomaly scores generated by the two models were positively correlated (r = 0.68; p < 0.001). At Tier 2, the Z-score and LSTM-based anomaly indices also demonstrated a statistically significant relationship (r = 0.5782; p < 0.001).

Approximately 14% of the real workflow executions contained at least one strong process-level deviation.

Following integration of the two analytical levels, the real-world executions were classified as:

  • 73.3% Normal;
  • 13.6% Environment Stress;
  • 8.0% Latent Logic Issue;
  • 5.1% Systemic Failure.

These findings indicate that combining run-level and process-level information provides a more differentiated representation of abnormal workflow behaviour than analysing either level in isolation.

Conclusion

The study demonstrates the potential value of hierarchical anomaly detection for monitoring and diagnosing complex Robotic Process Automation workflows.

By combining environmental and aggregate execution characteristics with information about individual process actions and temporal sequences, the proposed framework enables differentiation between infrastructure-related stress, latent internal workflow problems, and failures affecting both levels simultaneously.

The integration of SHAP-based explainability and the External Influence Index further strengthens the diagnostic value of the framework by transforming anomaly scores into interpretable information about the factors contributing to abnormal behaviour.

This approach may support more effective process monitoring, problem prioritisation, remediation, and governance of RPA environments by helping decision-makers move from simple anomaly detection toward more informed root-cause attribution.

The principal limitation of the study is that the main validation dataset is synthetic. Although the pilot analysis of 325 real workflow executions provides initial evidence of practical applicability, further validation is required across larger, heterogeneous, and production-scale RPA environments.

Scientific Keywords

Robotic Process AutomationRPA workflowshierarchical anomaly detectionanomaly detectionroot-cause attributionroot cause analysisexplainable artificial intelligenceexplainable AISHAPIsolation ForestAutoencoderLong Short-Term Memory AutoencoderLSTM-AEprocess monitoringworkflow diagnosticsoperational diagnosticsExternal Influence Indexautomated business processes