Skip to content

Anwendungsmonitoring

CareOn Early Warning System

Verstehen, wie es der Anwendung geht.

CareOn EWS (Shomer-AI) überwacht Leistung und Auffälligkeiten der CareOn-Anwendung. Es unterstützt den technischen Betrieb dabei, Abweichungen und betroffene Arbeitsabläufe einzuordnen.

Medicine. Information. Technology.

Context

At a glance.

01

Überblick

Systemzustand und betroffene Anwendungsbereiche zusammenführen.

02

Einordnung

Auffälligkeiten mit historischen Vergleichswerten betrachten.

03

Betrieb

Vorfälle verfolgen und die zuständigen Teams informieren.

CareOn EWS beschreibt die Gesundheit einer Anwendung. Klinische Risikobewertung wird hier separat unter Soteria erklärt. Zahlen und Statusanzeigen im ursprünglichen Tutorial sind Beispiele aus dessen Quellenstand, keine Live-Anzeige.

How it works

Follow the workflow.

  1. Anwendungsdaten

    Leistungs- und Nutzungsdaten

  2. Analyse

    Abweichungen erkennen

  3. Vorfälle

    Auswirkungen einordnen

  4. Betriebsteam

    Untersuchen und reagieren

Conceptual diagram · not live data

Explore the details

The original article.

The original service-desk article is preserved below. The editorial context above explains the service; statements in the source reflect its publication date.

Source updated: · Mirrored:

Read original article Content & documentation

CareOn EWS Tutorial

Shomer-AI Early Warning System — Getting Started Guide

v3.0 March 2026 eMedys AG / DigitalON

PART 1

Getting Started

Everything you need to log in, navigate the dashboard, and understand system health at a glance.

What is CareOn EWS?

The Shomer-AI CareOn Early Warning System (EWS) is an automated monitoring platform that watches the CareOn clinical application around the clock. Shomer (שומר) is Hebrew for "Guardian" — and that's exactly what it does.

Every 15 minutes, the system extracts performance data from the PIMEDONE database (over 1.4 billion rows across 7 API tables), runs anomaly detection using 7+ parallel algorithms, tracks incidents across analysis runs, and sends alerts via email and WhatsApp when problems are detected.

ℹ️ Read-Only & Safe

The dashboard is completely read-only. It does not make any changes to the production Oracle iMedOne database. All data comes from the PIMEDONE PostgreSQL replica, and the system's own state is stored in a local SQLite database. There is zero impact on production database performance.

Logging In

Access the dashboard at careon-alerts.digitalon.co.za. Authentication uses a secure, passwordless email OTP flow — no passwords to remember.

  1. Open the URL in any modern browser (Chrome, Firefox, Safari, Edge). You'll see the login page with the Shomer-AI shield logo.
  2. Enter your corporate email address. Your email domain must be from an authorized organization: netcare.co.za, emedys.com, digitalontech.com, or first2lead.com.
  3. Check your inbox for a 6-digit verification code. The code arrives within seconds and is valid for 10 minutes.
  4. Enter the code on the verification screen. You have up to 5 attempts. After successful verification, a session cookie keeps you logged in for 30 days.

???? Pro Tip

After logging in, you'll land on the Main page — the executive overview. Bookmark this URL for quick access. Your session lasts 30 days, so you won't need to log in often.

The Main Page

The Main page (/management) is your executive command center. It's the default landing page and is designed to give non-technical stakeholders an instant picture of system health.

Hero Banner

At the top of the page, a large banner shows the current system state with an animated pulse indicator. When everything is healthy, you'll see a green pulse with "All Systems Operational." The banner also displays three key metrics:

97.7%
SLA Compliance
100%
Effective Uptime
93/100
Performance Index

SLA Compliance shows what percentage of APIs respond under 2 seconds. Effective Uptime tracks how often the system ran clean in the last 24 hours (amber/known-issue runs also count as uptime). The Performance Index (PI) is a 0–100 health score based on the aggregate P95 latency.

Top APIs by Volume

Below the hero banner, a strip shows the top 5 busiest APIs with their P95 response times, call counts, and unique user counts. This tells you at a glance which APIs are handling the most traffic and how fast they're responding.

System Health Chart

A 24-hour chart shows throughput, active users, clinical activities, and issue counts over time. Delta pills above the chart compare current values against the day-of-week baseline — for example, THROUGHPUT +17% means traffic is 17% higher than typical for this day and hour.

Clinical Subsystems Strip

Ten clinical subsystem monitors are listed with their current metrics, a HEALTHY or DEGRADED badge, percentage change versus baseline, and a mini sparkline trend chart. The subsystems include Pharmacy, DrugSearch, Micromedex D2D, DMS Documents, Bedside Monitors, qSOFA Scores, CareOn MFA, HL7 Providers, SAP Entities, and CareOn Replication.

User Experience Chart

A dual-axis chart at the bottom shows per-group activity, user impact, and average wait times — broken down by clinical role (Doctor, Nurse, HCO, Billing, and more).

The 5-Level Traffic Light System

The hero banner's pulse color reflects a 5-level traffic light that the system calculates after every analysis run. This is the single most important indicator on the dashboard.

State Meaning What To Do
Green No anomalies — all systems operational No action needed
Yellow Warnings only, below critical thresholds Monitor — no immediate action
Amber Critical anomalies, but ALL within known baseline tolerance (DOW hourly +/- 2σ or accepted baseline x2) No action — known slow APIs within envelope
Orange 1–2 APIs breaching baselines beyond tolerance Investigate the affected APIs; limited impact
Red 3+ APIs breaching baselines — widespread degradation Immediate investigation; escalate to DTCS

ℹ️ Executive View

On the Main page, the banner color is simplified for executives: green, yellow, and amber all show as green. Only orange and red will change the banner color. For the full 5-level traffic light, switch to the DevOps page.

Clinical Subsystems

The Subsystems page (/subsystems) gives a dedicated, detailed view of all 10 clinical subsystem monitors. Each subsystem has its own card with a sparkline chart, HEALTHY/DEGRADED badge, key metrics, and data point counts.

Subsystem What It Monitors
Pharmacy DispensePrescription dispensing: labels, signatures, removals per interval
DrugSearchMedication search operations: queries, errors per interval
Micromedex D2DDrug-drug interaction checks, error rates, P95 response
DMS DocumentsClinical document generation: docs, cases, units
Bedside MonitorsVital sign transmissions: readings, patients, beds
qSOFA ScoresSepsis early warning: scores computed, high-risk patients
CareOn MFAMulti-factor authentication: binds, users, failures
HL7 ProvidersHL7 message transmission to Soteria: providers, incidents
SAP EntitiesSAP document transfers: types, volumes, latency
DB ReplicationOracle to PIMEDONE replication lag and connection health

Each subsystem shows an ONLINE or STALE freshness badge. Data is considered stale if the metrics are older than 15 minutes — which would indicate the analysis pipeline may have missed a cycle.

Navigation & Auto-Refresh

The top navigation bar provides one-click access to all 18 dashboard pages. The currently active page is highlighted. On the right side of the nav bar, you'll find:

  1. Who's Online — a dropdown showing all active authenticated sessions, with email addresses, country flags (from Cloudflare GeoIP), and BYPASS badges for monitoring servers.
  2. Auto-Refresh control — choose from 30s, 60s (default), 2m, 5m, or Off. A 3px progress bar at the bottom shows the countdown: blue → yellow (under 2 min remaining) → red pulsing (overdue).

???? Pro Tip

The auto-refresh setting is saved in your browser's localStorage, so it persists between sessions. Set it to 60s for routine monitoring, or Off when investigating a specific incident to prevent the page from refreshing while you're reading.

PART 2

Daily Operations

For support staff and IT operations: investigate issues, run diagnostics, and manage incidents.

The DevOps Dashboard

The DevOps page (/dashboard) is the operations nerve center. Unlike the executive Main page, it shows the full 5-level traffic light, embedded Grafana panels, pipeline operations, Oracle error panels, and detailed incident/alert sections.

Key Sections

Hero Banner (Full Traffic Light) — Shows the true 5-level state including amber. Displays Effective Uptime and Performance Index on the right, with a live countdown timer showing "Next analysis in M:SS".

Summary Cards — 7–9 KPI cards showing anomalies, API calls, active users, P95 latency, and more — all with 24-hour delta comparisons.

Embedded Grafana Panels — Three live Grafana panels update every minute: End-to-End Response Times (3h), Oracle RAC CPU Load (8h), and App Worker Connections (8h). These provide server-level infrastructure context that complements the application-level metrics from Shomer-AI.

Pipeline Operations Log — Shows recent analysis runs with their start time, duration, and outcome. Useful for confirming the pipeline is running on schedule.

Oracle DB Errors Panel — Summarizes recent ORA error codes correlated with API anomalies.

Active Incidents & Top Degraded APIs — Quick reference tables showing current incidents and the APIs experiencing the worst performance.

Using "Check Now" Diagnostics

When a clinician calls IT saying "CareOn is slow", the Check Now button gives your support team an answer in under 10 seconds — no manual SQL required.

  1. Click the "Check Now" button on the DevOps dashboard.
  2. Optionally enter a specific API name (e.g., Servicespending) and/or a user identifier to narrow the investigation.
  3. The system queries the last 15 minutes of live performance data in real-time.
  4. Within less than 10 seconds, you receive: a traffic light indicator (green/yellow/red) for the current state, a comparison table showing current P95 vs. hourly baseline, copy-paste diagnostic SQL queries for immediate PIMEDONE investigation, and if a user was specified, their recent call history with response times.

???? Pro Tip

Check Now results are persisted in the diagnostic_checks table and visible in the Check History section. Use this to compare multiple checks over time when investigating an evolving issue.

Reading Incidents

The Incidents page (/incidents) shows all tracked incidents — both active and resolved. Each anomaly becomes a tracked incident with a unique ID that persists across analysis runs.

Incident Lifecycle

Incidents flow through a clear lifecycle: OPENONGOINGRESOLVED. If the same API re-opens within 4 hours of resolution, the system marks it as flapping and auto-escalates after 3 flaps.

Cascade Groups

When 3+ active incidents share the same server or ORA error code, the system groups them into a cascade with a synthesized root cause — for example, "Database Issue: ORA-12170 affecting 10 APIs." Look for the cascade banner at the top of the Incidents page.

Incident Detail

Click any incident ID to open its detail page. Here you'll find: Root cause hypothesis — the system's best guess at what's causing the issue. Affected user count and user groups impacted. Cascade group membership linking related APIs. Pre-built diagnostic SQL queries you can paste directly into PIMEDONE. Deviation factor chart showing how the anomaly evolved over time. Operator feedback buttons — mark as False Positive, Confirmed, or Expected Maintenance.

⚠️ Operator Feedback Matters

When you mark an incident as False Positive, the system learns and loosens thresholds for that API. When you mark it as Confirmed (True Positive), thresholds are tightened. This feedback loop directly improves detection accuracy over time.

Understanding Baselines

The Baselines page (/baselines) visualizes the system's learned performance profiles as a 24-hour heatmap. Each row is an API, each column is an hour (0–23), and the cell color represents the expected P95 response time.

The 4-Level Baseline System

Level 1: Day-of-Week Hourly (most specific) — Wednesday at 10 AM has a different expected profile than Sunday at 10 AM. Requires at least 2 samples. Currently tracking 22,000+ DOW-hourly baseline records.

Level 2: Hourly Rolling — Per-hour baselines aggregated across all days. Captures consistent time-of-day patterns like the 07:00–09:00 ward round spike.

Level 3: Accepted Baselines — When an operator confirms an API is operating as expected, that baseline is "accepted" with a 1.25× tolerance factor. This prevents alert fatigue on inherently slow endpoints.

Level 4: 7-Day Rolling Median — A robust statistical fallback using IQR outlier rejection to prevent single bad days from inflating baselines.

ℹ️ Adaptive Thresholds

The system automatically tunes detection sensitivity per API based on the coefficient of variation (CV). Low-variance APIs get tight thresholds (2× spike triggers an alert). High-variance APIs get looser thresholds (up to 5×). Currently 208+ adaptive adjustments are active.

Problems & Clinical Impact

The Problems page (/problems) translates technical anomalies into clinical impact metrics that management and clinical stakeholders can understand.

Key sections include: Clinical Impact Strip — Total collective wait time, average wait per user, and the worst-performing API. Time Lost by Department — Which clinical groups (Doctors, Nurses, Pharmacists) were most affected. Action Items — Prioritized recommendations with ESCALATE/INVESTIGATE badges, so your team knows exactly what to do first.

Alert Channels

CareOn EWS delivers alerts through two channels simultaneously:

Email Alerts

Critical Alerts fire immediately when critical anomalies are detected (60-minute per-incident cooldown). They include the top anomalies, root cause hypothesis, ORA correlations, diagnostic SQL, and affected users. Both internal and client-filtered HTML reports are attached.

Daily Digest is sent at 07:00 SAST with a 24-hour summary: health trajectory, top degraded APIs, most affected users, incident summary, and subsystem health.

WhatsApp / Matrix Alerts

Seven message types are delivered to the support group chat: Critical Alert — anomalies with recommended actions. Recovery — previous critical anomalies cleared. All-Clear — all incidents resolved. Heartbeat — periodic proof-of-life every 4 hours (suppressed if critical sent recently). Shift Summary — end-of-shift carry-forward. Daily Digest — 24-hour metrics. Dashboard Screenshot — visual snapshot attached with critical/digest alerts.

???? Join the WhatsApp Group

Join the alert group at chat.whatsapp.com/B9IHGoFSICQDVflG3WKhTT to receive real-time critical alerts and recovery notifications on your phone.

PART 3

Advanced Features

For power users: deep performance analytics, historical replay, and system configuration.

API Stats & Grades

The API Stats page (/api-stats) shows a performance matrix for all ~300 monitored endpoints with P50, P90, P95, P99 response times, call counts, unique users, and trend indicators. Click any row to expand a 7-day P95 trend sparkline chart. Data is exportable as CSV.

The API Grades page (/api-grades) assigns letter grades (A–F) to each API based on P95 performance versus baseline, call volume, and false positive rate — giving you a quick at-a-glance health summary.

User Performance

The Users page (/user-performance) shows the top 20 most-impacted users per clinical role group (11 groups), ranked by cumulative wait time over a 7-day window. Each entry shows the user's rank, obfuscated ID, total API calls, P95 latency, total wait time, and a freshness badge.

Use the search box to look up any individual user by their ID code — you'll see their personal experience over the last 4 hours. This is how you answer the question: "Was Doctor Schmidt actually affected?"

ORA Errors & Exceptions

The ORA Errors page (/oracle-errors) provides dedicated Oracle database error analysis: summary cards (total errors, unique codes, affected hosts), an error table with ORA code, description, count, and severity, a trend chart showing errors over time, and most importantly — incident correlation that automatically links ORA errors to specific API anomalies.

The Exceptions page (/exceptions) covers non-ORA application exceptions, categorized into three types: API Exception, SQL Exception, and Application Exception. Each has its own trend chart and incident correlation.

Historical Replay

The Replay page (/replay) shows results from historical replay runs — where the system re-runs its entire detection pipeline against past production data, hour by hour.

Replay is used to validate detection accuracy against known outages (like the February 14 incident), test threshold changes retroactively, and calculate false positive rates. The page shows step counts, outcome distributions (True Positive / False Positive / Unknown), and detection quality metrics.

ℹ️ Building Trust

Historical replay is how the system proves its value before going live. By re-running against known incidents like the February 14 crisis (618-second response times, 73× normal), the team can verify the system would have detected the problem within 15 minutes — not hours.

SysOps Controls

The SysOps page (/sysops) is the system control panel. From here, operators can:

  • Toggle alert channels — Enable/disable email and Matrix/WhatsApp alerts independently
  • Switch Eval Mode — Toggle between LIVE (alerts are delivered) and EVAL mode (alerts generated and visible on dashboard but delivery suppressed, marked with [EVAL] prefix)
  • View pipeline health — Confirm the analysis pipeline is running on schedule
  • Accept/Clear baselines — Accept current baselines for specific APIs to prevent alert fatigue, or clear accepted baselines to reset
  • Trigger maintenance — Run database cleanup and optimization tasks

⚠️ Eval Mode

When testing new thresholds or detection strategies, use Eval Mode. Alerts will still appear on the dashboard with an [EVAL] prefix, but no emails or WhatsApp messages will be sent. This lets you validate changes without disturbing the team.

Quick Reference

7 Detection Algorithms

Algorithm What It Detects Warning Critical
SpikeSudden P95 increase exceeding baseline>3× baseline>5× baseline
SustainedElevated P95 persisting for 15+ minutes>2× for 15 min>3× for 15 min
DriftGradual P95 increase over the 7-day windowTrending upConfirmed regression
Server HotspotSingle app server's P95 far exceeds peers>2× peers>3× peers
Throughput DropAPI call volume dropped vs. DOW baseline>60% drop>90% drop
Tail LatencyP99 significantly exceeds P95 (long-tail outliers)P99/P95 >5×P99/P95 >7.5×
High VarianceResponse time StdDev spiked vs. baselineStdDev >3×StdDev >6×

All 18 Dashboard Pages

Page Purpose
MainExecutive KPI overview, clinical subsystems, user experience
Subsystems10 clinical subsystem cards with sparklines and health badges
DevOpsOperations center: traffic light, Grafana, Check Now, pipeline log
ProblemsClinical impact translation, time lost by department, action items
IncidentsIncident lifecycle, cascade groups, operator feedback
AlertsEmail and WhatsApp alert history with delivery status
QualityDetection quality: FP/TP rates, operator feedback stats
Baselines24-hour P95 heatmap with DOW-aware baselines
ReportsInternal + client HTML report archive back to Feb 14
API StatsPerformance matrix for ~300 endpoints with CSV export
UsersTop 20 impacted users per role group, user lookup
WorkflowsPer-run user impact, activity breakdown, trend charts
ORA ErrorsOracle error analysis with incident correlation
ExceptionsAPI/SQL/App exception categories with trend charts
ReplayHistorical replay validation and detection quality metrics
API GradesA–F letter grades for all ~300 endpoints
SysOpsAlert toggles, eval mode, pipeline health, maintenance
GuideBuilt-in usage guide and troubleshooting reference

Key URLs

Escalation Path

  1. Dashboard — Check CareOn Alert Dashboard for current status and anomaly details
  2. Grafana — Drill into the Athena Platform for server-level metrics and API logs
  3. IM4HC — Verify CareOn service health at IM4HC Health Check
  4. ELK Stack — Check application logs at Elasticsearch/Kibana (elkfra)
  5. CheckMK — Server hardware metrics (CPU, memory, disk)
  6. DBA — For Oracle errors: check AWR, table statistics, index health
  7. Escalate — Contact [email protected] for critical issues requiring vendor intervention

Shomer-AI CareOn Early Warning System — Copyright © 2026 eMedys AG and OPEN MEDYS GMBH. All Rights Reserved.

Contact: Fabian Berger (CEO/Vorstand) — [email protected]

This tutorial was generated for the CareOn EWS v3 dashboard. For the latest documentation, see the built-in Guide page.

← All companies & solutions
eMedys AG

Company and product information
Source mirror dated · Content responsibility: eMedys AG

eMedys overview ↗