Anwendungsmonitoring
CareOn Early Warning System
Verstehen, wie es der Anwendung geht.
CareOn EWS (Shomer-AI) überwacht Leistung und Auffälligkeiten der CareOn-Anwendung. Es unterstützt den technischen Betrieb dabei, Abweichungen und betroffene Arbeitsabläufe einzuordnen.

Context
At a glance.
Überblick
Systemzustand und betroffene Anwendungsbereiche zusammenführen.
Einordnung
Auffälligkeiten mit historischen Vergleichswerten betrachten.
Betrieb
Vorfälle verfolgen und die zuständigen Teams informieren.
CareOn EWS beschreibt die Gesundheit einer Anwendung. Klinische Risikobewertung wird hier separat unter Soteria erklärt. Zahlen und Statusanzeigen im ursprünglichen Tutorial sind Beispiele aus dessen Quellenstand, keine Live-Anzeige.
How it works
Follow the workflow.
Anwendungsdaten
Leistungs- und Nutzungsdaten
Analyse
Abweichungen erkennen
Vorfälle
Auswirkungen einordnen
Betriebsteam
Untersuchen und reagieren
Conceptual diagram · not live data
Explore the details
The original article.
The original service-desk article is preserved below. The editorial context above explains the service; statements in the source reflect its publication date.
Read original article Content & documentation
CareOn EWS Tutorial
Shomer-AI Early Warning System — Getting Started Guide
v3.0 March 2026 eMedys AG / DigitalON
Table of Contents
Part 1 — Getting Started
Part 2 — Daily Operations
Part 3 — Advanced Features
PART 1
Getting Started
Everything you need to log in, navigate the dashboard, and understand system health at a glance.
What is CareOn EWS?
The Shomer-AI CareOn Early Warning System (EWS) is an automated monitoring platform that watches the CareOn clinical application around the clock. Shomer (שומר) is Hebrew for "Guardian" — and that's exactly what it does.
Every 15 minutes, the system extracts performance data from the PIMEDONE database (over 1.4 billion rows across 7 API tables), runs anomaly detection using 7+ parallel algorithms, tracks incidents across analysis runs, and sends alerts via email and WhatsApp when problems are detected.
ℹ️ Read-Only & Safe
The dashboard is completely read-only. It does not make any changes to the production Oracle iMedOne database. All data comes from the PIMEDONE PostgreSQL replica, and the system's own state is stored in a local SQLite database. There is zero impact on production database performance.
Logging In
Access the dashboard at careon-alerts.digitalon.co.za. Authentication uses a secure, passwordless email OTP flow — no passwords to remember.
- Open the URL in any modern browser (Chrome, Firefox, Safari, Edge). You'll see the login page with the Shomer-AI shield logo.
- Enter your corporate email address. Your email domain must be from an authorized organization:
netcare.co.za,emedys.com,digitalontech.com, orfirst2lead.com. - Check your inbox for a 6-digit verification code. The code arrives within seconds and is valid for 10 minutes.
- Enter the code on the verification screen. You have up to 5 attempts. After successful verification, a session cookie keeps you logged in for 30 days.
???? Pro Tip
After logging in, you'll land on the Main page — the executive overview. Bookmark this URL for quick access. Your session lasts 30 days, so you won't need to log in often.
The Main Page
The Main page (/management) is your executive command center. It's the default landing page and is designed to give non-technical stakeholders an instant picture of system health.
Hero Banner
At the top of the page, a large banner shows the current system state with an animated pulse indicator. When everything is healthy, you'll see a green pulse with "All Systems Operational." The banner also displays three key metrics:
| 97.7% SLA Compliance | 100% Effective Uptime | 93/100 Performance Index |
SLA Compliance shows what percentage of APIs respond under 2 seconds. Effective Uptime tracks how often the system ran clean in the last 24 hours (amber/known-issue runs also count as uptime). The Performance Index (PI) is a 0–100 health score based on the aggregate P95 latency.
Top APIs by Volume
Below the hero banner, a strip shows the top 5 busiest APIs with their P95 response times, call counts, and unique user counts. This tells you at a glance which APIs are handling the most traffic and how fast they're responding.
System Health Chart
A 24-hour chart shows throughput, active users, clinical activities, and issue counts over time. Delta pills above the chart compare current values against the day-of-week baseline — for example, THROUGHPUT +17% means traffic is 17% higher than typical for this day and hour.
Clinical Subsystems Strip
Ten clinical subsystem monitors are listed with their current metrics, a HEALTHY or DEGRADED badge, percentage change versus baseline, and a mini sparkline trend chart. The subsystems include Pharmacy, DrugSearch, Micromedex D2D, DMS Documents, Bedside Monitors, qSOFA Scores, CareOn MFA, HL7 Providers, SAP Entities, and CareOn Replication.
User Experience Chart
A dual-axis chart at the bottom shows per-group activity, user impact, and average wait times — broken down by clinical role (Doctor, Nurse, HCO, Billing, and more).
The 5-Level Traffic Light System
The hero banner's pulse color reflects a 5-level traffic light that the system calculates after every analysis run. This is the single most important indicator on the dashboard.
| State | Meaning | What To Do |
|---|---|---|
| Green | No anomalies — all systems operational | No action needed |
| Yellow | Warnings only, below critical thresholds | Monitor — no immediate action |
| Amber | Critical anomalies, but ALL within known baseline tolerance (DOW hourly +/- 2σ or accepted baseline x2) | No action — known slow APIs within envelope |
| Orange | 1–2 APIs breaching baselines beyond tolerance | Investigate the affected APIs; limited impact |
| Red | 3+ APIs breaching baselines — widespread degradation | Immediate investigation; escalate to DTCS |
ℹ️ Executive View
On the Main page, the banner color is simplified for executives: green, yellow, and amber all show as green. Only orange and red will change the banner color. For the full 5-level traffic light, switch to the DevOps page.
Clinical Subsystems
The Subsystems page (/subsystems) gives a dedicated, detailed view of all 10 clinical subsystem monitors. Each subsystem has its own card with a sparkline chart, HEALTHY/DEGRADED badge, key metrics, and data point counts.
| Subsystem | What It Monitors |
|---|---|
| Pharmacy Dispense | Prescription dispensing: labels, signatures, removals per interval |
| DrugSearch | Medication search operations: queries, errors per interval |
| Micromedex D2D | Drug-drug interaction checks, error rates, P95 response |
| DMS Documents | Clinical document generation: docs, cases, units |
| Bedside Monitors | Vital sign transmissions: readings, patients, beds |
| qSOFA Scores | Sepsis early warning: scores computed, high-risk patients |
| CareOn MFA | Multi-factor authentication: binds, users, failures |
| HL7 Providers | HL7 message transmission to Soteria: providers, incidents |
| SAP Entities | SAP document transfers: types, volumes, latency |
| DB Replication | Oracle to PIMEDONE replication lag and connection health |
Each subsystem shows an ONLINE or STALE freshness badge. Data is considered stale if the metrics are older than 15 minutes — which would indicate the analysis pipeline may have missed a cycle.
Navigation & Auto-Refresh
The top navigation bar provides one-click access to all 18 dashboard pages. The currently active page is highlighted. On the right side of the nav bar, you'll find:
- Who's Online — a dropdown showing all active authenticated sessions, with email addresses, country flags (from Cloudflare GeoIP), and BYPASS badges for monitoring servers.
- Auto-Refresh control — choose from 30s, 60s (default), 2m, 5m, or Off. A 3px progress bar at the bottom shows the countdown: blue → yellow (under 2 min remaining) → red pulsing (overdue).
???? Pro Tip
The auto-refresh setting is saved in your browser's localStorage, so it persists between sessions. Set it to 60s for routine monitoring, or Off when investigating a specific incident to prevent the page from refreshing while you're reading.
PART 2
Daily Operations
For support staff and IT operations: investigate issues, run diagnostics, and manage incidents.
The DevOps Dashboard
The DevOps page (/dashboard) is the operations nerve center. Unlike the executive Main page, it shows the full 5-level traffic light, embedded Grafana panels, pipeline operations, Oracle error panels, and detailed incident/alert sections.
Key Sections
Hero Banner (Full Traffic Light) — Shows the true 5-level state including amber. Displays Effective Uptime and Performance Index on the right, with a live countdown timer showing "Next analysis in M:SS".
Summary Cards — 7–9 KPI cards showing anomalies, API calls, active users, P95 latency, and more — all with 24-hour delta comparisons.
Embedded Grafana Panels — Three live Grafana panels update every minute: End-to-End Response Times (3h), Oracle RAC CPU Load (8h), and App Worker Connections (8h). These provide server-level infrastructure context that complements the application-level metrics from Shomer-AI.
Pipeline Operations Log — Shows recent analysis runs with their start time, duration, and outcome. Useful for confirming the pipeline is running on schedule.
Oracle DB Errors Panel — Summarizes recent ORA error codes correlated with API anomalies.
Active Incidents & Top Degraded APIs — Quick reference tables showing current incidents and the APIs experiencing the worst performance.
Using "Check Now" Diagnostics
When a clinician calls IT saying "CareOn is slow", the Check Now button gives your support team an answer in under 10 seconds — no manual SQL required.
- Click the "Check Now" button on the DevOps dashboard.
- Optionally enter a specific API name (e.g.,
Servicespending) and/or a user identifier to narrow the investigation. - The system queries the last 15 minutes of live performance data in real-time.
- Within less than 10 seconds, you receive: a traffic light indicator (green/yellow/red) for the current state, a comparison table showing current P95 vs. hourly baseline, copy-paste diagnostic SQL queries for immediate PIMEDONE investigation, and if a user was specified, their recent call history with response times.
???? Pro Tip
Check Now results are persisted in the diagnostic_checks table and visible in the Check History section. Use this to compare multiple checks over time when investigating an evolving issue.
Reading Incidents
The Incidents page (/incidents) shows all tracked incidents — both active and resolved. Each anomaly becomes a tracked incident with a unique ID that persists across analysis runs.
Incident Lifecycle
Incidents flow through a clear lifecycle: OPEN → ONGOING → RESOLVED. If the same API re-opens within 4 hours of resolution, the system marks it as flapping and auto-escalates after 3 flaps.
Cascade Groups
When 3+ active incidents share the same server or ORA error code, the system groups them into a cascade with a synthesized root cause — for example, "Database Issue: ORA-12170 affecting 10 APIs." Look for the cascade banner at the top of the Incidents page.
Incident Detail
Click any incident ID to open its detail page. Here you'll find: Root cause hypothesis — the system's best guess at what's causing the issue. Affected user count and user groups impacted. Cascade group membership linking related APIs. Pre-built diagnostic SQL queries you can paste directly into PIMEDONE. Deviation factor chart showing how the anomaly evolved over time. Operator feedback buttons — mark as False Positive, Confirmed, or Expected Maintenance.
⚠️ Operator Feedback Matters
When you mark an incident as False Positive, the system learns and loosens thresholds for that API. When you mark it as Confirmed (True Positive), thresholds are tightened. This feedback loop directly improves detection accuracy over time.
Understanding Baselines
The Baselines page (/baselines) visualizes the system's learned performance profiles as a 24-hour heatmap. Each row is an API, each column is an hour (0–23), and the cell color represents the expected P95 response time.
The 4-Level Baseline System
Level 1: Day-of-Week Hourly (most specific) — Wednesday at 10 AM has a different expected profile than Sunday at 10 AM. Requires at least 2 samples. Currently tracking 22,000+ DOW-hourly baseline records.
Level 2: Hourly Rolling — Per-hour baselines aggregated across all days. Captures consistent time-of-day patterns like the 07:00–09:00 ward round spike.
Level 3: Accepted Baselines — When an operator confirms an API is operating as expected, that baseline is "accepted" with a 1.25× tolerance factor. This prevents alert fatigue on inherently slow endpoints.
Level 4: 7-Day Rolling Median — A robust statistical fallback using IQR outlier rejection to prevent single bad days from inflating baselines.
ℹ️ Adaptive Thresholds
The system automatically tunes detection sensitivity per API based on the coefficient of variation (CV). Low-variance APIs get tight thresholds (2× spike triggers an alert). High-variance APIs get looser thresholds (up to 5×). Currently 208+ adaptive adjustments are active.
Problems & Clinical Impact
The Problems page (/problems) translates technical anomalies into clinical impact metrics that management and clinical stakeholders can understand.
Key sections include: Clinical Impact Strip — Total collective wait time, average wait per user, and the worst-performing API. Time Lost by Department — Which clinical groups (Doctors, Nurses, Pharmacists) were most affected. Action Items — Prioritized recommendations with ESCALATE/INVESTIGATE badges, so your team knows exactly what to do first.
Alert Channels
CareOn EWS delivers alerts through two channels simultaneously:
Email Alerts
Critical Alerts fire immediately when critical anomalies are detected (60-minute per-incident cooldown). They include the top anomalies, root cause hypothesis, ORA correlations, diagnostic SQL, and affected users. Both internal and client-filtered HTML reports are attached.
Daily Digest is sent at 07:00 SAST with a 24-hour summary: health trajectory, top degraded APIs, most affected users, incident summary, and subsystem health.
WhatsApp / Matrix Alerts
Seven message types are delivered to the support group chat: Critical Alert — anomalies with recommended actions. Recovery — previous critical anomalies cleared. All-Clear — all incidents resolved. Heartbeat — periodic proof-of-life every 4 hours (suppressed if critical sent recently). Shift Summary — end-of-shift carry-forward. Daily Digest — 24-hour metrics. Dashboard Screenshot — visual snapshot attached with critical/digest alerts.
???? Join the WhatsApp Group
Join the alert group at chat.whatsapp.com/B9IHGoFSICQDVflG3WKhTT to receive real-time critical alerts and recovery notifications on your phone.
PART 3
Advanced Features
For power users: deep performance analytics, historical replay, and system configuration.
API Stats & Grades
The API Stats page (/api-stats) shows a performance matrix for all ~300 monitored endpoints with P50, P90, P95, P99 response times, call counts, unique users, and trend indicators. Click any row to expand a 7-day P95 trend sparkline chart. Data is exportable as CSV.
The API Grades page (/api-grades) assigns letter grades (A–F) to each API based on P95 performance versus baseline, call volume, and false positive rate — giving you a quick at-a-glance health summary.
User Performance
The Users page (/user-performance) shows the top 20 most-impacted users per clinical role group (11 groups), ranked by cumulative wait time over a 7-day window. Each entry shows the user's rank, obfuscated ID, total API calls, P95 latency, total wait time, and a freshness badge.
Use the search box to look up any individual user by their ID code — you'll see their personal experience over the last 4 hours. This is how you answer the question: "Was Doctor Schmidt actually affected?"
ORA Errors & Exceptions
The ORA Errors page (/oracle-errors) provides dedicated Oracle database error analysis: summary cards (total errors, unique codes, affected hosts), an error table with ORA code, description, count, and severity, a trend chart showing errors over time, and most importantly — incident correlation that automatically links ORA errors to specific API anomalies.
The Exceptions page (/exceptions) covers non-ORA application exceptions, categorized into three types: API Exception, SQL Exception, and Application Exception. Each has its own trend chart and incident correlation.
Historical Replay
The Replay page (/replay) shows results from historical replay runs — where the system re-runs its entire detection pipeline against past production data, hour by hour.
Replay is used to validate detection accuracy against known outages (like the February 14 incident), test threshold changes retroactively, and calculate false positive rates. The page shows step counts, outcome distributions (True Positive / False Positive / Unknown), and detection quality metrics.
ℹ️ Building Trust
Historical replay is how the system proves its value before going live. By re-running against known incidents like the February 14 crisis (618-second response times, 73× normal), the team can verify the system would have detected the problem within 15 minutes — not hours.
SysOps Controls
The SysOps page (/sysops) is the system control panel. From here, operators can:
- Toggle alert channels — Enable/disable email and Matrix/WhatsApp alerts independently
- Switch Eval Mode — Toggle between LIVE (alerts are delivered) and EVAL mode (alerts generated and visible on dashboard but delivery suppressed, marked with [EVAL] prefix)
- View pipeline health — Confirm the analysis pipeline is running on schedule
- Accept/Clear baselines — Accept current baselines for specific APIs to prevent alert fatigue, or clear accepted baselines to reset
- Trigger maintenance — Run database cleanup and optimization tasks
⚠️ Eval Mode
When testing new thresholds or detection strategies, use Eval Mode. Alerts will still appear on the dashboard with an [EVAL] prefix, but no emails or WhatsApp messages will be sent. This lets you validate changes without disturbing the team.
Quick Reference
7 Detection Algorithms
| Algorithm | What It Detects | Warning | Critical |
|---|---|---|---|
| Spike | Sudden P95 increase exceeding baseline | >3× baseline | >5× baseline |
| Sustained | Elevated P95 persisting for 15+ minutes | >2× for 15 min | >3× for 15 min |
| Drift | Gradual P95 increase over the 7-day window | Trending up | Confirmed regression |
| Server Hotspot | Single app server's P95 far exceeds peers | >2× peers | >3× peers |
| Throughput Drop | API call volume dropped vs. DOW baseline | >60% drop | >90% drop |
| Tail Latency | P99 significantly exceeds P95 (long-tail outliers) | P99/P95 >5× | P99/P95 >7.5× |
| High Variance | Response time StdDev spiked vs. baseline | StdDev >3× | StdDev >6× |
All 18 Dashboard Pages
| Page | Purpose |
|---|---|
| Main | Executive KPI overview, clinical subsystems, user experience |
| Subsystems | 10 clinical subsystem cards with sparklines and health badges |
| DevOps | Operations center: traffic light, Grafana, Check Now, pipeline log |
| Problems | Clinical impact translation, time lost by department, action items |
| Incidents | Incident lifecycle, cascade groups, operator feedback |
| Alerts | Email and WhatsApp alert history with delivery status |
| Quality | Detection quality: FP/TP rates, operator feedback stats |
| Baselines | 24-hour P95 heatmap with DOW-aware baselines |
| Reports | Internal + client HTML report archive back to Feb 14 |
| API Stats | Performance matrix for ~300 endpoints with CSV export |
| Users | Top 20 impacted users per role group, user lookup |
| Workflows | Per-run user impact, activity breakdown, trend charts |
| ORA Errors | Oracle error analysis with incident correlation |
| Exceptions | API/SQL/App exception categories with trend charts |
| Replay | Historical replay validation and detection quality metrics |
| API Grades | A–F letter grades for all ~300 endpoints |
| SysOps | Alert toggles, eval mode, pipeline health, maintenance |
| Guide | Built-in usage guide and troubleshooting reference |
Key URLs
| Resource | URL |
|---|---|
| Dashboard | careon-alerts.digitalon.co.za |
| WhatsApp Alert Group | chat.whatsapp.com/B9IHGoFSICQDVflG3WKhTT |
| Support Portal | support.digitalon.co.za |
Escalation Path
- Dashboard — Check CareOn Alert Dashboard for current status and anomaly details
- Grafana — Drill into the Athena Platform for server-level metrics and API logs
- IM4HC — Verify CareOn service health at IM4HC Health Check
- ELK Stack — Check application logs at Elasticsearch/Kibana (elkfra)
- CheckMK — Server hardware metrics (CPU, memory, disk)
- DBA — For Oracle errors: check AWR, table statistics, index health
- Escalate — Contact [email protected] for critical issues requiring vendor intervention
Shomer-AI CareOn Early Warning System — Copyright © 2026 eMedys AG and OPEN MEDYS GMBH. All Rights Reserved.
Contact: Fabian Berger (CEO/Vorstand) — [email protected]
This tutorial was generated for the CareOn EWS v3 dashboard. For the latest documentation, see the built-in Guide page.