Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Prometheus does not train a machine-learning anomaly model in its core server. It collects timestamped metrics, evaluates PromQL recording and alerting rules, and sends firing alerts to Alertmanager. You can detect anomalies with fixed thresholds or statistical baselines inside Prometheus, then add a separately hosted model or a managed service such as Amazon Managed Service for Prometheus when seasonality and gradual drift make static rules inadequate.
What Prometheus does—and does not—provide
Prometheus stores labeled numeric time series and evaluates PromQL expressions on a schedule. Recording rules precompute expensive or frequently used expressions; alerting rules turn a true expression into an alert. Neither rule type silently learns a normal pattern from your history. A “Prometheus anomaly detector” built only with the open-source server is therefore a query-and-rules design, not an embedded ML system.
Alertmanager is a separate part of the stack. Prometheus sends it alerts, while Alertmanager groups related alerts, suppresses duplicates, applies inhibition and silencing, and delivers notifications to chat, ticketing systems, or paging services. A sophisticated detector can still create a bad on-call experience if this routing and noise-control layer is misconfigured.
Three practical ways to detect unusual behavior
| Approach | Best fit | Advantages | Costs and limits | Typical delivery |
|---|---|---|---|---|
| Fixed PromQL threshold | A service-level objective or a clear failure boundary, such as an error ratio above an agreed limit | Fast, transparent, inexpensive, and easy to test in a query browser | Needs manual limits and can misfire when traffic, seasonality, or capacity changes | Alerting rule to Alertmanager; page only when the symptom is actionable |
| Statistical baseline in PromQL | Metrics with a reasonably stable recent distribution or a known comparison period | Adapts to normal variation without a separate model service; the baseline and deviation are inspectable | Requires enough history and careful treatment of missing data, low traffic, deploys, and changing variance | Recording rules calculate the baseline; an alert compares the current value with it |
| Learned or managed detector | Seasonal, drifting, or multi-dimensional behavior that fixed rules cannot represent reliably | Can learn normal patterns and produce a deviation score or bands | Introduces data, model, sensitivity, hosting, and operational dependencies; explainability is lower than a simple query | Detector output is exposed to Prometheus or its alert path, then routed by Alertmanager |
Compare candidates on detection quality, time to detection, explainability, adaptation to growth and seasonality, history and cardinality requirements, operating cost, and the final delivery path. There is no established independent precision, recall, latency, or cost benchmark for these choices in the available evidence, so validate them against your own incident history instead of quoting a universal accuracy number.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitches#1 Best Overall
Build a useful PromQL detector first
1. Start with user-facing symptoms
Instrument latency, error rate, availability, and workload throughput before trying to model internal causes. A page should answer “what is the user experiencing and what should the responder do?” CPU, queue depth, or a single database counter can be valuable diagnostic context, but they should not page by themselves unless they are directly tied to an imminent, actionable failure.
2. Record stable, aggregated series
Aggregate dimensions that are not needed for the decision. This keeps dashboards and anomaly queries from repeatedly scanning expensive raw series and avoids teaching a model the noise of every URL, user, pod, or request ID. Preserve the labels needed to identify the affected service or region, and keep high-cardinality detail available for drill-down rather than for the paging expression.
groups:
- name: service-slIs
interval: 30s
rules:
- record: service:http_requests:rate5m
expr: sum by (service) (rate(http_requests_total[5m]))
- record: service:http_errors:rate5m
expr: sum by (service) (rate(http_requests_total{status=~"5.."}[5m]))
- record: service:http_error_ratio:5m
expr: service:http_errors:rate5m / service:http_requests:rate5m
Protect ratios from empty or near-zero traffic in your own environment. For low-volume services, a percentage can swing wildly; an availability or request-count condition may be a more honest signal.
3. Use a threshold with a persistence window
The for clause keeps an alert pending until its expression remains true for the configured duration. That is the simplest defense against a one-scrape spike. keep_firing_for, where supported by your Prometheus version, can keep an alert firing through a short data gap or a flapping resolution.
- alert: ServiceHighErrorRatio
expr: service:http_error_ratio:5m > 0.02
for: 10m
keep_firing_for: 5m
labels:
severity: page
annotations:
summary: "{{ $labels.service }} error ratio is elevated"
description: "The five-minute error ratio has exceeded the service limit for ten minutes."
Choose the limit from the service’s SLO, capacity test, or incident experience. Put the response procedure in an annotation or linked runbook maintained by your team; do not page on a condition for which nobody has an immediate action.
4. Add a statistical baseline when a fixed limit is the wrong shape
PromQL provides range functions such as avg_over_time and stddev_over_time. A recording rule can calculate a rolling mean and spread, and an alert can fire when the current value is several deviations away. The window and multiplier are policy choices, not accuracy guarantees.
Rank #4
- record: service:http_error_ratio:mean1h
expr: avg_over_time(service:http_error_ratio:5m[1h])
- record: service:http_error_ratio:stddev1h
expr: stddev_over_time(service:http_error_ratio:5m[1h])
- alert: ServiceErrorRatioOutsideBaseline
expr: service:http_error_ratio:5m > service:http_error_ratio:mean1h + 3 * clamp_min(service:http_error_ratio:stddev1h, 0.001)
for: 10m
labels:
severity: warning
annotations:
summary: "{{ $labels.service }} error ratio is outside its recent baseline"
description: "Investigate the service, recent deploys, and traffic volume before escalating."
For daily or weekly seasonality, compare with a corresponding historical window using PromQL’s offset modifier rather than averaging unrelated hours together. A baseline built from sparse samples, a recent outage, or a major migration will encode the wrong “normal”; exclude or segment those periods and inspect the series on a dashboard.
Stop short spikes from paging
Noise control has two independent parts:
- Rule persistence: use an appropriate range such as
rate(...[5m]), then require the condition withfor. The duration should exceed the blip you deliberately want to ignore but remain short enough for the service’s response objective. - Notification policy: let Alertmanager group alerts from the same incident, inhibit dependent symptoms when a higher-level outage is firing, and provide silences for planned work. Route warning-level observations to a dashboard or ticket instead of a page.
Do not solve every spike by making for extremely long. That can hide a real incident. Check scrape gaps, exporter failures, counter resets, deploy windows, and denominator changes when an alert flaps. If the metric is inherently bursty, alert on an aggregate over a longer window or on a user-facing symptom rather than the raw event count.
Best Value
When a learned detector is justified
Add a model only after a fixed threshold or transparent baseline has shown a specific limitation: for example, traffic follows a predictable daily pattern, grows steadily, or has several interacting seasonal signals. The model can run in a separate service, exporter, rule pipeline, or managed Prometheus capability. Prometheus remains the collection, label, query, and alert-transport layer; it does not become an ML trainer merely because a model consumes its data.
Amazon Managed Service for Prometheus option
AWS documents anomaly detection for Amazon Managed Service for Prometheus using the Random Cut Forest algorithm. The service learns normal behavior and seasonal variation, handles missing data, and returns four values: upper_band, lower_band, score, and value. Treat those outputs as signals that still need an operational policy.
AWS recommends at least 14 days of consistent metric history before enabling anomaly detection for optimal results. This is a setup guideline, not a measured accuracy guarantee. Start with stable, aggregated averages or sums rather than raw high-cardinality series, tune sensitivity for the trade-off between false positives and missed anomalies, and review the detector as traffic and application behavior change.
The documented workflow exposes CreateAnomalyDetector to create a detector in a workspace and PreviewAnomalyDetector to evaluate a Prometheus query over a selected historical period before implementation. Preview the bands and scores, compare them with known incidents and benign deploys, and route only a tested condition to paging.
Validate before putting humans in the loop
- Choose one user-facing metric and define the incident that should result from a detection.
- Graph the raw and aggregated series, including scrape gaps and deploy times.
- Run the PromQL rule or detector in a non-paging route and review historical periods containing both incidents and normal peaks.
- Measure operational outcomes: meaningful incidents found, noisy notifications, time to detection, and whether responders had a clear action.
- Adjust the window, threshold, baseline, or detector sensitivity; document the reason for each change.
- Promote only alerts that are urgent, important, actionable, and real. Keep diagnostic causes as dashboard panels or lower-urgency notifications unless they independently require immediate action.
Operational decision
Use PromQL thresholds for clear service limits, rolling or seasonal baselines when normal variation is understandable, and a managed or separately hosted detector only for behavior those methods cannot model cleanly. In every case, aggregate stable data, require persistence, validate historically, and let Alertmanager control delivery. An anomaly score without a defined response belongs on a dashboard—not in an on-call person’s pager.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




