Investigation

SRE incident
D

Orders API latency p95 above 2s

open · high · orders · prod

RCA

1. Root cause cannot be determined from available telemetry because the orders service has no observable metrics or logs in Detective. i. Prometheus returned no series for service=~"orders|orders.*" [E1], indicating a likely evidence gap or connector/configuration issue rather than confirmed service behavior. ii. Loki returned no log streams matching error patterns for the orders app [E2], and no active Prometheus alerts or rules matched the incident context [E3]. iii. Prior postmortems for checkout and payment have low similarity, so they provide no applicable context for this orders latency incident. 2. The high-severity alert cannot be validated or investigated without obtaining authoritative monitoring data from the source systems.

Status: complete · Confidence: 29%

pendinglow risk

Prior verified fix: Fix for Orders API latency p95 above 2s

Target
orders
Executor
not set
Permissions
not set
Rollback
Review prior fix context before repeating the same remediation.
Expires
2026-09-03T08:01:50.089635Z
pendinglow risk

Prior verified fix: Fix for Orders API latency p95 above 2s

Target
orders
Executor
not set
Permissions
not set
Rollback
Review prior fix context before repeating the same remediation.
Expires
2026-09-03T08:01:53.069791Z

Resolve incident (2nd)

Record what fixed the incident. Each resolution is stored as Resolution #N · timestamp · "message" for postmortem search.

Previous resolutions

Resolution #1 · 2026-09-01T20:47:17Z · "test and fixed"

Chat

Ask Detective about the incident, evidence, approvals, or next diagnostics. Write actions are not executed.

detective

1. Root cause cannot be determined from available telemetry because the orders service has no observable metrics or logs in Detective. i. Prometheus returned no series for service=~"orders|orders.*" [E1], indicating a likely evidence gap or connector/configuration issue rather than confirmed service behavior. ii. Loki returned no log streams matching error patterns for the orders app [E2], and no active Prometheus alerts or rules matched the incident context [E3]. iii. Prior postmortems for checkout and payment have low similarity, so they provide no applicable context for this orders latency incident. 2. The high-severity alert cannot be validated or investigated without obtaining authoritative monitoring data from the source systems.

detective

1. Telemetry is absent for the orders service in the incident window. i. Prometheus returns zero series for up{service=~"orders|orders.*"}, so service availability cannot be confirmed [E1]. ii. Loki returns zero log streams for orders/* labels, so error, exception, crash, and panic signals cannot be evaluated [E2]. iii. No matching Prometheus alerting rules or active alerts exist for this incident context [E3]. 2. The incident title reports orders API p95 latency above 2s, but no metric, log, or rule evidence is available to verify which component, deployment, or dependency is the source. 3. Given the complete absence of observability data, the immediate actionable finding is an evidence gap or connector/configuration issue: Prometheus/Loki query selectors, namespaces, or endpoints may be misconfigured, unavailable, or returning no matching streams. Prior postmortems for checkout and payment are low-similarity and do not inform this incident.

Application Context

Deployment-supplied mapping for owners, repositories, pipelines, runtime labels, and infrastructure topology.

not mapped

No application context matched this incident. Add an app context mapping file and set APP_CONTEXT_PATH to enable repo, pipeline, and infrastructure correlation.

Infrastructure Topology

Where to investigate for this application — CDN, load balancers, Kubernetes, data stores, and custom hints from app-context.json.

not mapped

Add an infrastructure block to the matched application in your context file (for example CloudFront, ALB, Kubernetes, Redis, Kafka) so Detective knows where to collect evidence.

Ownership And Escalation

Team routing and escalation hints from resolved application context.

not mapped

No ownership metadata matched this incident. Add owner and escalation fields to the application context mapping.

Runbook Suggestions

Matching local runbooks and Confluence knowledge collected during investigation.

0 suggestion(s)

No runbook evidence yet. Mount `RUNBOOKS_PATH` or configure Confluence to collect runbook guidance during investigation.

Recent Changes

Commits, merged PRs/MRs, and CI/CD runs collected from application context near the alert window.

0 item(s)

No SCM change evidence yet. Enable application context, set GITHUB_TOKEN or GITLAB_TOKEN, and rerun investigation when the incident maps to a repository.

Similar Past Incidents

Vector-ranked historical RCA memory plus related incidents by fingerprint or prior occurrence.

2 match(es)
checkoutprodsimilarity 8%

Postmortem: Checkout pod is crash looping

Open incident

fixed

paymentprodsimilarity 8%

Postmortem: Payment pod is not running

Open incident

fixed

Connector Policy For This Incident

Select connectors from alert labels, application context, and infrastructure topology.

auto

Selected: loki, prometheus

Investigation Coverage

Connector selection, evidence collection, and gaps for the latest investigation.

22 connector(s)
Selected2
Succeeded2
Skipped20
Failed/Timed Out0

Evidence by source: loki: 2, prometheus: 4

Selected connectors: Prometheus, Loki

Connector Timeline

skipped

GitHub

github
Capability
evidence
Configured
no
Active
no
Items
0
Duration
0 ms

not configured

skipped

GitHub Docs

github_docs
Capability
evidence
Configured
no
Active
no
Items
0
Duration
0 ms

not configured

skipped

SCM Change Correlation

scm_changes
Capability
evidence
Configured
no
Active
no
Items
0
Duration
0 ms

not configured

skipped

Datadog

datadog
Capability
evidence
Configured
no
Active
no
Items
0
Duration
0 ms

not configured

skipped

Grafana

grafana
Capability
evidence
Configured
no
Active
no
Items
0
Duration
0 ms

not configured

skipped

Tempo

tempo
Capability
evidence
Configured
no
Active
no
Items
0
Duration
0 ms

not configured

skipped

Elasticsearch

elasticsearch
Capability
evidence
Configured
no
Active
no
Items
0
Duration
0 ms

not configured

skipped

AWS

aws
Capability
evidence
Configured
no
Active
no
Items
0
Duration
0 ms

not configured

skipped

AWS EKS

aws_eks
Capability
evidence
Configured
no
Active
no
Items
0
Duration
0 ms

not configured

skipped

Azure

azure
Capability
evidence
Configured
no
Active
no
Items
0
Duration
0 ms

not configured

skipped

Kubernetes

kubernetes
Capability
evidence
Configured
yes
Active
yes
Items
0
Duration
0 ms

not selected by investigation policy

skipped

Postgres

postgres
Capability
evidence
Configured
no
Active
no
Items
0
Duration
0 ms

not configured

skipped

MySQL

mysql
Capability
evidence
Configured
no
Active
no
Items
0
Duration
0 ms

not configured

skipped

MongoDB

mongodb
Capability
evidence
Configured
no
Active
no
Items
0
Duration
0 ms

not configured

skipped

Confluence

confluence
Capability
evidence
Configured
no
Active
no
Items
0
Duration
0 ms

not configured

skipped

Runbooks

runbooks
Capability
evidence
Configured
no
Active
no
Items
0
Duration
0 ms

not configured

skipped

Jenkins

jenkins
Capability
evidence
Configured
no
Active
no
Items
0
Duration
0 ms

not configured

skipped

Argo CD

argocd
Capability
evidence
Configured
no
Active
no
Items
0
Duration
0 ms

not configured

skipped

Flux

flux
Capability
evidence
Configured
yes
Active
yes
Items
0
Duration
0 ms

not selected by investigation policy

skipped

Bash Diagnostics

bash
Capability
evidence
Configured
no
Active
no
Items
0
Duration
0 ms

not configured

success

Prometheus

prometheus
Capability
evidence
Configured
yes
Active
yes
Items
2
Duration
374 ms

configured

success

Loki

loki
Capability
evidence
Configured
yes
Active
yes
Items
1
Duration
375 ms

configured

Evidence

E1prometheus

Prometheus service availability

No Prometheus series returned for query: up{service=~"orders|orders.*"}

E2loki

Loki logs around alert window

No Loki log streams returned. Attempted queries: {app=~"orders|orders.*"} |~ "(?i)(error|exception|fail|crash|panic|fatal|back-off)"

E3prometheus

Prometheus alert/rule metadata

No matching Prometheus active alerts or alerting rules were found for this incident context.

E1prometheus

Prometheus service availability

No Prometheus series returned for query: up{service=~"orders|orders.*"}

E2loki

Loki logs around alert window

No Loki log streams returned. Attempted queries: {app=~"orders|orders.*"} |~ "(?i)(error|exception|fail|crash|panic|fatal|back-off)"

E3prometheus

Prometheus alert/rule metadata

No matching Prometheus active alerts or alerting rules were found for this incident context.

Alerts