Cram Sheet280 words

Topic 5.2 — Analyzing metrics — cram sheet

Topic 5.2 — Analyzing metrics · cram sheet

Infrastructure indicators

IndicatorWatch
CPUSaturation — and low CPU with high latency
MemoryGrowth without plateau (leak); paging
DiskIOPS and queue depth
NetworkThroughput, retransmits, connection exhaustion
  • Per resource ask: utilisation · saturation (queue) · errors.
  • Queueing starts before saturation — 60% with a deep queue is already a bottleneck.
  • Low CPU + high latency = waiting, not computing. The shape teams miss.

Analysing telemetry

Order: which operations are slow (p95/p99) → what are they waiting on (dependencies) → where in the code (traces). Starting at profiling answers the wrong question in detail.

Aggregates hide people. A healthy 0.1% global error rate can conceal one tenant failing every request. Segment by operation, region, client version, tenant. Correlate with release annotations — most regressions follow a deployment.

Distributed tracing

  • A correlation id propagated across service boundaries stitches spans into one operation.
  • A trace that stops partway = missing propagation, not a fast service.
  • Tracing separates the service that reports a failure from the one that caused it.

KQL

Shape: source → wheresummarizeorder bytake

OperatorDoes
whereFilter
summarizeAggregate — count(), avg(), percentile()
bin()Bucket timestamps for time series
project / extendChoose / add columns
joinCombine tables

Filter early, especially on time — less data scanned, cheaper and faster. percentile(duration, 95) is the latency view averages hide.

Ready to study Designing and Implementing Microsoft DevOps Solutions (AZ-400)?

Practice tests, flashcards, and all study notes — free, no sign-up needed.

Start Studying — Free