Skip to main content

Storage IO

2readings
2conditions
2themes fed
Livestatus

Storage throughput, queues and response time per volume.

Topic id: storage_io.

Reported every 5 minutes on clock boundaries for the window ending at t. sparklogs.window_coverage_pct is below 100 when collection covered only part of the window. Use chart buckets at least 5 minutes wide.

Fields​

Read/write throughput, queues and response times for each volume.

FieldTypeUnitMeaning
sparklogs.data.storage_io.volumestringThe volume's stable identity (its GUID), the same identity disk_volumes uses.
sparklogs.data.storage_io.display_namestringThe volume's drive letter or mount path, for display.
sparklogs.data.storage_io.volume_rolestringWhat this volume is for, the same value disk_volumes writes: os, fixed_data, or removable. Absent when the volume has no role.
sparklogs.data.storage_io.volume_role_codeintegervolume_role as the numeric code disk_volumes uses for the same role.
sparklogs.data.storage_io.device_classstringThe latency class of the volume's backing device: hdd or flash. Absent when the backing device is not fully resolved.
sparklogs.data.storage_io.device_class_codefloatdevice_class as the number the latency rules compare.
sparklogs.data.storage_io.latency_warn_msfloatmillisecondsLatency threshold in milliseconds for degraded performance, chosen by device class and host class.
sparklogs.data.storage_io.latency_error_msfloatmillisecondsLatency threshold in milliseconds at or above which this volume is considered severely degraded.
sparklogs.data.storage_io.latency_recover_msfloatmillisecondsThe latency this volume must fall back under before a latency episode is considered closed.
sparklogs.data.storage_io.busy_pct_avgfloatpercentThe average share of the window this volume's backing store spent busy.
sparklogs.data.storage_io.busy_pct_p90_10sfloatpercentThe 90th percentile, across the window's 10-second samples, of busy share.
sparklogs.data.storage_io.busy_pct_max_10sfloatpercentThe highest 10-second busy share the window observed.
sparklogs.data.storage_io.read_iops_avgfloatiopsThe average read rate over the window.
sparklogs.data.storage_io.read_iops_max_10sfloatiopsThe highest 10-second read rate the window observed.
sparklogs.data.storage_io.write_iops_avgfloatiopsThe average write rate over the window.
sparklogs.data.storage_io.write_iops_max_10sfloatiopsThe highest 10-second write rate the window observed.
sparklogs.data.storage_io.read_mb_per_s_avgfloatmegabytes_per_secondThe average read throughput over the window, in 1024-based MB per second.
sparklogs.data.storage_io.read_mb_per_s_max_10sfloatmegabytes_per_secondThe highest 10-second read throughput the window observed, in 1024-based MB per second.
sparklogs.data.storage_io.write_mb_per_s_avgfloatmegabytes_per_secondThe average write throughput over the window, in 1024-based MB per second.
sparklogs.data.storage_io.write_mb_per_s_max_10sfloatmegabytes_per_secondThe highest 10-second write throughput the window observed, in 1024-based MB per second.
sparklogs.data.storage_io.read_latency_ms_avgfloatmillisecondsThe count-weighted average read latency over the window.
sparklogs.data.storage_io.write_latency_ms_avgfloatmillisecondsThe count-weighted average write latency over the window.
sparklogs.data.storage_io.latency_ms_p90_10sfloatmillisecondsThe 90th percentile, across the window's 10-second mean-latency samples, of latency. Chosen over a per-IO percentile because the tail matters and a median converges on the mean.
sparklogs.data.storage_io.latency_ms_max_10sfloatmillisecondsThe highest 10-second mean latency the window observed.
sparklogs.data.storage_io.avg_read_kbfloatkilobytesThe average size of a read over the window.
sparklogs.data.storage_io.avg_write_kbfloatkilobytesThe average size of a write over the window.
sparklogs.data.storage_io.queue_depth_avgfloatThe average outstanding IO queue depth over the window.
sparklogs.data.storage_io.queue_depth_max_10sfloatThe highest 10-second queue depth the window observed.
sparklogs.data.storage_io.topology_segment_countintegercountHow many distinct topology signatures the window's samples held. More than two in one window marks topology_mixed and holds until every sample shares one topology again.
sparklogs.data.storage_io.topology_mixedboolWhether the topology changed too many times within one window to trust a single reduction. Absent when the topology held steady.
sparklogs.data.storage_io.disk_latency_degraded_age_basisstringonset: witnessed start. observed: already present when first seen, making age a lower bound. unknown_ongoing: no meaningful onset time.
sparklogs.data.storage_io.disk_latency_degraded_age_hfloathoursHow long this condition has been open, in hours.
sparklogs.data.storage_io.disk_saturated_age_basisstringonset: witnessed start. observed: already present when first seen, making age a lower bound. unknown_ongoing: no meaningful onset time.
sparklogs.data.storage_io.disk_saturated_age_hfloathoursHow long this condition has been open, in hours.

Conditions​

A condition is a state that holds for a while. The agent opens it when the host enters it, keeps it open while it lasts, and closes it when the host comes back out, so one episode answers for the whole stretch instead of one alert per sample.

ConditionSeverityHow an episode ends
disk latency degraded (disk_latency_degraded)Warning to ErrorIt closes on a recovery rule written for this condition, which reads more than one measurement together.
disk saturated (disk_saturated)NoticeIt closes when the measurement falls back past its recovery point.

Example​

Inventory (every 5 minutes)

1 volume; "C" 0.8 ms p90, 240 IOPS.

sparklogs.data.storage_io.volume: volume:3b1a9c1e-0000-0000-0000-100000000001
sparklogs.data.storage_io.busy_pct_avg: 22.0
sparklogs.data.storage_io.queue_depth_avg: 1.0
sparklogs.data.storage_io.latency_ms_p90_10s: 0.8

SparkLogs: CONTEXT, Info, storage_io: INVENTORY: 1 volume; "C" 0.8 ms p90, 240 IOPS.

Selected conditions​

disk_latency_degraded​

Storage latency is severe while the disk is busy.

Also reported by: Storage IO

Impact: Workloads above the storage stack may stall or time out.

Example

started; volume "C" latency p90 320

sparklogs.instance: volume:3b1a9c1e-0000-0000-0000-100000000001
sparklogs.data.storage_io.volume: volume:3b1a9c1e-0000-0000-0000-100000000001
sparklogs.data.storage_io.disk_latency_degraded_age_h: 0.0

SparkLogs: disk_latency_degraded, Warning, storage_io: disk_latency_degraded: NOTABLE: started; volume "C" latency p90 320

CaseSeverityTicket class
onsetTrace to Fatalstorage
heldTrace to Fatalstorage
recoveredTrace to Fatalstorage

disk_saturated​

The disk is busy, queueing and slow to respond.

Also reported by: Storage IO

Impact: Workloads above the storage stack may wait on IO.

Example

started; volume "C" busy 94% (threshold 90%)

sparklogs.instance: volume:3b1a9c1e-0000-0000-0000-100000000001
sparklogs.data.storage_io.volume: volume:3b1a9c1e-0000-0000-0000-100000000001
sparklogs.data.storage_io.busy_pct_avg: 94.0
sparklogs.data.storage_io.queue_depth_avg: 6.0
sparklogs.data.storage_io.latency_ms_p90_10s: 48.0
sparklogs.data.storage_io.disk_saturated_age_h: 0.0

SparkLogs: disk_saturated, Notice, storage_io: disk_saturated: NOTABLE: started; volume "C" busy 94% (threshold 90%)

CaseSeverityTicket class
onsetTrace to Fatalstorage
heldTrace to Fatalstorage
recoveredTrace to Fatalstorage