Storage IO
Storage throughput, queues and response time per volume.
Topic id: storage_io.
Reported every 5 minutes on clock boundaries for the window ending at t. sparklogs.window_coverage_pct is below 100 when collection covered only part of the window. Use chart buckets at least 5 minutes wide.
Fields
Read/write throughput, queues and response times for each volume.
| Field | Type | Unit | Meaning |
|---|---|---|---|
sparklogs.data.storage_io.volume | string | The volume's stable identity (its GUID), the same identity disk_volumes uses. | |
sparklogs.data.storage_io.display_name | string | The volume's drive letter or mount path, for display. | |
sparklogs.data.storage_io.volume_role | string | What this volume is for, the same value disk_volumes writes: os, fixed_data, or removable. Absent when the volume has no role. | |
sparklogs.data.storage_io.volume_role_code | integer | volume_role as the numeric code disk_volumes uses for the same role. | |
sparklogs.data.storage_io.device_class | string | The latency class of the volume's backing device: hdd or flash. Absent when the backing device is not fully resolved. | |
sparklogs.data.storage_io.device_class_code | float | device_class as the number the latency rules compare. | |
sparklogs.data.storage_io.latency_warn_ms | float | milliseconds | Latency threshold in milliseconds for degraded performance, chosen by device class and host class. |
sparklogs.data.storage_io.latency_error_ms | float | milliseconds | Latency threshold in milliseconds at or above which this volume is considered severely degraded. |
sparklogs.data.storage_io.latency_recover_ms | float | milliseconds | The latency this volume must fall back under before a latency episode is considered closed. |
sparklogs.data.storage_io.busy_pct_avg | float | percent | The average share of the window this volume's backing store spent busy. |
sparklogs.data.storage_io.busy_pct_p90_10s | float | percent | The 90th percentile, across the window's 10-second samples, of busy share. |
sparklogs.data.storage_io.busy_pct_max_10s | float | percent | The highest 10-second busy share the window observed. |
sparklogs.data.storage_io.read_iops_avg | float | iops | The average read rate over the window. |
sparklogs.data.storage_io.read_iops_max_10s | float | iops | The highest 10-second read rate the window observed. |
sparklogs.data.storage_io.write_iops_avg | float | iops | The average write rate over the window. |
sparklogs.data.storage_io.write_iops_max_10s | float | iops | The highest 10-second write rate the window observed. |
sparklogs.data.storage_io.read_mb_per_s_avg | float | megabytes_per_second | The average read throughput over the window, in 1024-based MB per second. |
sparklogs.data.storage_io.read_mb_per_s_max_10s | float | megabytes_per_second | The highest 10-second read throughput the window observed, in 1024-based MB per second. |
sparklogs.data.storage_io.write_mb_per_s_avg | float | megabytes_per_second | The average write throughput over the window, in 1024-based MB per second. |
sparklogs.data.storage_io.write_mb_per_s_max_10s | float | megabytes_per_second | The highest 10-second write throughput the window observed, in 1024-based MB per second. |
sparklogs.data.storage_io.read_latency_ms_avg | float | milliseconds | The count-weighted average read latency over the window. |
sparklogs.data.storage_io.write_latency_ms_avg | float | milliseconds | The count-weighted average write latency over the window. |
sparklogs.data.storage_io.latency_ms_p90_10s | float | milliseconds | The 90th percentile, across the window's 10-second mean-latency samples, of latency. Chosen over a per-IO percentile because the tail matters and a median converges on the mean. |
sparklogs.data.storage_io.latency_ms_max_10s | float | milliseconds | The highest 10-second mean latency the window observed. |
sparklogs.data.storage_io.avg_read_kb | float | kilobytes | The average size of a read over the window. |
sparklogs.data.storage_io.avg_write_kb | float | kilobytes | The average size of a write over the window. |
sparklogs.data.storage_io.queue_depth_avg | float | The average outstanding IO queue depth over the window. | |
sparklogs.data.storage_io.queue_depth_max_10s | float | The highest 10-second queue depth the window observed. | |
sparklogs.data.storage_io.topology_segment_count | integer | count | How many distinct topology signatures the window's samples held. More than two in one window marks topology_mixed and holds until every sample shares one topology again. |
sparklogs.data.storage_io.topology_mixed | bool | Whether the topology changed too many times within one window to trust a single reduction. Absent when the topology held steady. | |
sparklogs.data.storage_io.disk_latency_degraded_age_basis | string | onset: witnessed start. observed: already present when first seen, making age a lower bound. unknown_ongoing: no meaningful onset time. | |
sparklogs.data.storage_io.disk_latency_degraded_age_h | float | hours | How long this condition has been open, in hours. |
sparklogs.data.storage_io.disk_saturated_age_basis | string | onset: witnessed start. observed: already present when first seen, making age a lower bound. unknown_ongoing: no meaningful onset time. | |
sparklogs.data.storage_io.disk_saturated_age_h | float | hours | How long this condition has been open, in hours. |
Conditions
A condition is a state that holds for a while. The agent opens it when the host enters it, keeps it open while it lasts, and closes it when the host comes back out, so one episode answers for the whole stretch instead of one alert per sample.
Example
Inventory (every 5 minutes)
1 volume; "C" 0.8 ms p90, 240 IOPS.
sparklogs.data.storage_io.volume: volume:3b1a9c1e-0000-0000-0000-100000000001
sparklogs.data.storage_io.busy_pct_avg: 22.0
sparklogs.data.storage_io.queue_depth_avg: 1.0
sparklogs.data.storage_io.latency_ms_p90_10s: 0.8
SparkLogs: CONTEXT, Info, storage_io: INVENTORY: 1 volume; "C" 0.8 ms p90, 240 IOPS.
Selected conditions
disk_latency_degraded
Storage latency is severe while the disk is busy.
Also reported by: Storage IO
Impact: Workloads above the storage stack may stall or time out.
Example
started; volume "C" latency p90 320
sparklogs.instance: volume:3b1a9c1e-0000-0000-0000-100000000001
sparklogs.data.storage_io.volume: volume:3b1a9c1e-0000-0000-0000-100000000001
sparklogs.data.storage_io.disk_latency_degraded_age_h: 0.0
SparkLogs: disk_latency_degraded, Warning, storage_io: disk_latency_degraded: NOTABLE: started; volume "C" latency p90 320
| Case | Severity | Ticket class |
|---|---|---|
onset | Trace to Fatal | storage |
held | Trace to Fatal | storage |
recovered | Trace to Fatal | storage |
disk_saturated
The disk is busy, queueing and slow to respond.
Also reported by: Storage IO
Impact: Workloads above the storage stack may wait on IO.
Example
started; volume "C" busy 94% (threshold 90%)
sparklogs.instance: volume:3b1a9c1e-0000-0000-0000-100000000001
sparklogs.data.storage_io.volume: volume:3b1a9c1e-0000-0000-0000-100000000001
sparklogs.data.storage_io.busy_pct_avg: 94.0
sparklogs.data.storage_io.queue_depth_avg: 6.0
sparklogs.data.storage_io.latency_ms_p90_10s: 48.0
sparklogs.data.storage_io.disk_saturated_age_h: 0.0
SparkLogs: disk_saturated, Notice, storage_io: disk_saturated: NOTABLE: started; volume "C" busy 94% (threshold 90%)
| Case | Severity | Ticket class |
|---|---|---|
onset | Trace to Fatal | storage |
held | Trace to Fatal | storage |
recovered | Trace to Fatal | storage |