Skip to main content

Host performance

6readings
6conditions
1themes fed
Livestatus

Whole-host CPU and memory posture over a measured window.

Topic id: performance.

Reported every 5 minutes on clock boundaries for the window ending at t. sparklogs.window_coverage_pct is below 100 when collection covered only part of the window. Use chart buckets at least 5 minutes wide.

Fields​

CPU load, queueing and memory readings over the reporting window.

FieldTypeUnitMeaning
sparklogs.data.performance.cpu_pct_time_over_90floatpercentPercent of the window's sample pairs with CPU busy over 90%.
sparklogs.data.performance.cpu_pct_time_over_70floatpercentPercent of the window's sample pairs with CPU busy over 70%.
sparklogs.data.performance.cpu_busy_pct_avgfloatpercentThe host's average CPU busy share over the window, core-normalized.
sparklogs.data.performance.cpu_busy_pct_max_10sfloatpercentThe highest 10-second CPU busy share the window observed, core-normalized: the burst the average hides.
sparklogs.data.performance.cpu_busy_pct_p90_10sfloatpercentThe 90th percentile, across the window's 10-second samples, of the CPU busy share, core-normalized.
sparklogs.data.performance.cpu_kernel_pct_of_busy_avgfloatpercentThe average share of busy CPU time spent in kernel mode.
sparklogs.data.performance.cpu_user_pct_of_busy_avgfloatpercentThe average share of busy CPU time spent in user mode.
sparklogs.data.performance.cpu_kernel_excl_drivers_pct_of_busy_avgfloatpercentAverage share of busy CPU time spent in kernel mode, excluding interrupt and deferred procedure call time.
sparklogs.data.performance.cpu_interrupt_pct_avgfloatpercentThe average share of busy CPU time spent servicing interrupts.
sparklogs.data.performance.cpu_interrupt_pct_max_10sfloatpercentThe highest 10-second interrupt share the window observed.
sparklogs.data.performance.cpu_dpc_pct_avgfloatpercentThe average share of busy CPU time spent in deferred procedure calls.
sparklogs.data.performance.cpu_dpc_pct_max_10sfloatpercentThe highest 10-second DPC share the window observed.
sparklogs.data.performance.cpu_interrupt_dpc_pct_avgfloatpercentAverage share of busy CPU time spent servicing interrupts and deferred procedure calls.
sparklogs.data.performance.cpu_clock_pct_of_base_avgfloatpercentAverage CPU clock as a percentage of rated non-turbo base frequency. Can exceed 100 under turbo.
sparklogs.data.performance.cpu_base_clock_mhzintegerThe rated non-turbo base clock, in MHz. Absent when the host did not report one.
sparklogs.data.performance.cpu_clock_mhz_avgfloatAverage effective clock in MHz: rated base frequency multiplied by cpu_clock_pct_of_base_avg / 100. Absent without a base frequency.
sparklogs.data.performance.run_queue_p90_10sfloatThe 90th percentile, across the window's 10-second samples, of how many threads were ready to run but waiting for a core. On a virtual machine the queue can be threads waiting on host dispatch, which the guest does not count as busy.
sparklogs.data.performance.run_queue_per_core_p90_10sfloatrun_queue_p90_10s divided by the logical core count. On a VM, queued threads may be waiting for host scheduling without appearing as busy guest CPU.
sparklogs.data.performance.logical_core_countintegercountHow many logical processors the host has, which the per-core queue reading is divided by.
sparklogs.data.performance.commit_pctfloatpercentThe current commit charge as a percentage of the commit limit.
sparklogs.data.performance.commit_pct_max_windowfloatpercentThe highest commit charge percentage the window observed.
sparklogs.data.performance.hard_faults_per_sfloatper_secondHard page faults per second, latest 10s sample.
sparklogs.data.performance.ram_pct_time_in_hard_fault_stormfloatpercentPercentage of the window's hard-fault readings above the configured storm-rate threshold.
sparklogs.data.performance.ram_total_bytesintegerbytesInstalled physical RAM, the same figure system_info reports, at the end of the window.
sparklogs.data.performance.ram_available_bytesintegerbytesPhysical RAM available to new allocations (free, zeroed and standby pages) at the end of the window.
sparklogs.data.performance.ram_standby_bytesintegerbytesRAM holding cached pages the memory manager can repurpose (the standby list, all priorities) at the end of the window. Part of available.
sparklogs.data.performance.ram_modified_bytesintegerbytesRAM holding dirty pages waiting to be written out (the modified list) at the end of the window.
sparklogs.data.performance.ram_compressed_bytesintegerbytesResident memory held by the memory-compression store, from the latest successful process enumeration (at most one process-table generation old). Zero when compression is off. Absent until the first successful enumeration, or when no process-table topic is collected.
sparklogs.data.performance.cpu_busy_age_basisstringonset: witnessed start. observed: already present when first seen, making age a lower bound. unknown_ongoing: no meaningful onset time.
sparklogs.data.performance.cpu_busy_age_hfloathoursHow long this condition has been open, in hours.
sparklogs.data.performance.cpu_interrupt_storm_age_basisstringonset: witnessed start. observed: already present when first seen, making age a lower bound. unknown_ongoing: no meaningful onset time.
sparklogs.data.performance.cpu_interrupt_storm_age_hfloathoursHow long this condition has been open, in hours.
sparklogs.data.performance.cpu_kernel_dominated_age_basisstringonset: witnessed start. observed: already present when first seen, making age a lower bound. unknown_ongoing: no meaningful onset time.
sparklogs.data.performance.cpu_kernel_dominated_age_hfloathoursHow long this condition has been open, in hours.
sparklogs.data.performance.cpu_throttled_under_load_age_basisstringonset: witnessed start. observed: already present when first seen, making age a lower bound. unknown_ongoing: no meaningful onset time.
sparklogs.data.performance.cpu_throttled_under_load_age_hfloathoursHow long this condition has been open, in hours.
sparklogs.data.performance.ram_commit_near_cap_age_basisstringonset: witnessed start. observed: already present when first seen, making age a lower bound. unknown_ongoing: no meaningful onset time.
sparklogs.data.performance.ram_commit_near_cap_age_hfloathoursHow long this condition has been open, in hours.
sparklogs.data.performance.ram_hard_fault_storm_age_basisstringonset: witnessed start. observed: already present when first seen, making age a lower bound. unknown_ongoing: no meaningful onset time.
sparklogs.data.performance.ram_hard_fault_storm_age_hfloathoursHow long this condition has been open, in hours.

Conditions​

A condition is a state that holds for a while. The agent opens it when the host enters it, keeps it open while it lasts, and closes it when the host comes back out, so one episode answers for the whole stretch instead of one alert per sample.

ConditionSeverityHow an episode ends
cpu busy (cpu_busy)Display to NoticeIt closes when the measurement falls back past its recovery point.
cpu interrupt storm (cpu_interrupt_storm)Warning to ErrorIt closes when the measurement falls back past its recovery point.
cpu kernel dominated (cpu_kernel_dominated)NoticeIt closes when any one of the several recovery conditions is met.
cpu throttled under load (cpu_throttled_under_load)Warning to ErrorIt closes on a recovery rule written for this condition, which reads more than one measurement together.
committed memory near its limit (ram_commit_near_cap)Display to CriticalIt closes when the measurement falls back past its recovery point.
hard fault storm (ram_hard_fault_storm)SeriousIt closes when any one of the several recovery conditions is met.

Example​

Inventory (every 5 minutes)

cpu 12% avg, commit 48%, run queue 0.3/core, interrupts 1.2%, hard faults 6/s.

sparklogs.data.performance.cpu_pct_time_over_90: 0.0
sparklogs.data.performance.run_queue_per_core_p90_10s: 0.3
sparklogs.data.performance.cpu_interrupt_dpc_pct_avg: 1.2
sparklogs.data.performance.cpu_busy_pct_avg: 12.0
sparklogs.data.performance.cpu_kernel_excl_drivers_pct_of_busy_avg: 18.0
sparklogs.data.performance.commit_pct: 48.0
sparklogs.data.performance.ram_pct_time_in_hard_fault_storm: 0.0

SparkLogs: CONTEXT, Info, performance: INVENTORY: cpu 12% avg, commit 48%, run queue 0.3/core, interrupts 1.2%, hard faults 6/s.

Selected conditions​

cpu_busy​

CPU is busy.

Also reported by: Host performance

Example

started; time over 90% CPU 100% (threshold 80%), run queue 5/core (boundary 4/core)

sparklogs.data.performance.cpu_pct_time_over_90: 100.0
sparklogs.data.performance.run_queue_per_core_p90_10s: 5.0
sparklogs.data.performance.cpu_busy_age_h: 0.0

SparkLogs: cpu_busy, Notice, performance: cpu_busy: NOTABLE: started; time over 90% CPU 100% (threshold 80%), run queue 5/core (boundary 4/core)

Example

cleared after 0h51m, peaked Notice, relapses 1; time over 90% CPU 0% (clears at 70%)

sparklogs.data.performance.cpu_pct_time_over_90: 0.0
sparklogs.data.performance.run_queue_per_core_p90_10s: 0.13

SparkLogs: cpu_busy, Info, performance: cpu_busy: RECOVERED: cleared after 0h51m, peaked Notice, relapses 1; time over 90% CPU 0% (clears at 70%)

Example

subsiding, not yet cleared; time over 90% CPU 76.67% (clears at 70%), open for 0h25m

sparklogs.data.performance.cpu_pct_time_over_90: 76.67
sparklogs.data.performance.run_queue_per_core_p90_10s: 0.13
sparklogs.data.performance.cpu_busy_age_h: 0.42

SparkLogs: cpu_busy, Display, performance: cpu_busy: ELEVATED: subsiding, not yet cleared; time over 90% CPU 76.67% (clears at 70%), open for 0h25m

Example

relapsed; time over 90% CPU 100% (threshold 80%), run queue 5/core (boundary 4/core), open for 0h40m, relapses 1

sparklogs.data.performance.cpu_pct_time_over_90: 100.0
sparklogs.data.performance.run_queue_per_core_p90_10s: 5.0
sparklogs.data.performance.cpu_busy_age_h: 0.67

SparkLogs: cpu_busy, Notice, performance: cpu_busy: ELEVATED: relapsed; time over 90% CPU 100% (threshold 80%), run queue 5/core (boundary 4/core), open for 0h40m, relapses 1

CaseSeverityTicket class
onsetTrace to Fatalperformance
heldTrace to Fatalperformance
recoveredTrace to Fatalperformance

cpu_interrupt_storm​

CPU time is dominated by interrupt and DPC handling.

Also reported by: Host performance

Impact: Device or driver interrupt load can starve ordinary work on the host.

Example

started; interrupt+DPC 34% (threshold 30%)

sparklogs.data.performance.cpu_interrupt_dpc_pct_avg: 34.0
sparklogs.data.performance.cpu_interrupt_storm_age_h: 0.0

SparkLogs: cpu_interrupt_storm, Warning, performance: cpu_interrupt_storm: NOTABLE: started; interrupt+DPC 34% (threshold 30%)

Example

cleared after 0h50m, peaked Warning, relapses 1; interrupt+DPC 2% (clears at 20%)

sparklogs.data.performance.cpu_interrupt_dpc_pct_avg: 2.0

SparkLogs: cpu_interrupt_storm, Info, performance: cpu_interrupt_storm: RECOVERED: cleared after 0h50m, peaked Warning, relapses 1; interrupt+DPC 2% (clears at 20%)

Example

subsiding, not yet cleared; interrupt+DPC 25% (clears at 20%), open for 0h26m

sparklogs.data.performance.cpu_interrupt_dpc_pct_avg: 25.0
sparklogs.data.performance.cpu_interrupt_storm_age_h: 0.43

SparkLogs: cpu_interrupt_storm, Notice, performance: cpu_interrupt_storm: ELEVATED: subsiding, not yet cleared; interrupt+DPC 25% (clears at 20%), open for 0h26m

Example

relapsed; interrupt+DPC 34% (threshold 30%), open for 0h42m, relapses 1

sparklogs.data.performance.cpu_interrupt_dpc_pct_avg: 34.0
sparklogs.data.performance.cpu_interrupt_storm_age_h: 0.7

SparkLogs: cpu_interrupt_storm, Warning, performance: cpu_interrupt_storm: ELEVATED: relapsed; interrupt+DPC 34% (threshold 30%), open for 0h42m, relapses 1

CaseSeverityTicket class
onsetTrace to Fatalperformance
heldTrace to Fatalperformance
recoveredTrace to Fatalperformance

cpu_kernel_dominated​

Busy CPU time is mostly kernel work, excluding interrupts and deferred procedure calls.

Also reported by: Host performance

Impact: Application throughput can be lower than the busy figure alone suggests.

Example

started; CPU 84% (threshold 70%), kernel 85.26% of busy (threshold 70%)

sparklogs.data.performance.cpu_busy_pct_avg: 84.0
sparklogs.data.performance.cpu_kernel_excl_drivers_pct_of_busy_avg: 85.26
sparklogs.data.performance.cpu_kernel_dominated_age_h: 0.0

SparkLogs: cpu_kernel_dominated, Notice, performance: cpu_kernel_dominated: NOTABLE: started; CPU 84% (threshold 70%), kernel 85.26% of busy (threshold 70%)

Example

cleared after 0h50m, peaked Notice, relapses 1; CPU 40%, kernel 25% of busy (clears below 60% CPU or 60% kernel)

sparklogs.data.performance.cpu_busy_pct_avg: 40.0
sparklogs.data.performance.cpu_kernel_excl_drivers_pct_of_busy_avg: 25.0

SparkLogs: cpu_kernel_dominated, Info, performance: cpu_kernel_dominated: RECOVERED: cleared after 0h50m, peaked Notice, relapses 1; CPU 40%, kernel 25% of busy (clears below 60% CPU or 60% kernel)

Example

subsiding, not yet cleared; CPU 65%, kernel 61.92% of busy (clears below 60% CPU or 60% kernel), open for 0h27m

sparklogs.data.performance.cpu_busy_pct_avg: 65.0
sparklogs.data.performance.cpu_kernel_excl_drivers_pct_of_busy_avg: 61.92
sparklogs.data.performance.cpu_kernel_dominated_age_h: 0.45

SparkLogs: cpu_kernel_dominated, Notice, performance: cpu_kernel_dominated: ELEVATED: subsiding, not yet cleared; CPU 65%, kernel 61.92% of busy (clears below 60% CPU or 60% kernel), open for 0h27m

Example

relapsed; CPU 84% (threshold 70%), kernel 85.62% of busy (threshold 70%), open for 0h41m, relapses 1

sparklogs.data.performance.cpu_busy_pct_avg: 84.0
sparklogs.data.performance.cpu_kernel_excl_drivers_pct_of_busy_avg: 85.62
sparklogs.data.performance.cpu_kernel_dominated_age_h: 0.68

SparkLogs: cpu_kernel_dominated, Notice, performance: cpu_kernel_dominated: ELEVATED: relapsed; CPU 84% (threshold 70%), kernel 85.62% of busy (threshold 70%), open for 0h41m, relapses 1

CaseSeverityTicket class
onsetTrace to Fatalperformance
heldTrace to Fatalperformance
recoveredTrace to Fatalperformance

cpu_throttled_under_load​

The CPU is running below its rated frequency while under load.

Also reported by: Host performance

Impact: Work takes longer than the hardware would otherwise allow; thermal, power or firmware limits are the usual cause.

Example

started; CPU 82% (threshold 70%), performance 45% (below 50%)

sparklogs.data.performance.cpu_throttled_under_load_age_h: 0.0

SparkLogs: cpu_throttled_under_load, Error, performance: cpu_throttled_under_load: NOTABLE: started; CPU 82% (threshold 70%), performance 45% (below 50%)

Example

cleared after 0h50m, peaked Error, relapses 1; CPU 40%, performance 100% (clears below 60% CPU or at 80% performance)

SparkLogs: cpu_throttled_under_load, Info, performance: cpu_throttled_under_load: RECOVERED: cleared after 0h50m, peaked Error, relapses 1; CPU 40%, performance 100% (clears below 60% CPU or at 80% performance)

Example

subsiding, not yet cleared; CPU 65%, performance 70% (clears below 60% CPU or at 80% performance), open for 0h27m

sparklogs.data.performance.cpu_throttled_under_load_age_h: 0.45

SparkLogs: cpu_throttled_under_load, Notice, performance: cpu_throttled_under_load: ELEVATED: subsiding, not yet cleared; CPU 65%, performance 70% (clears below 60% CPU or at 80% performance), open for 0h27m

Example

relapsed; CPU 82% (threshold 70%), performance 45% (below 50%), open for 0h41m, relapses 1

sparklogs.data.performance.cpu_throttled_under_load_age_h: 0.68

SparkLogs: cpu_throttled_under_load, Warning, performance: cpu_throttled_under_load: ELEVATED: relapsed; CPU 82% (threshold 70%), performance 45% (below 50%), open for 0h41m, relapses 1

CaseSeverityTicket class
onsetTrace to Fatalperformance
heldTrace to Fatalperformance
recoveredTrace to Fatalperformance

ram_commit_near_cap​

Committed memory is high.

Also reported by: Host performance

Example

started; commit 92% (threshold 90%)

sparklogs.data.performance.commit_pct: 92.0
sparklogs.data.performance.ram_commit_near_cap_age_h: 0.0

SparkLogs: ram_commit_near_cap, Notice, performance: ram_commit_near_cap: NOTABLE: started; commit 92% (threshold 90%)

Example

cleared after 0h44m, peaked Notice, relapses 1; commit 70% (clears at 80%)

sparklogs.data.performance.commit_pct: 70.0

SparkLogs: ram_commit_near_cap, Info, performance: ram_commit_near_cap: RECOVERED: cleared after 0h44m, peaked Notice, relapses 1; commit 70% (clears at 80%)

Example

subsiding, not yet cleared; commit 82% (clears at 80%), open for 0h18m

sparklogs.data.performance.commit_pct: 82.0
sparklogs.data.performance.ram_commit_near_cap_age_h: 0.3

SparkLogs: ram_commit_near_cap, Display, performance: ram_commit_near_cap: ELEVATED: subsiding, not yet cleared; commit 82% (clears at 80%), open for 0h18m

Example

relapsed; commit 92% (threshold 90%), open for 0h39m, relapses 1

sparklogs.data.performance.commit_pct: 92.0
sparklogs.data.performance.ram_commit_near_cap_age_h: 0.65

SparkLogs: ram_commit_near_cap, Notice, performance: ram_commit_near_cap: ELEVATED: relapsed; commit 92% (threshold 90%), open for 0h39m, relapses 1

CaseSeverityTicket class
onsetTrace to Fatalperformance
heldTrace to Fatalperformance
recoveredTrace to Fatalperformance

ram_hard_fault_storm​

High hard-page-fault activity persists.

Also reported by: Host performance

Impact: Memory-related disk reads may slow workloads.

Example

started; commit 97% (threshold 95%), hard-fault duty 100% (threshold 50%)

sparklogs.data.performance.commit_pct: 97.0
sparklogs.data.performance.ram_pct_time_in_hard_fault_storm: 100.0
sparklogs.data.performance.ram_hard_fault_storm_age_h: 0.0

SparkLogs: ram_hard_fault_storm, Serious, performance: ram_hard_fault_storm: NOTABLE: started; commit 97% (threshold 95%), hard-fault duty 100% (threshold 50%)

Example

cleared after 0h49m, peaked Serious, relapses 1; commit 70%, hard-fault duty 0% (clears below 85% commit or 40% duty)

sparklogs.data.performance.commit_pct: 70.0
sparklogs.data.performance.ram_pct_time_in_hard_fault_storm: 0.0

SparkLogs: ram_hard_fault_storm, Info, performance: ram_hard_fault_storm: RECOVERED: cleared after 0h49m, peaked Serious, relapses 1; commit 70%, hard-fault duty 0% (clears below 85% commit or 40% duty)

Example

subsiding, not yet cleared; commit 90%, hard-fault duty 45.16% (clears below 85% commit or 40% duty), open for 0h23m

sparklogs.data.performance.commit_pct: 90.0
sparklogs.data.performance.ram_pct_time_in_hard_fault_storm: 45.16
sparklogs.data.performance.ram_hard_fault_storm_age_h: 0.38

SparkLogs: ram_hard_fault_storm, Notice, performance: ram_hard_fault_storm: ELEVATED: subsiding, not yet cleared; commit 90%, hard-fault duty 45.16% (clears below 85% commit or 40% duty), open for 0h23m

Example

relapsed; commit 97% (threshold 95%), hard-fault duty 100% (threshold 50%), open for 0h40m, relapses 1

sparklogs.data.performance.commit_pct: 97.0
sparklogs.data.performance.ram_pct_time_in_hard_fault_storm: 100.0
sparklogs.data.performance.ram_hard_fault_storm_age_h: 0.67

SparkLogs: ram_hard_fault_storm, Serious, performance: ram_hard_fault_storm: ELEVATED: relapsed; commit 97% (threshold 95%), hard-fault duty 100% (threshold 50%), open for 0h40m, relapses 1

CaseSeverityTicket class
onsetTrace to Fatalperformance
heldTrace to Fatalperformance
recoveredTrace to Fatalperformance