Host performance
Whole-host CPU and memory posture over a measured window.
Topic id: performance.
Reported every 5 minutes on clock boundaries for the window ending at t. sparklogs.window_coverage_pct is below 100 when collection covered only part of the window. Use chart buckets at least 5 minutes wide.
Fields
CPU load, queueing and memory readings over the reporting window.
| Field | Type | Unit | Meaning |
|---|---|---|---|
sparklogs.data.performance.cpu_pct_time_over_90 | float | percent | Percent of the window's sample pairs with CPU busy over 90%. |
sparklogs.data.performance.cpu_pct_time_over_70 | float | percent | Percent of the window's sample pairs with CPU busy over 70%. |
sparklogs.data.performance.cpu_busy_pct_avg | float | percent | The host's average CPU busy share over the window, core-normalized. |
sparklogs.data.performance.cpu_busy_pct_max_10s | float | percent | The highest 10-second CPU busy share the window observed, core-normalized: the burst the average hides. |
sparklogs.data.performance.cpu_busy_pct_p90_10s | float | percent | The 90th percentile, across the window's 10-second samples, of the CPU busy share, core-normalized. |
sparklogs.data.performance.cpu_kernel_pct_of_busy_avg | float | percent | The average share of busy CPU time spent in kernel mode. |
sparklogs.data.performance.cpu_user_pct_of_busy_avg | float | percent | The average share of busy CPU time spent in user mode. |
sparklogs.data.performance.cpu_kernel_excl_drivers_pct_of_busy_avg | float | percent | Average share of busy CPU time spent in kernel mode, excluding interrupt and deferred procedure call time. |
sparklogs.data.performance.cpu_interrupt_pct_avg | float | percent | The average share of busy CPU time spent servicing interrupts. |
sparklogs.data.performance.cpu_interrupt_pct_max_10s | float | percent | The highest 10-second interrupt share the window observed. |
sparklogs.data.performance.cpu_dpc_pct_avg | float | percent | The average share of busy CPU time spent in deferred procedure calls. |
sparklogs.data.performance.cpu_dpc_pct_max_10s | float | percent | The highest 10-second DPC share the window observed. |
sparklogs.data.performance.cpu_interrupt_dpc_pct_avg | float | percent | Average share of busy CPU time spent servicing interrupts and deferred procedure calls. |
sparklogs.data.performance.cpu_clock_pct_of_base_avg | float | percent | Average CPU clock as a percentage of rated non-turbo base frequency. Can exceed 100 under turbo. |
sparklogs.data.performance.cpu_base_clock_mhz | integer | The rated non-turbo base clock, in MHz. Absent when the host did not report one. | |
sparklogs.data.performance.cpu_clock_mhz_avg | float | Average effective clock in MHz: rated base frequency multiplied by cpu_clock_pct_of_base_avg / 100. Absent without a base frequency. | |
sparklogs.data.performance.run_queue_p90_10s | float | The 90th percentile, across the window's 10-second samples, of how many threads were ready to run but waiting for a core. On a virtual machine the queue can be threads waiting on host dispatch, which the guest does not count as busy. | |
sparklogs.data.performance.run_queue_per_core_p90_10s | float | run_queue_p90_10s divided by the logical core count. On a VM, queued threads may be waiting for host scheduling without appearing as busy guest CPU. | |
sparklogs.data.performance.logical_core_count | integer | count | How many logical processors the host has, which the per-core queue reading is divided by. |
sparklogs.data.performance.commit_pct | float | percent | The current commit charge as a percentage of the commit limit. |
sparklogs.data.performance.commit_pct_max_window | float | percent | The highest commit charge percentage the window observed. |
sparklogs.data.performance.hard_faults_per_s | float | per_second | Hard page faults per second, latest 10s sample. |
sparklogs.data.performance.ram_pct_time_in_hard_fault_storm | float | percent | Percentage of the window's hard-fault readings above the configured storm-rate threshold. |
sparklogs.data.performance.ram_total_bytes | integer | bytes | Installed physical RAM, the same figure system_info reports, at the end of the window. |
sparklogs.data.performance.ram_available_bytes | integer | bytes | Physical RAM available to new allocations (free, zeroed and standby pages) at the end of the window. |
sparklogs.data.performance.ram_standby_bytes | integer | bytes | RAM holding cached pages the memory manager can repurpose (the standby list, all priorities) at the end of the window. Part of available. |
sparklogs.data.performance.ram_modified_bytes | integer | bytes | RAM holding dirty pages waiting to be written out (the modified list) at the end of the window. |
sparklogs.data.performance.ram_compressed_bytes | integer | bytes | Resident memory held by the memory-compression store, from the latest successful process enumeration (at most one process-table generation old). Zero when compression is off. Absent until the first successful enumeration, or when no process-table topic is collected. |
sparklogs.data.performance.cpu_busy_age_basis | string | onset: witnessed start. observed: already present when first seen, making age a lower bound. unknown_ongoing: no meaningful onset time. | |
sparklogs.data.performance.cpu_busy_age_h | float | hours | How long this condition has been open, in hours. |
sparklogs.data.performance.cpu_interrupt_storm_age_basis | string | onset: witnessed start. observed: already present when first seen, making age a lower bound. unknown_ongoing: no meaningful onset time. | |
sparklogs.data.performance.cpu_interrupt_storm_age_h | float | hours | How long this condition has been open, in hours. |
sparklogs.data.performance.cpu_kernel_dominated_age_basis | string | onset: witnessed start. observed: already present when first seen, making age a lower bound. unknown_ongoing: no meaningful onset time. | |
sparklogs.data.performance.cpu_kernel_dominated_age_h | float | hours | How long this condition has been open, in hours. |
sparklogs.data.performance.cpu_throttled_under_load_age_basis | string | onset: witnessed start. observed: already present when first seen, making age a lower bound. unknown_ongoing: no meaningful onset time. | |
sparklogs.data.performance.cpu_throttled_under_load_age_h | float | hours | How long this condition has been open, in hours. |
sparklogs.data.performance.ram_commit_near_cap_age_basis | string | onset: witnessed start. observed: already present when first seen, making age a lower bound. unknown_ongoing: no meaningful onset time. | |
sparklogs.data.performance.ram_commit_near_cap_age_h | float | hours | How long this condition has been open, in hours. |
sparklogs.data.performance.ram_hard_fault_storm_age_basis | string | onset: witnessed start. observed: already present when first seen, making age a lower bound. unknown_ongoing: no meaningful onset time. | |
sparklogs.data.performance.ram_hard_fault_storm_age_h | float | hours | How long this condition has been open, in hours. |
Conditions
A condition is a state that holds for a while. The agent opens it when the host enters it, keeps it open while it lasts, and closes it when the host comes back out, so one episode answers for the whole stretch instead of one alert per sample.
Example
Inventory (every 5 minutes)
cpu 12% avg, commit 48%, run queue 0.3/core, interrupts 1.2%, hard faults 6/s.
sparklogs.data.performance.cpu_pct_time_over_90: 0.0
sparklogs.data.performance.run_queue_per_core_p90_10s: 0.3
sparklogs.data.performance.cpu_interrupt_dpc_pct_avg: 1.2
sparklogs.data.performance.cpu_busy_pct_avg: 12.0
sparklogs.data.performance.cpu_kernel_excl_drivers_pct_of_busy_avg: 18.0
sparklogs.data.performance.commit_pct: 48.0
sparklogs.data.performance.ram_pct_time_in_hard_fault_storm: 0.0
SparkLogs: CONTEXT, Info, performance: INVENTORY: cpu 12% avg, commit 48%, run queue 0.3/core, interrupts 1.2%, hard faults 6/s.
Selected conditions
cpu_busy
CPU is busy.
Also reported by: Host performance
Example
started; time over 90% CPU 100% (threshold 80%), run queue 5/core (boundary 4/core)
sparklogs.data.performance.cpu_pct_time_over_90: 100.0
sparklogs.data.performance.run_queue_per_core_p90_10s: 5.0
sparklogs.data.performance.cpu_busy_age_h: 0.0
SparkLogs: cpu_busy, Notice, performance: cpu_busy: NOTABLE: started; time over 90% CPU 100% (threshold 80%), run queue 5/core (boundary 4/core)
Example
cleared after 0h51m, peaked Notice, relapses 1; time over 90% CPU 0% (clears at 70%)
sparklogs.data.performance.cpu_pct_time_over_90: 0.0
sparklogs.data.performance.run_queue_per_core_p90_10s: 0.13
SparkLogs: cpu_busy, Info, performance: cpu_busy: RECOVERED: cleared after 0h51m, peaked Notice, relapses 1; time over 90% CPU 0% (clears at 70%)
Example
subsiding, not yet cleared; time over 90% CPU 76.67% (clears at 70%), open for 0h25m
sparklogs.data.performance.cpu_pct_time_over_90: 76.67
sparklogs.data.performance.run_queue_per_core_p90_10s: 0.13
sparklogs.data.performance.cpu_busy_age_h: 0.42
SparkLogs: cpu_busy, Display, performance: cpu_busy: ELEVATED: subsiding, not yet cleared; time over 90% CPU 76.67% (clears at 70%), open for 0h25m
Example
relapsed; time over 90% CPU 100% (threshold 80%), run queue 5/core (boundary 4/core), open for 0h40m, relapses 1
sparklogs.data.performance.cpu_pct_time_over_90: 100.0
sparklogs.data.performance.run_queue_per_core_p90_10s: 5.0
sparklogs.data.performance.cpu_busy_age_h: 0.67
SparkLogs: cpu_busy, Notice, performance: cpu_busy: ELEVATED: relapsed; time over 90% CPU 100% (threshold 80%), run queue 5/core (boundary 4/core), open for 0h40m, relapses 1
| Case | Severity | Ticket class |
|---|---|---|
onset | Trace to Fatal | performance |
held | Trace to Fatal | performance |
recovered | Trace to Fatal | performance |
cpu_interrupt_storm
CPU time is dominated by interrupt and DPC handling.
Also reported by: Host performance
Impact: Device or driver interrupt load can starve ordinary work on the host.
Example
started; interrupt+DPC 34% (threshold 30%)
sparklogs.data.performance.cpu_interrupt_dpc_pct_avg: 34.0
sparklogs.data.performance.cpu_interrupt_storm_age_h: 0.0
SparkLogs: cpu_interrupt_storm, Warning, performance: cpu_interrupt_storm: NOTABLE: started; interrupt+DPC 34% (threshold 30%)
Example
cleared after 0h50m, peaked Warning, relapses 1; interrupt+DPC 2% (clears at 20%)
sparklogs.data.performance.cpu_interrupt_dpc_pct_avg: 2.0
SparkLogs: cpu_interrupt_storm, Info, performance: cpu_interrupt_storm: RECOVERED: cleared after 0h50m, peaked Warning, relapses 1; interrupt+DPC 2% (clears at 20%)
Example
subsiding, not yet cleared; interrupt+DPC 25% (clears at 20%), open for 0h26m
sparklogs.data.performance.cpu_interrupt_dpc_pct_avg: 25.0
sparklogs.data.performance.cpu_interrupt_storm_age_h: 0.43
SparkLogs: cpu_interrupt_storm, Notice, performance: cpu_interrupt_storm: ELEVATED: subsiding, not yet cleared; interrupt+DPC 25% (clears at 20%), open for 0h26m
Example
relapsed; interrupt+DPC 34% (threshold 30%), open for 0h42m, relapses 1
sparklogs.data.performance.cpu_interrupt_dpc_pct_avg: 34.0
sparklogs.data.performance.cpu_interrupt_storm_age_h: 0.7
SparkLogs: cpu_interrupt_storm, Warning, performance: cpu_interrupt_storm: ELEVATED: relapsed; interrupt+DPC 34% (threshold 30%), open for 0h42m, relapses 1
| Case | Severity | Ticket class |
|---|---|---|
onset | Trace to Fatal | performance |
held | Trace to Fatal | performance |
recovered | Trace to Fatal | performance |
cpu_kernel_dominated
Busy CPU time is mostly kernel work, excluding interrupts and deferred procedure calls.
Also reported by: Host performance
Impact: Application throughput can be lower than the busy figure alone suggests.
Example
started; CPU 84% (threshold 70%), kernel 85.26% of busy (threshold 70%)
sparklogs.data.performance.cpu_busy_pct_avg: 84.0
sparklogs.data.performance.cpu_kernel_excl_drivers_pct_of_busy_avg: 85.26
sparklogs.data.performance.cpu_kernel_dominated_age_h: 0.0
SparkLogs: cpu_kernel_dominated, Notice, performance: cpu_kernel_dominated: NOTABLE: started; CPU 84% (threshold 70%), kernel 85.26% of busy (threshold 70%)
Example
cleared after 0h50m, peaked Notice, relapses 1; CPU 40%, kernel 25% of busy (clears below 60% CPU or 60% kernel)
sparklogs.data.performance.cpu_busy_pct_avg: 40.0
sparklogs.data.performance.cpu_kernel_excl_drivers_pct_of_busy_avg: 25.0
SparkLogs: cpu_kernel_dominated, Info, performance: cpu_kernel_dominated: RECOVERED: cleared after 0h50m, peaked Notice, relapses 1; CPU 40%, kernel 25% of busy (clears below 60% CPU or 60% kernel)
Example
subsiding, not yet cleared; CPU 65%, kernel 61.92% of busy (clears below 60% CPU or 60% kernel), open for 0h27m
sparklogs.data.performance.cpu_busy_pct_avg: 65.0
sparklogs.data.performance.cpu_kernel_excl_drivers_pct_of_busy_avg: 61.92
sparklogs.data.performance.cpu_kernel_dominated_age_h: 0.45
SparkLogs: cpu_kernel_dominated, Notice, performance: cpu_kernel_dominated: ELEVATED: subsiding, not yet cleared; CPU 65%, kernel 61.92% of busy (clears below 60% CPU or 60% kernel), open for 0h27m
Example
relapsed; CPU 84% (threshold 70%), kernel 85.62% of busy (threshold 70%), open for 0h41m, relapses 1
sparklogs.data.performance.cpu_busy_pct_avg: 84.0
sparklogs.data.performance.cpu_kernel_excl_drivers_pct_of_busy_avg: 85.62
sparklogs.data.performance.cpu_kernel_dominated_age_h: 0.68
SparkLogs: cpu_kernel_dominated, Notice, performance: cpu_kernel_dominated: ELEVATED: relapsed; CPU 84% (threshold 70%), kernel 85.62% of busy (threshold 70%), open for 0h41m, relapses 1
| Case | Severity | Ticket class |
|---|---|---|
onset | Trace to Fatal | performance |
held | Trace to Fatal | performance |
recovered | Trace to Fatal | performance |
cpu_throttled_under_load
The CPU is running below its rated frequency while under load.
Also reported by: Host performance
Impact: Work takes longer than the hardware would otherwise allow; thermal, power or firmware limits are the usual cause.
Example
started; CPU 82% (threshold 70%), performance 45% (below 50%)
sparklogs.data.performance.cpu_throttled_under_load_age_h: 0.0
SparkLogs: cpu_throttled_under_load, Error, performance: cpu_throttled_under_load: NOTABLE: started; CPU 82% (threshold 70%), performance 45% (below 50%)
Example
cleared after 0h50m, peaked Error, relapses 1; CPU 40%, performance 100% (clears below 60% CPU or at 80% performance)
SparkLogs: cpu_throttled_under_load, Info, performance: cpu_throttled_under_load: RECOVERED: cleared after 0h50m, peaked Error, relapses 1; CPU 40%, performance 100% (clears below 60% CPU or at 80% performance)
Example
subsiding, not yet cleared; CPU 65%, performance 70% (clears below 60% CPU or at 80% performance), open for 0h27m
sparklogs.data.performance.cpu_throttled_under_load_age_h: 0.45
SparkLogs: cpu_throttled_under_load, Notice, performance: cpu_throttled_under_load: ELEVATED: subsiding, not yet cleared; CPU 65%, performance 70% (clears below 60% CPU or at 80% performance), open for 0h27m
Example
relapsed; CPU 82% (threshold 70%), performance 45% (below 50%), open for 0h41m, relapses 1
sparklogs.data.performance.cpu_throttled_under_load_age_h: 0.68
SparkLogs: cpu_throttled_under_load, Warning, performance: cpu_throttled_under_load: ELEVATED: relapsed; CPU 82% (threshold 70%), performance 45% (below 50%), open for 0h41m, relapses 1
| Case | Severity | Ticket class |
|---|---|---|
onset | Trace to Fatal | performance |
held | Trace to Fatal | performance |
recovered | Trace to Fatal | performance |
ram_commit_near_cap
Committed memory is high.
Also reported by: Host performance
Example
started; commit 92% (threshold 90%)
sparklogs.data.performance.commit_pct: 92.0
sparklogs.data.performance.ram_commit_near_cap_age_h: 0.0
SparkLogs: ram_commit_near_cap, Notice, performance: ram_commit_near_cap: NOTABLE: started; commit 92% (threshold 90%)
Example
cleared after 0h44m, peaked Notice, relapses 1; commit 70% (clears at 80%)
sparklogs.data.performance.commit_pct: 70.0
SparkLogs: ram_commit_near_cap, Info, performance: ram_commit_near_cap: RECOVERED: cleared after 0h44m, peaked Notice, relapses 1; commit 70% (clears at 80%)
Example
subsiding, not yet cleared; commit 82% (clears at 80%), open for 0h18m
sparklogs.data.performance.commit_pct: 82.0
sparklogs.data.performance.ram_commit_near_cap_age_h: 0.3
SparkLogs: ram_commit_near_cap, Display, performance: ram_commit_near_cap: ELEVATED: subsiding, not yet cleared; commit 82% (clears at 80%), open for 0h18m
Example
relapsed; commit 92% (threshold 90%), open for 0h39m, relapses 1
sparklogs.data.performance.commit_pct: 92.0
sparklogs.data.performance.ram_commit_near_cap_age_h: 0.65
SparkLogs: ram_commit_near_cap, Notice, performance: ram_commit_near_cap: ELEVATED: relapsed; commit 92% (threshold 90%), open for 0h39m, relapses 1
| Case | Severity | Ticket class |
|---|---|---|
onset | Trace to Fatal | performance |
held | Trace to Fatal | performance |
recovered | Trace to Fatal | performance |
ram_hard_fault_storm
High hard-page-fault activity persists.
Also reported by: Host performance
Impact: Memory-related disk reads may slow workloads.
Example
started; commit 97% (threshold 95%), hard-fault duty 100% (threshold 50%)
sparklogs.data.performance.commit_pct: 97.0
sparklogs.data.performance.ram_pct_time_in_hard_fault_storm: 100.0
sparklogs.data.performance.ram_hard_fault_storm_age_h: 0.0
SparkLogs: ram_hard_fault_storm, Serious, performance: ram_hard_fault_storm: NOTABLE: started; commit 97% (threshold 95%), hard-fault duty 100% (threshold 50%)
Example
cleared after 0h49m, peaked Serious, relapses 1; commit 70%, hard-fault duty 0% (clears below 85% commit or 40% duty)
sparklogs.data.performance.commit_pct: 70.0
sparklogs.data.performance.ram_pct_time_in_hard_fault_storm: 0.0
SparkLogs: ram_hard_fault_storm, Info, performance: ram_hard_fault_storm: RECOVERED: cleared after 0h49m, peaked Serious, relapses 1; commit 70%, hard-fault duty 0% (clears below 85% commit or 40% duty)
Example
subsiding, not yet cleared; commit 90%, hard-fault duty 45.16% (clears below 85% commit or 40% duty), open for 0h23m
sparklogs.data.performance.commit_pct: 90.0
sparklogs.data.performance.ram_pct_time_in_hard_fault_storm: 45.16
sparklogs.data.performance.ram_hard_fault_storm_age_h: 0.38
SparkLogs: ram_hard_fault_storm, Notice, performance: ram_hard_fault_storm: ELEVATED: subsiding, not yet cleared; commit 90%, hard-fault duty 45.16% (clears below 85% commit or 40% duty), open for 0h23m
Example
relapsed; commit 97% (threshold 95%), hard-fault duty 100% (threshold 50%), open for 0h40m, relapses 1
sparklogs.data.performance.commit_pct: 97.0
sparklogs.data.performance.ram_pct_time_in_hard_fault_storm: 100.0
sparklogs.data.performance.ram_hard_fault_storm_age_h: 0.67
SparkLogs: ram_hard_fault_storm, Serious, performance: ram_hard_fault_storm: ELEVATED: relapsed; commit 97% (threshold 95%), hard-fault duty 100% (threshold 50%), open for 0h40m, relapses 1
| Case | Severity | Ticket class |
|---|---|---|
onset | Trace to Fatal | performance |
held | Trace to Fatal | performance |
recovered | Trace to Fatal | performance |