Agent overhead
What the SparkLogs agent stack itself costs the host: CPU, working set and handles.
Topic id: agent_overhead.
Reported every 15 minutes on clock boundaries for the window ending at t. sparklogs.window_coverage_pct is below 100 when collection covered only part of the window. Use chart buckets at least 15 minutes wide.
Fields
CPU, memory and handle readings for the agent and its event-processing companion.
| Field | Type | Unit | Meaning |
|---|---|---|---|
sparklogs.data.agent_overhead.row | string | Which component this row is: agent for the agent process itself, vector for its Vector child. | |
sparklogs.data.agent_overhead.agent_version | string | The agent version, plain semver, without build metadata. | |
sparklogs.data.agent_overhead.agent_build | string | The agent build: semver plus the commit this binary was built from. .dirty means the tree was modified. | |
sparklogs.data.agent_overhead.pid | integer | This component's process ID. Combine with create_time_ts to tell one run of it from the next. | |
sparklogs.data.agent_overhead.create_time_ts | string | timestamp | When this component's process started, as RFC3339 UTC with a Z suffix, converted from the Windows FILETIME. |
sparklogs.data.agent_overhead.service_name | string | The Windows service this component runs as: SparkLogsAgent or SparkLogsVector on the default install, the instance-suffixed name on a side-by-side instance. | |
sparklogs.data.agent_overhead.instance_name | string | The side-by-side instance this install is. Absent on the default install. | |
sparklogs.data.agent_overhead.cpu_cycles_delta | integer | CPU cycles between the first and last readings retained in the window. Absent without two readings of the same process identity. | |
sparklogs.data.agent_overhead.user_time_delta_ms | float | milliseconds | User-mode CPU time between the first and last readings retained in the window. Absent without two readings of the same process identity. |
sparklogs.data.agent_overhead.kernel_time_delta_ms | float | milliseconds | Kernel-mode CPU time between the first and last readings retained in the window. Absent without two readings of the same process identity. |
sparklogs.data.agent_overhead.cpu_pct_of_one_core_avg | float | percent | Mean of CPU percentages measured between consecutive captures, relative to one logical core. Each percentage uses the measured elapsed time. |
sparklogs.data.agent_overhead.cpu_pct_of_one_core_p95 | float | percent | 95th percentile of this process's CPU percentages between captures, relative to one logical core. |
sparklogs.data.agent_overhead.working_set_avg_bytes | integer | bytes | This process's mean resident working set over the window. |
sparklogs.data.agent_overhead.working_set_peak_bytes | integer | bytes | The highest working set this process reached in the window. |
sparklogs.data.agent_overhead.private_bytes | integer | bytes | This process's private (non-shared) memory at the most recent reading in the window. |
sparklogs.data.agent_overhead.handle_count_peak | integer | count | The highest handle count this process reached in the window. |
sparklogs.data.agent_overhead.handle_count_avg | float | count | Mean handle count over the window, used to assess the handle budget. |
sparklogs.data.agent_overhead.io_read_bytes_delta | integer | bytes | Bytes read between the first and last readings retained in the window. Absent without two readings of the same process identity. |
sparklogs.data.agent_overhead.io_write_bytes_delta | integer | bytes | Bytes written between the first and last readings retained in the window. Absent without two readings of the same process identity. |
sparklogs.data.agent_overhead.io_other_bytes_delta | integer | bytes | Bytes of I/O other than reads or writes between the first and last readings retained in the window. Absent without two readings of the same process identity. |
sparklogs.data.agent_overhead.restart_count_1h | integer | count | How many times this row's own process restarted in the last hour. |
sparklogs.data.agent_overhead.uptime_s | integer | seconds | How long this process has been running. |
sparklogs.data.agent_overhead.combined_working_set_avg_mb | float | megabytes | Sum of the agent and Vector mean working sets in MiB (1,048,576 bytes), reported on the agent row and used to assess the RAM budget. |
sparklogs.data.agent_overhead.private_bytes_monotonic_windows | integer | count | How many consecutive windows this process's private memory has risen without a drop. |
sparklogs.data.agent_overhead.private_bytes_slope_mb_per_h | float | megabytes_per_hour | This process's private-memory growth rate in MiB (1,048,576 bytes) per hour, fitted from its recent readings. Present once two readings at different times exist, whether memory rose or not. |
sparklogs.data.agent_overhead.handle_monotonic_windows | integer | count | How many consecutive windows this process's handle count has risen without a drop. |
sparklogs.data.agent_overhead.sparklogs_agent_cpu_over_budget_age_basis | string | onset: witnessed start. observed: already present when first seen, making age a lower bound. unknown_ongoing: no meaningful onset time. | |
sparklogs.data.agent_overhead.sparklogs_agent_cpu_over_budget_age_h | float | hours | How long this condition has been open, in hours. |
sparklogs.data.agent_overhead.sparklogs_agent_handle_over_budget_age_basis | string | onset: witnessed start. observed: already present when first seen, making age a lower bound. unknown_ongoing: no meaningful onset time. | |
sparklogs.data.agent_overhead.sparklogs_agent_handle_over_budget_age_h | float | hours | How long this condition has been open, in hours. |
sparklogs.data.agent_overhead.sparklogs_agent_memory_over_budget_age_basis | string | onset: witnessed start. observed: already present when first seen, making age a lower bound. unknown_ongoing: no meaningful onset time. | |
sparklogs.data.agent_overhead.sparklogs_agent_memory_over_budget_age_h | float | hours | How long this condition has been open, in hours. |
Conditions
A condition is a state that holds for a while. The agent opens it when the host enters it, keeps it open while it lasts, and closes it when the host comes back out, so one episode answers for the whole stretch instead of one alert per sample.
Example
Inventory (every 15 minutes)
agent+Vector WS 96.0 MB; component "agent" cpu 1.40% of one core, WS 62.0 MB, 1800 handles; component "vector" cpu 2.60% of one core, WS 34.0 MB, 640 handles.
sparklogs.data.agent_overhead.row: agent
sparklogs.data.agent_overhead.cpu_pct_of_one_core_avg: 1.4
sparklogs.data.agent_overhead.handle_count_avg: 1800.0
sparklogs.data.agent_overhead.combined_working_set_avg_mb: 96.0
SparkLogs: CONTEXT, Info, agent_overhead: INVENTORY: agent+Vector WS 96.0 MB; component "agent" cpu 1.40% of one core, WS 62.0 MB, 1800 handles; component "vector" cpu 2.60% of one core, WS 34.0 MB, 640 handles.
Selected conditions
sparklogs_agent_cpu_over_budget
SparkLogs Agent CPU usage exceeds its budget.
Also reported by: Agent overhead
Impact: Monitoring overhead may be higher than expected on this host.
Example
started; component "agent" agent CPU of one core 145% (threshold 20%)
sparklogs.instance: component:agent
sparklogs.data.agent_overhead.row: agent
sparklogs.data.agent_overhead.cpu_pct_of_one_core_avg: 145.0
sparklogs.data.agent_overhead.sparklogs_agent_cpu_over_budget_age_h: 0.0
SparkLogs: sparklogs_agent_cpu_over_budget, Notice, agent_overhead: sparklogs_agent_cpu_over_budget: NOTABLE: started; component "agent" agent CPU of one core 145% (threshold 20%)
Example
cleared after 0h26m, peaked Notice, relapses 1; component "agent" agent CPU of one core 4% (clears at 16%)
sparklogs.instance: component:agent
sparklogs.data.agent_overhead.row: agent
sparklogs.data.agent_overhead.cpu_pct_of_one_core_avg: 4.0
SparkLogs: sparklogs_agent_cpu_over_budget, Info, agent_overhead: sparklogs_agent_cpu_over_budget: RECOVERED: cleared after 0h26m, peaked Notice, relapses 1; component "agent" agent CPU of one core 4% (clears at 16%)
Example
subsiding, not yet cleared; component "agent" agent CPU of one core 9% (clears at 16%), open for 0h26m
sparklogs.instance: component:agent
sparklogs.data.agent_overhead.row: agent
sparklogs.data.agent_overhead.cpu_pct_of_one_core_avg: 9.0
sparklogs.data.agent_overhead.sparklogs_agent_cpu_over_budget_age_h: 0.43
SparkLogs: sparklogs_agent_cpu_over_budget, Notice, agent_overhead: sparklogs_agent_cpu_over_budget: ELEVATED: subsiding, not yet cleared; component "agent" agent CPU of one core 9% (clears at 16%), open for 0h26m
Example
relapsed; component "agent" agent CPU of one core 59.13% (threshold 20%), open for 0h29m, relapses 1
sparklogs.instance: component:agent
sparklogs.data.agent_overhead.row: agent
sparklogs.data.agent_overhead.cpu_pct_of_one_core_avg: 59.13
sparklogs.data.agent_overhead.sparklogs_agent_cpu_over_budget_age_h: 0.48
SparkLogs: sparklogs_agent_cpu_over_budget, Notice, agent_overhead: sparklogs_agent_cpu_over_budget: ELEVATED: relapsed; component "agent" agent CPU of one core 59.13% (threshold 20%), open for 0h29m, relapses 1
| Case | Severity | Ticket class |
|---|---|---|
onset | Trace to Fatal | rmm |
held | Trace to Fatal | rmm |
recovered | Trace to Fatal | rmm |
sparklogs_agent_handle_over_budget
SparkLogs Agent handle usage exceeds its budget.
Also reported by: Agent overhead
Impact: Monitoring overhead may be higher than expected on this host.
Example
started; component "agent" agent handles 4687.5 (threshold 4000)
sparklogs.instance: component:agent
sparklogs.data.agent_overhead.row: agent
sparklogs.data.agent_overhead.handle_count_avg: 4687.5
sparklogs.data.agent_overhead.sparklogs_agent_handle_over_budget_age_h: 0.0
SparkLogs: sparklogs_agent_handle_over_budget, Notice, agent_overhead: sparklogs_agent_handle_over_budget: NOTABLE: started; component "agent" agent handles 4687.5 (threshold 4000)
Example
cleared after 0h15m, peaked Notice, relapses 1; component "agent" agent handles 1800 (clears at 3200)
sparklogs.instance: component:agent
sparklogs.data.agent_overhead.row: agent
sparklogs.data.agent_overhead.handle_count_avg: 1800.0
SparkLogs: sparklogs_agent_handle_over_budget, Info, agent_overhead: sparklogs_agent_handle_over_budget: RECOVERED: cleared after 0h15m, peaked Notice, relapses 1; component "agent" agent handles 1800 (clears at 3200)
Example
subsiding, not yet cleared; component "agent" agent handles 3500 (clears at 3200), open for 0h17m
sparklogs.instance: component:agent
sparklogs.data.agent_overhead.row: agent
sparklogs.data.agent_overhead.handle_count_avg: 3500.0
sparklogs.data.agent_overhead.sparklogs_agent_handle_over_budget_age_h: 0.28
SparkLogs: sparklogs_agent_handle_over_budget, Notice, agent_overhead: sparklogs_agent_handle_over_budget: ELEVATED: subsiding, not yet cleared; component "agent" agent handles 3500 (clears at 3200), open for 0h17m
Example
relapsed; component "agent" agent handles 4437.5 (threshold 4000), open for 0h33m, relapses 1
sparklogs.instance: component:agent
sparklogs.data.agent_overhead.row: agent
sparklogs.data.agent_overhead.handle_count_avg: 4437.5
sparklogs.data.agent_overhead.sparklogs_agent_handle_over_budget_age_h: 0.55
SparkLogs: sparklogs_agent_handle_over_budget, Notice, agent_overhead: sparklogs_agent_handle_over_budget: ELEVATED: relapsed; component "agent" agent handles 4437.5 (threshold 4000), open for 0h33m, relapses 1
| Case | Severity | Ticket class |
|---|---|---|
onset | Trace to Fatal | rmm |
held | Trace to Fatal | rmm |
recovered | Trace to Fatal | rmm |
sparklogs_agent_memory_over_budget
Combined SparkLogs Agent and Vector memory usage exceeds its budget.
Also reported by: Agent overhead
Impact: Monitoring overhead may be higher than expected on this host.
Example
started; component "agent" agent and Vector working set 382 MB (threshold 350 MB)
sparklogs.instance: component:agent
sparklogs.data.agent_overhead.row: agent
sparklogs.data.agent_overhead.combined_working_set_avg_mb: 382.01
sparklogs.data.agent_overhead.sparklogs_agent_memory_over_budget_age_h: 0.0
SparkLogs: sparklogs_agent_memory_over_budget, Notice, agent_overhead: sparklogs_agent_memory_over_budget: NOTABLE: started; component "agent" agent and Vector working set 382 MB (threshold 350 MB)
Example
cleared after 0h14m, peaked Notice, relapses 1; component "agent" agent and Vector working set 220 MB (clears at 280 MB)
sparklogs.instance: component:agent
sparklogs.data.agent_overhead.row: agent
sparklogs.data.agent_overhead.combined_working_set_avg_mb: 219.5
SparkLogs: sparklogs_agent_memory_over_budget, Info, agent_overhead: sparklogs_agent_memory_over_budget: RECOVERED: cleared after 0h14m, peaked Notice, relapses 1; component "agent" agent and Vector working set 220 MB (clears at 280 MB)
Example
subsiding, not yet cleared; component "agent" agent and Vector working set 300 MB (clears at 280 MB), open for 0h15m
sparklogs.instance: component:agent
sparklogs.data.agent_overhead.row: agent
sparklogs.data.agent_overhead.combined_working_set_avg_mb: 300.0
sparklogs.data.agent_overhead.sparklogs_agent_memory_over_budget_age_h: 0.25
SparkLogs: sparklogs_agent_memory_over_budget, Notice, agent_overhead: sparklogs_agent_memory_over_budget: ELEVATED: subsiding, not yet cleared; component "agent" agent and Vector working set 300 MB (clears at 280 MB), open for 0h15m
Example
relapsed; component "agent" agent and Vector working set 366 MB (threshold 350 MB), open for 0h32m, relapses 1
sparklogs.instance: component:agent
sparklogs.data.agent_overhead.row: agent
sparklogs.data.agent_overhead.combined_working_set_avg_mb: 366.25
sparklogs.data.agent_overhead.sparklogs_agent_memory_over_budget_age_h: 0.53
SparkLogs: sparklogs_agent_memory_over_budget, Notice, agent_overhead: sparklogs_agent_memory_over_budget: ELEVATED: relapsed; component "agent" agent and Vector working set 366 MB (threshold 350 MB), open for 0h32m, relapses 1
| Case | Severity | Ticket class |
|---|---|---|
onset | Trace to Fatal | rmm |
held | Trace to Fatal | rmm |
recovered | Trace to Fatal | rmm |