Feed health
0readings
0conditions
0themes fed
Livestatus
Per-feed health for the inputs the agent is shipping.
Topic id: feed_health.
Full inventory every 2 hours. Supported changes are reported when the agent observes them.
Fields
Collection and delivery status for each Data Feed.
| Field | Type | Unit | Meaning |
|---|---|---|---|
sparklogs.data.feed_health.module | string | Which data-feed module this row reports on (the pack module id). | |
sparklogs.data.feed_health.health | string | The debounced health verdict: on collector_health, fresh, stale (also a topic that has not reported at all within three of its intervals), stalled, frozen, waiting_prereq or unknown (the topic has ticked this process but produced no valid generation yet); on feed_health, onboarding, onboarding_stuck, current, behind, stuck, blocked or unknown. Until the topic or module reports in this process the row carries the verdict it last reported; absent when it has never reported one. | |
sparklogs.data.feed_health.reason | string | Why health is not the healthy value, from a closed vocabulary specific to the topic (for example capture_failed, not_declared, or missing_required_channel). Absent when health needs no explanation. | |
sparklogs.data.feed_health.since | string | timestamp | When the current health verdict began. Carried from the last report until the topic or module reports in this process, then reset, because the debounce that tracks it is in-memory. |
sparklogs.data.feed_health.lag_value | integer | How far this module is behind the head, in whatever lag_unit names. | |
sparklogs.data.feed_health.lag_unit | string | What lag_value counts: records (an estimate that can undershoot) or bytes (an exact measure). | |
sparklogs.data.feed_health.lag_is_upper_bound | bool | true when lag_value counts records the collector asked Windows to drop before delivery (suppressed event ids) as behind, so it overstates how far this Windows Event Log module is behind; the record estimate can also undershoot where record ids have holes. Present only when true. | |
sparklogs.data.feed_health.data_skips | integer | count | How many spans of permanently lost events this module has recorded: a deliberate discard to escape a poisoned resume position, never an ordinary delay. |
sparklogs.data.feed_health.withholding | object_array | The bound channels holding this module back, worst first, each with the channel name, why (skipped, unavailable, never_drained or no_record), since, and where known the channel_type and the last win32 last_error. Absent when the module is withholding on nothing. | |
sparklogs.data.feed_health.withholding_omitted | integer | count | How many withheld channels did not fit the list. Absent when the list is complete. |
sparklogs.data.feed_health.files_discovered | integer | count | How many files this module's file source has discovered. Present for file sources only. |
sparklogs.data.feed_health.files_unreadable | integer | count | How many of the discovered files could not be read. Present for file sources only. |
sparklogs.data.feed_health.error_kind | string | Why the files of a blocked file source could not be read: permission_denied (the agent's account may not read them; reason access_denied), sharing_violation (another process has held a file open without sharing read access for at least two samples; reason unreadable) or other (see win32_error; reason unreadable). Present only while the row reads blocked. | |
sparklogs.data.feed_health.win32_error | integer | The Windows error code the read failed with, for the file behind a blocked verdict (5 is access denied). Present only while the row reads blocked and Windows gave a code. | |
sparklogs.data.feed_health.polls_at_head | integer | count | How many of the collector's polls of this Windows Event Log module's channels ended with the channel at the head over the agent's last liveness sample interval (heartbeat). Present only while the row reads behind or stuck, from the second sample of a collector run on; beside polls_on_budget it says whether the reader kept reaching the head or kept stopping on its pull budget. |
sparklogs.data.feed_health.polls_on_budget | integer | count | How many of the collector's polls of this Windows Event Log module's channels stopped on the pull budget with events still queued over the agent's last liveness sample interval (heartbeat). Present only while the row reads behind or stuck, from the second sample of a collector run on. |
sparklogs.data.feed_health.records_read_per_min | integer | How many records per minute the collector read from this Windows Event Log module's channels, measured between its last two status reports. Suppressed records are never read and are not counted. Present only while the row reads behind or stuck, once a collector run has reported twice, and absent while the collector has not rewritten its status for three write intervals (a stalled collector reports no rate rather than its last one); a behind row reading at a high rate is catching up, one reading near zero is not. | |
sparklogs.data.feed_health.delta_rate_limited | bool | true while this row's transitions are capped by the daily rate limit: its reported health is held at the last one reported until the window drains. false otherwise. Engaging and releasing the cap are each reported once. Absent only on a row with no verdict yet this process, which carries the last reported value. |
This topic reports what the host looks like rather than scoring a condition on it. The values it carries are queryable on their own.
Example
Inventory (every 2 hours)
4 feeds, 2 current, 1 behind 1h00m (collector_down) "win.eventlog.security", 1 blocked 0h45m (channel_unavailable) "win.eventlog.storage".
sparklogs.data.feed_health.module: win.eventlog.storage
sparklogs.data.feed_health.delta_rate_limited: false
sparklogs.data.feed_health.health: blocked
sparklogs.data.feed_health.reason: channel_unavailable
SparkLogs: CONTEXT, Info, feed_health: INVENTORY: 4 feeds, 2 current, 1 behind 1h00m (collector_down) "win.eventlog.security", 1 blocked 0h45m (channel_unavailable) "win.eventlog.storage".