Tags

, , , ,

Observability is usually looked at from the perspective of applications and their needs and the cloud native platforms. Nothing wrong with that, but it is also a rather narrow view of the telemetry we are already collecting or are very capable of collecting. Our observability agents (Fluent Bit, Fluentd, Elastic Agent …) can collect signals (logs, metrics and traces) that can contribute to the security operations within an enterprise. More importantly, the same open-source tooling can help answer a deceptively simple question such as has someone changed the system, or changed the tooling that tells us what the system is doing?

When it comes to security, that second part matters. An attacker does not always need to exploit a new vulnerability. If they can alter a service definition, scheduled task, access-control file, logging configuration or telemetry destination, they can prevent us from sensing what damage they maybe doing, including opening up simpler or more invasive routes into your systems. MITRE ATT&CK describes this kind of action as a defence-evasion technique. Research from Palo Alto Networks Unit 42 has also described cloud logging services as an attack surface, including changes that stop, redirect or selectively reduce audit collection. Not only that getting at logs, means they can potentially determine when systems that are running, and find new attack vectors.

As you can see being able to detect configuration and file manipulation is an important factor for security. Many may react to this with concern about introducing cost for expensive security tools, but if there are simple means to listen to OS activities that relate to file manipulation then we can leverage our observability estate for this aspect of security.

The wider open-source stack

That neutrality for our observability agents is valuable for security. A host or cluster can emit operational and security telemetry through the same managed pipeline, while routing data to one or more destinations such as OpenSearch, Loki, ClickHouse, a SIEM, object storage or an OTLP-compatible service. Not to mention if the event is for an application configuration, we can evaluate it against application and service viewpoints such as managing configuration drift.

The roles are complementary:

  • AIDE is a open-source Linux utility that provides the file-integrity check. It builds a known-good database and reports additions, removals and changes to selected files and attributes. Aide achieves this through the use of the iNodes API, which if you preferred could also be tapped into even more directly.
  • Fluent Bit or other OTel compliant tools collect the result close to the workload, adds context, buffers during a network interruption and routes the event. Our backend services which are in all likelihood OTLP capable, means the events can be send nearly anywhere for analysis and alerting
  • The analytical backend correlates the integrity event with logins, process starts, deployments, traces and other activity, then drives an alert or investigation.

AIDE is deliberately focused. It does not need to become a SIEM and Fluent Bit does not need to become a file-integrity engine. Combining small tools through open interfaces lets each component do the job it is good at.

We can use Aide not only to monitor our configurations, but when the deployed binaries are disturbed. We can extend this setup further. By exploiting custom commands with OpAMP (particularly for a supervisor) can be used to not only manage the observability agent, but also manage the deployment of configuration to tools like Aide. Now if we think something is being abused, the option to change our file integrity monitoring is easily in our reach.

Why file integrity belongs in the observability conversation

The available incident statistics do not normally include a neat category called “configuration-file manipulation”. The behaviour is spread across techniques such as creating or modifying system processes, scheduled tasks, startup mechanisms, accounts, policies and logging controls. That makes a single adoption or attack percentage difficult to defend. It does not make the risk rare.

For example, Google Cloud’s M-Trends 2025 reporting included techniques such as Create or Modify System Process (20.6%), Account Manipulation (19.9%), Scheduled Task/Job (14.3%), Registry Run Keys/Startup Folder (11.5%) and Web Shell (7.2%) among observed investigations. These are broader ATT&CK behaviours, not a count of configuration-file changes, but many are enabled or persisted through changes to system configuration.

There is also a compliance driver. PCI DSS v4.0.1 Requirement 11.5.2 explicitly requires a change-detection mechanism, such as file-integrity monitoring, for critical files, with comparisons at least weekly and alerts for unauthorised modification. NIST SP 800-53 Revision 5 control SI-7 requires monitoring for unauthorised changes to software, firmware and information. Other standards may describe the required outcome as configuration control, integrity, audit protection or system monitoring rather than naming a particular product.

FIM helps provide the evidence behind those outcomes:

  • what changed;
  • which host or workload reported it;
  • when the change was detected;
  • which attributes or content hashes changed;
  • whether a change matched an approved deployment window;
  • whether the alert was acknowledged and investigated; and
  • whether the integrity baseline was updated through a controlled process.

The last point is easily overlooked. Automatically accepting every change into the baseline produces a very tidy report and a very weak control.

What should be monitored?

Starting with every file on every host usually creates cost and noise. I would start with the files that can change control flow, identity, collection or evidence retention:

  • Fluent Bit configuration, parser files and included fragments;
  • Vector configuration and transformation definitions;
  • OpenTelemetry Collector configuration;
  • systemd units and drop-ins for telemetry agents;
  • log rotation and retention configuration;
  • TLS trust stores, client certificates and references to credentials;
  • auditd, syslog and journald configuration;
  • scheduled tasks, startup scripts and authorised SSH keys;
  • alert rules, dashboards and detection content managed as files; and
  • the AIDE configuration and baseline database itself.

Runtime state, offset databases, buffers and normal log files should normally be excluded. They are expected to change and monitoring them obscures the files that should remain stable.

There is an architectural wrinkle: if AIDE reports through Fluent Bit, and an attacker changes both AIDE and Fluent Bit, the local signal can be suppressed. The practical response is defence in depth. Restrict write access, ship events off-host quickly, monitor the agent configuration and service unit, protect or sign the AIDE baseline, and alert when expected heartbeats or scheduled check results disappear. Silence is also a signal.

Example: AIDE reporting through Fluent Bit and OTLP

The following is a starting point rather than a universal production configuration. Package paths and AIDE defaults differ between Linux distributions, so it should be merged with the supplied configuration rather than replacing it blindly.

First, add an AIDE rule group for the observability configuration. For example, in /etc/aide/aide.conf.d/observability.conf:

# Permissions, inode, links, ownership, size, timestamps and SHA-256 content.
OBS_CONFIG = p+i+n+u+g+s+m+c+sha256
​
/etc/fluent-bit                         OBS_CONFIG
/etc/vector                             OBS_CONFIG
/etc/otelcol-contrib                   OBS_CONFIG
/etc/systemd/system/fluent-bit.service.d OBS_CONFIG
/etc/systemd/system/vector.service.d     OBS_CONFIG
/etc/systemd/system/otelcol.service.d   OBS_CONFIG
/etc/aide                               OBS_CONFIG
​
# These are runtime data and should not be treated as static configuration.
!/var/lib/fluent-bit
!/var/lib/vector
!/var/lib/otelcol
!/var/log

The main AIDE configuration can direct its report to a file that Fluent Bit is permitted to read:

database=file:/var/lib/aide/aide.db.gz
database_out=file:/var/lib/aide/aide.db.new.gz
gzip_dbout=true
​
report_url=file:/var/log/aide/aide.log
report_level=changed_attributes

Initialise and approve the baseline using the procedure provided by the operating-system package. The baseline database should be writable only by the account performing the controlled update; consider retaining a signed or read-only copy outside the host.

Run aide --check from a systemd timer or the organisation’s scheduler. An hourly check provides better detection latency than the weekly minimum stated by PCI DSS, although the right frequency depends on estate size, I/O cost and risk. The scheduler should also emit a success or failure event so that a missing check can be detected.

Fluent Bit can then tail the report, add security context and forward the records over OTLP. The Tail input uses a database to retain offsets, and the OpenTelemetry output supports OTLP/HTTP and OTLP/gRPC.

service:
flush: 5
log_level: info
storage.path: /var/lib/fluent-bit/storage
pipeline:
inputs:
- name: tail
tag: security.fim.aide
path: /var/log/aide/aide.log
key: body
path_key: log.file.path
db: /var/lib/fluent-bit/aide-tail.db
db.sync: normal
read_from_head: false
refresh_interval: 5
rotate_wait: 30
skip_empty_lines: true
skip_long_lines: true
storage.type: filesystem
filters:
- name: modify
match: security.fim.aide
Add:
- event.domain file
- event.category file
- event.type change
- security.control file_integrity
- service.name aide
outputs:
- name: opentelemetry
match: security.fim.aide
host: otel-collector.internal
port: 4318
logs_uri: /v1/logs
tls: on
tls.verify: on
compress: gzip

This example emits each AIDE report line as a log record. That is simple and robust, but the backend will need to group lines belonging to the same check. For a more mature deployment I would place a small wrapper around the AIDE invocation that converts each detected change into JSON Lines and adds a check identifier, hostname, baseline version and execution result. Fluent Bit can then use its JSON parser and every changed file becomes a structured event.

Vector can implement the same pattern with its file source and an OTLP-capable destination or intermediate Collector. The OpenTelemetry Collector can also collect files using the contributed filelog receiver. The choice is less important than applying consistent fields, reliable buffering, secure transport and an independently managed destination.

From event collection to a security control

Collecting the AIDE output is only the first half of the solution. To turn it into a control, the backend needs useful detection logic and operational ownership.

I would start with alerts for:

  • a change to an observability agent destination, filter or exclusion rule;
  • a change to a telemetry service unit or executable;
  • deletion or replacement of the AIDE database;
  • changes outside an approved deployment window;
  • a burst of changes across several hosts;
  • a configuration change followed by a fall in log volume; and
  • a missing AIDE run, agent heartbeat or expected integrity summary.

Correlating the FIM event with deployment records is important. A checksum change shortly after an approved GitOps rollout is very different from the same change following an interactive root login at 02:00. Observability provides the surrounding evidence that allows the integrity event to be assessed rather than simply counted.

There is also value in feeding an integrity event back into operational telemetry. A changed collector configuration can annotate dashboards, traces and incident timelines. The question changes from “why did the logs stop?” to “did the logs stop when this configuration changed?” That is a much better starting point for both an SRE and a security analyst.

Final thoughts

The observability and security domains have spent years building separate pipelines for data that frequently starts on the same machine. Open standards and focused open-source tools give us an opportunity to reduce that duplication.

AIDE supplies a mature integrity signal. Fluent Bit and Vector make that signal transportable and enrichable. OpenTelemetry makes the route to the wider platform portable. None of these components replaces access control, endpoint detection or a SIEM, but together they add a useful security capability to infrastructure many organisations already operate.

Perhaps the most important configuration to observe is the configuration that decides what we are allowed to observe.

Resources

File Integrity Management using Aide and Fluent Bit