I've been self-hosting for about six years with very few problems, but over the past couple of months several unrelated things have started failing in small, frustrating ways. Certificate renewal stopped working silently, an unchanged Docker workload began getting killed for running out of memory, and a cron job that had logged successfully since 2021 stopped producing logs without reporting an error.
A few friends are seeing similar issues on completely different setups, including intermittent DKIM failures and WireGuard tunnels dropping after kernel updates. None of us made any major intentional changes. I use BeAdmin for management, but it has otherwise been stable, so there doesn't seem to be one shared platform involved.
Does this sound like system drift caused by recent kernel or package updates, or is it more likely a combination of aging hardware, configuration rot, and coincidence?
4 Answers
DKIM is not necessarily evidence of a kernel problem. Sending services, selectors, SPF includes, subdomains, and DMARC alignment can drift when someone adds an email provider or makes a temporary configuration change. Periodically verify the published records and selectors against the systems that are actually sending mail. For the host issues, I’d still inspect package histories, hardware health, memory errors, disk space, and the vendor’s kernel configuration before blaming upstream kernel changes.
The cron logging issue could be caused by log rotation or a size and file-count limit rather than cron itself. Some logging setups stop behaving once a directory accumulates too many files or reaches an unexpected limit. Check the rotation rules, disk usage, inode usage, and whether the job is still running but writing somewhere else.
I’d compare kernel and package versions from before the problems started instead of treating every failure as unrelated. Several different issues appearing in the same time window makes the package update history worth investigating, even if no one changed anything manually. Automatic updates are still changes.
The bigger pattern is that everything failed silently. Add external checks based on outcomes rather than just process status: alert before certificates expire, use a heartbeat or dead-man’s switch for cron, monitor container restart and OOM counts, and test the actual WireGuard tunnel. That makes drift visible soon after it happens.

Related Questions
Can't Load PhpMyadmin On After Server Update
Redirect www to non-www in Apache Conf
How To Check If Your SSL Cert Is SHA 1
Windows TrackPad Gestures