Skip to content

periphery: ignore host ZFS ARC when /proc/meminfo comes from lxcfs - #1635

Open
phedoreanu wants to merge 1 commit into
moghtech:mainfrom
phedoreanu:fix-lxc-zfs-arc-memory
Open

phedoreanu wants to merge 1 commit into
moghtech:mainfrom
phedoreanu:fix-lxc-zfs-arc-memory

Conversation

@phedoreanu

@phedoreanu phedoreanu commented Sep 16, 2026 •

Copy link
Copy Markdown

Problem

Since #1489 (shipped in 2.3.0), Periphery subtracts the ZFS ARC size from used memory. Inside an LXC container on a ZFS Proxmox host that goes wrong: lxcfs serves a cgroup-scoped /proc/meminfo, but /proc/spl/kstat/zfs/arcstats is not virtualized and still shows the host's ARC. The subtraction saturates to zero, so every LXC guest smaller than the host's ARC reports 0.00 GB used.

This is the mechanism described in #1139 (comment). On my fleet, four LXC containers (1 to 2 GiB each) on a host with a 6.3 GiB ARC have reported mem_used_gb: 0 in every stats record since they picked up 2.3.x:

$ awk '$1=="size"{print $3/2^30 " GiB"}' /proc/spl/kstat/zfs/arcstats   # inside a 1.5 GiB CT
6.27 GiB
$ grep ' /proc/meminfo ' /proc/self/mountinfo
1206 1149 0:61 /proc/meminfo /proc/meminfo rw,nosuid,nodev,relatime shared:593 master:81 - fuse.lxcfs lxcfs rw,user_id=0,group_id=0,allow_other

Fix

Check /proc/self/mountinfo for a FUSE filesystem mounted on /proc/meminfo. When that is the case the meminfo numbers are container-scoped and the ARC cannot be part of them, so the ARC is treated as zero (both for used and for the reported mem_zfs_arc_gb).

I preferred this over the arc >= total clamp suggested in the issue because the clamp still subtracts the full host ARC from any container whose limit is larger than the ARC (for example an 8 GiB container on a host with a 6 GiB ARC), which underreports instead of zeroing. Bare-metal hosts and Periphery running in Docker directly on a ZFS host see the real /proc/meminfo and keep the ARC handling from #1489 unchanged.

The second scope mismatch from that comment (Periphery in Docker inside an LXC seeing its own cgroup) is a separate problem and not addressed here.

Testing

  • Unit tests for the mountinfo parsing (lxcfs line vs. plain procfs) and for parse_zfs_arc_size.
  • cargo test -p komodo_periphery stats::mem and cargo fmt --all -- --check pass.

Refs #1139, #125.

Inside an LXC container lxcfs serves a cgroup-scoped /proc/meminfo, but
/proc/spl/kstat/zfs/arcstats is not virtualized and still reports the
host's ARC. Subtracting that ARC from the container's used memory
saturates to zero on any container smaller than the host's ARC, so
Komodo shows 0% RAM for every LXC guest on a ZFS Proxmox host.

Detect the lxcfs case by checking /proc/self/mountinfo for a FUSE
filesystem mounted on /proc/meminfo and skip the ARC adjustment there.
Bare-metal and Docker-on-host setups keep the ARC handling from moghtech#1489.
@phedoreanu

Copy link
Copy Markdown
Author

Verified on a live host. Proxmox 9 (kernel 7.0.14-17-pve, lxcfs 7.0.0-pve1, rpool on ZFS with a 6.28 GiB ARC), unprivileged Ubuntu LXC with a 1.5 GiB limit, Periphery as a systemd binary (not in Docker). Consecutive stats records for that server in Core's Stats collection, 15 s apart, with the patched Periphery starting between the second and the first:

2026-09-16T11:17:45Z  mem_used_gb=0.000  mem_buff_cache_gb=0.279  mem_zfs_arc_gb=6.278  mem_total_gb=1.5   (stock 2.3.3)
2026-09-16T11:18:00Z  mem_used_gb=0.000  mem_buff_cache_gb=0.279  mem_zfs_arc_gb=6.278  mem_total_gb=1.5   (stock 2.3.3)
2026-09-16T11:18:15Z  mem_used_gb=0.698  mem_buff_cache_gb=0.281  mem_zfs_arc_gb=0.000  mem_total_gb=1.5   (this patch)

free -m inside the container reports 681 MiB used at the same time, so the number is now the container's real usage.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant