Files
orca/.ciagent/RESEARCH_v0.14.md
T
Jon Chery b925fda3aa docs(P00): research findings — v0.14 ingress bootstrap
7 research topics: nft SNAT masquerade syntax, pve-firewall
coexistence (priority collision fix), podman-in-LXC (fuse=1
requirement), traefik v3.3 TLS model (certResolver does not exist —
use dynamic tls.certificates), pct create floating-IP syntax, podman
restart persistence (podman-restart.service), SELinux :Z omission.

Key findings that change the plan:
- nft postrouting: ip saddr 127.0.0.0/8 oifname != lo masquerade
- nft first-apply: pre-create table before nft -f
- pve-firewall: shift orca input/forward to priority -10
- LXC features: nesting=1,keyctl=1,fuse=1 (fuse=1 for fuse-overlayfs)
- traefik TLS: drop certResolver: orca, use dynamic tls.certificates
- podman: --restart=unless-stopped + enable podman-restart.service
- volumes: omit :Z flag, use :ro on both mounts

---ci---
project: orca
phase: 0
milestone: v0.14
status: research
---/ci---
2026-08-10 18:12:59 +00:00

1036 lines
48 KiB
Markdown

# Research: v0.14 Ingress Bootstrap Completeness
**Milestone**: v0.14 — Ingress Bootstrap Completeness
**Role**: ci-researcher
**Date**: 2026-08-10
**Status**: Complete
This document records findings, recommended approaches, confidence
levels, and pitfalls for the seven technical topics blocking the v0.14
ingress bootstrap milestone. The milestone replaces the traefik
binary+systemd install (`internal/traefik/install.go`) with a podman
container running a custom `orca-traefik` image, and completes the nft
SNAT+DNAT ingress stack in `internal/emitter/nft.go`.
Each topic is structured as: (a) findings, (b) recommended approach,
(c) confidence level, (d) pitfalls.
---
## Topic 1: nft SNAT/MASQUERADE postrouting syntax
### (a) Findings
The existing emitter (`internal/emitter/nft.go:180-184`) defines a
`prerouting` nat chain with DNAT rules but has NO `postrouting` chain.
For the R-017 hybrid model (Traefik on 127.0.0.1:8443, nft DNATs public
:443 → 127.0.0.1:8443), the return path needs masquerade so replies
from 127.0.0.1:8443 appear to come from the public IP.
**Chain type and hook**: nft nat chains support `prerouting` and
`postrouting` hooks in the `inet` family (nftables wiki "Configuring
chains": nat chain type is "supported by the ip, ip6 and inet table
families"). The postrouting hook sees packets "after routing, just
before they leave the local system."
**`masquerade` vs `snat to`**: Per Gentoo wiki nftables/Examples "Basic
NAT": "If we have a static IP, it would be slightly faster to use source
nat (SNAT) instead of masquerade. This way the router would replace the
source with a predefined IP, instead of looking up the outgoing IP for
every packet." `masquerade` is correct when the outgoing IP is dynamic
or when there are multiple egress interfaces — which is the orca case
(the public IP may be on any interface). `masquerade` is available
since kernel 3.18.
**Match expression — `oifname != "lo"` vs `ip saddr 127.0.0.0/8`**:
The DNAT rewrites destination to 127.0.0.1:8443, so the reply source is
127.0.0.1. We must masquerade only traffic whose source was rewritten
to loopback by the DNAT — i.e. traffic leaving a non-loopback interface
with source 127.0.0.0/8. The correct expression is:
```
ip saddr 127.0.0.0/8 oifname != "lo" masquerade
```
A bare `oifname != "lo" masquerade` would masquerade ALL non-loopback
egress, which is over-broad (it would NAT legitimate traffic that
wasn't DNAT'd). Matching `ip saddr 127.0.0.0/8` scopes masquerade to
exactly the DNAT'd return path. This is the canonical pattern for
"hairpin NAT" / "loopback DNAT return path."
**`inet` family and IPv6**: `masquerade` in an `inet` table works for
both IPv4 and IPv6, BUT: (1) the `ip saddr 127.0.0.0/8` match is
IPv4-only, so the rule only applies to IPv4 packets; (2) IPv6 has no
NAT-masquerade analogue in common use (RFC 6296 is rarely deployed) and
orca's DNAT is IPv4-only (`dnat to 127.0.0.1:8443` is IPv4). For v0.14
the postrouting chain should be IPv4-scoped. If IPv6 ingress is added
later, a separate `ip6 saddr ::1/128 oifname != "lo" masquerade` rule
would be needed, but that is out of scope for v0.14.
**`flush table inet orca-ingress` + re-apply**: The existing emitter
does `flush table inet orca-ingress` then recreates the table
(`nft.go:144-145`). This is the documented idempotent pattern
(nftables wiki "Flushing chains": `flush table` deletes all rules in
the table). Adding a postrouting chain to the flushed+recreated table is
safe — `flush table` removes all chains/sets/rules in the table but
keeps the table itself; the declarative re-create reinstates everything.
The one caveat: `flush table` on a non-existent table errors. The
emitter currently relies on the table existing or on `nft -f` tolerating
the flush-then-create. Since the file is `#!/usr/sbin/nft -f` and is
applied as a single transaction, `flush table` failing on first apply
(non-existent table) would abort the whole file. **This is a latent
bug**: on first-ever apply, `flush table inet orca-ingress` errors with
"No such file or directory" and the table is never created. The fix is
`delete table inet orca-ingress` (which tolerates absence? — no, it also
errors) or `add table inet orca-ingress` first, or use
`flush ruleset inet orca-ingress` — actually the robust idiom is:
```
table inet orca-ingress
delete table inet orca-ingress
table inet orca-ingress {
...
}
```
But `delete table` on a missing table also errors. The truly idempotent
pattern is to wrap in `add table inet orca-ingress` (no-op if exists in
`nft -f`? — `add table` errors if exists). The cleanest solution: use
`flush ruleset` (wipes EVERYTHING — dangerous, conflicts with
pve-firewall, see Topic 2) OR omit the flush and rely on `nft -f`
replacing the table atomically. **Actually `nft -f` with a `table ...{
}` block does NOT atomically replace an existing table of the same
name — it errors "File exists"** unless preceded by `delete table` or
`flush table`. The working idiom that survives first-apply is:
```
#!/usr/sbin/nft -f
add table inet orca-ingress 2>/dev/null
flush table inet orca-ingress
table inet orca-ingress { ... }
```
But `nft -f` doesn't support shell redirect semantics inside the file.
The real fix: emit `delete table inet orca-ingress` and accept that
`nft -f` treats a missing-table delete as a warning (nft 1.0+ tolerates
this in `-f` mode? — needs verification). The safest cross-version
approach is to apply via `nft -f <file>` where the file begins with
`flush table inet orca-ingress` and the operator pre-creates the table
with `nft add table inet orca-ingress` if absent. For orca, the install
step should run `nft add table inet orca-ingress 2>/dev/null || true`
before the first `nft -f` apply. **This is a pitfall the v0.14 emitter
must address** — see pitfalls.
### (b) Recommended approach
Add a `postrouting` nat chain to `renderNftRuleset` in
`internal/emitter/nft.go`, after the `prerouting` chain:
```nft
chain postrouting {
type nat hook postrouting priority 100; policy accept;
ip saddr 127.0.0.0/8 oifname != "lo" masquerade
}
```
Priority `100` is the standard `srcnat` priority (nftables wiki "Netfilter
hooks": `NF_IP_PRI_SRCNAT = 100`). The existing `prerouting` uses
priority `-100` (`dstnat`), which is correct and consistent.
For the first-apply / flush-table pitfall: change the emitter to emit
`add table inet orca-ingress` is not valid inside a `table {}` block.
Instead, the apply command in the SSH-push transport should run:
```bash
nft list table inet orca-ingress >/dev/null 2>&1 || nft add table inet orca-ingress
nft -f /etc/nftables.d/orca.nft
```
OR change the rendered file to use the `delete table` idiom (nft ≥ 1.0
treats delete-of-missing as warning, not error, in `-f` mode — but this
is version-dependent). The pre-create approach is robust across all nft
versions.
### (c) Confidence level
**High**. The chain syntax, hook, priority, and `masquerade` keyword are
all documented and widely deployed. The `ip saddr 127.0.0.0/8 oifname !=
"lo"` match is the standard loopback-DNAT-return pattern. The only
medium-confidence item is the first-apply flush-table behavior, which
depends on nft version and should be verified on the target Proxmox
kernel.
### (d) Pitfalls
1. **First-apply `flush table` on non-existent table errors** and aborts
the `nft -f` transaction, leaving no table created. The installer
must pre-create the table or the emitter must use a tolerate-absence
idiom.
2. **Over-broad masquerade** (`oifname != "lo" masquerade` without the
`ip saddr 127.0.0.0/8` match) would NAT all egress traffic and break
non-ingress routing. Always scope to the loopback source.
3. **IPv6**: the `ip saddr` match is IPv4-only; do not assume the rule
covers IPv6. The DNAT itself is IPv4-only (`127.0.0.1`), so this is
consistent, but document it.
4. **`masquerade` vs `snat to <public-ip>`**: if the public IP is
static and known, `snat to <ip>` is slightly faster (no per-packet
interface lookup). For orca's generic case (IP may be dynamic,
multiple interfaces), `masquerade` is safer. If the cluster config
carries an explicit public IP, a future optimization can emit `snat
to`.
5. **Connection tracking**: `masquerade` only applies to the first
packet of a flow (nftables wiki: "Only the first packet of a given
flow hits this chain; subsequent packets bypass it"). This is correct
behavior — conntrack handles the rest — but means the chain must not
be used for filtering.
---
## Topic 2: Proxmox pve-firewall vs nft conflict
### (a) Findings
**pve-firewall uses iptables-nft, not native nft tables.** Per the
Proxmox VE Firewall docs (pve-docs chapter-pve-firewall, version 9.2.4):
"The firewall runs two service daemons on each node: pvefw-logger
(NFLOG daemon) and pve-firewall (updates iptables rules)." The docs
explicitly say `iptables-save` to inspect generated rules. So the stock
`pve-firewall` manages **iptables** rules (which on modern Proxmox
translate to the nft `inet` backend via `iptables-nft`), in the
`security` chains, NOT a custom nft table.
**Proxmox also offers `proxmox-firewall` (nftables-based, tech
preview).** The docs note: "As an alternative to pve-firewall we offer
proxmox-firewall, which is an implementation of the Proxmox VE firewall
based on the newer nftables rather than iptables." This is gated behind
the `nftables: <boolean>` option in `/etc/pve/nodes/<nodename>/host.fw`
(default `0`). When enabled, proxmox-firewall uses its own nft tables.
**Does pve-firewall flush ALL nft tables?** The stock pve-firewall
manages iptables rules via `iptables-restore`-style operations on its
own chains. It does NOT flush unrelated nft tables — `iptables-nft`
operates on the `xt` compat chains in nft, not on arbitrary user tables.
A separate `table inet orca-ingress` is invisible to pve-firewall and
will NOT be flushed by `pve-firewall restart` or `pve-firewall update`.
**However**, if `proxmox-firewall` (the nft tech-preview) is enabled,
its behavior is less certain — it may use `flush ruleset` or operate on
a specific table. The docs don't specify its flush scope, and it's a
tech preview, so the risk is low for v0.14 (most Proxmox 8/9 deployments
use stock pve-firewall).
**Priority conflicts**: pve-firewall's iptables chains run at standard
iptables priorities. orca's `orca-ingress` table uses `priority -100`
(prerouting nat) and `priority 100` (postrouting nat), and `priority
filter` (input/forward). These are standard priorities and nft executes
all base chains at a hook in priority order — pve-firewall's iptables
chains and orca's nft chains coexist (a packet traverses ALL base chains
at a hook in priority order, per nftables wiki "Base chain priority":
"packets will traverse all of the chains within the scope of a given
hook until they are either dropped or no more base chains exist"). So
there is no priority collision that would skip orca's rules — but there
IS a semantic interaction: if pve-firewall DROPs a packet in its input
chain (priority 0, after orca's input at priority `filter` which is also
0 — same priority means undefined order!), orca's accept verdict is not
final. **The orca `input` chain uses `priority filter` which is the
symbolic name for 0, the same as pve-firewall's INPUT chain**. This is a
real concern: two base chains at the same hook + same priority have
undefined evaluation order. The fix: give orca's chains a distinct
priority (e.g. `priority -10` for input, ahead of pve-firewall's 0) so
orca's SYN-flood filter runs deterministically before pve-firewall.
**Recommended approach for third-party nft rules on Proxmox**: The
Proxmox community consensus (forums, docs) is that a separate nft table
with a distinct name (`orca-ingress`) coexists fine with pve-firewall's
iptables-managed chains, as long as you don't touch pve-firewall's
chains. `pve-firewall restart` regenerates only its own iptables rules.
### (b) Recommended approach
1. Keep `table inet orca-ingress` as a separate, named table — do NOT
use `flush ruleset` anywhere in the orca emitter (that would wipe
pve-firewall's rules).
2. Shift orca's `input` and `forward` chains to a priority ahead of
pve-firewall's iptables chains to guarantee deterministic order.
Change `priority filter``priority -10` (or `priority mangle -10`)
for the `input` chain, and similarly for `forward`. Keep the nat
chains at `priority -100`/`100` (nat priorities don't collide with
pve-firewall's filter chains).
3. Document that orca's nft table is independent of pve-firewall and
survives `pve-firewall restart`/`update`.
4. If `proxmox-firewall` (nft tech-preview) is enabled, add a doctor
check warning that interaction is untested; recommend staying on
stock pve-firewall for v0.14.
### (c) Confidence level
**Medium-High** for stock pve-firewall coexistence (well-documented
iptables-based, separate table is safe). **Low** for the nft
tech-preview `proxmox-firewall` (docs don't specify flush scope; it's a
tech preview and uncommon). **Medium** for the priority-collision fix
(the nftables wiki confirms same-priority undefined order; shifting
priority is the textbook fix, but the exact pve-firewall iptables
priority values aren't in the docs — they're standard iptables
priorities which map to 0).
### (d) Pitfalls
1. **Same-priority base chains have undefined evaluation order**
orca's `input` (`priority filter` = 0) and pve-firewall's INPUT
(iptables priority 0) may run in either order. An `accept` from orca
does NOT prevent pve-firewall from later dropping the packet. Shift
orca's priority to be deterministic.
2. **`flush ruleset` would destroy pve-firewall** — never emit it. The
current emitter uses `flush table inet orca-ingress` (scoped), which
is safe.
3. **`proxmox-firewall` (nft tech-preview)**: untested interaction. Add
a doctor check.
4. **pve-firewall's `ipfilter-net*` IP sets** enforce source-IP
spoofing protection on VM/CT interfaces — if orca's DNAT'd traffic
egresses a bridge that has ipfilter enabled, the 127.0.0.1 source
may be dropped as spoofed. This is only relevant if orca runs inside
a VM/CT with firewall+ipfilter enabled on its net interface. For the
host-level ingress case (nft on the Proxmox host), this doesn't
apply.
5. **`pve-firewall stop` does NOT remove orca's table** (good) but
`pve-firewall start` re-adds its own rules (also fine).
---
## Topic 3: Podman inside Proxmox LXC (nesting/keyctl)
### (a) Findings
**`nesting=1,keyctl=1` is the documented requirement for running
containers (docker/podman) inside an unprivileged LXC.** Per Proxmox VE
Linux Container docs (`features` key, line 2174-2200 of the LXC doc):
- `keyctl=<boolean>`: "For unprivileged containers only: Allow the use
of the keyctl() system call. This is required to use docker inside a
container. By default unprivileged containers will see this system
call as non-existent."
- `nesting=<boolean>`: "Allow nesting. Best used with unprivileged
containers with additional id mapping. Note that this will expose
procfs and sysfs contents of the host to the guest. This is also
required by systemd to isolate services."
So `--features nesting=1,keyctl=1` on an **unprivileged** LXC is the
correct and sufficient configuration for podman. `keyctl` is explicitly
"required to use docker inside a container" and applies to podman
equally (podman uses the kernel keyring for storage creds).
**Is `--privileged` needed for rootful podman?** No. Rootful podman
inside an unprivileged LXC with `nesting=1,keyctl=1` works because the
LXC's "unprivileged" refers to the user-namespace mapping of the LXC
itself; rootful podman inside it runs as the LXC's root (mapped UID).
The Proxmox docs note privileged containers are "unsafe" and should be
avoided. **Recommendation: unprivileged LXC + `nesting=1,keyctl=1` +
rootful podman inside** (root inside the LXC is fine; the LXC is still
unprivileged from the host's view).
**Ubuntu 24.04 LXC template packages**: The Ubuntu 24.04 LXC template
is minimal. Beyond `podman`, you need podman's runtime dependencies
which are NOT all pulled in by the `podman` metapackage on Ubuntu 24.04:
- `conmon` (container monitor — sometimes a separate package)
- `crun` or `runc` (OCI runtime — `crun` preferred, `runc` ≥ 1.1.11)
- `fuse-overlayfs` (for rootless overlay; see next finding)
- `uidmap` (for rootless subuid/subgid — only needed if running rootless
podman; rootful podman inside the LXC doesn't need it)
- `netavark` or `containernetworking-plugins` (CNI networking — only if
using bridged container networking, NOT needed for `--network host`)
- `passt` (rootless networking — not needed for rootful/host network)
For orca's case (rootful podman, `--network host`), the minimal set is:
`podman`, `conmon`, `crun`, `fuse-overlayfs`. Install via
`apt-get install -y podman conmon crun fuse-overlayfs`.
**fuse-overlayfs inside LXC**: There IS a known issue. Podman's default
storage driver is `overlay`, which requires a backing filesystem that
supports overlay (the LXC's rootfs, if on ZFS subvolume or ext4 image,
may or may not support native overlay). Inside an LXC, the overlay
driver often fails because the kernel's overlay mount requires
`CONFIG_OVERLAY_FS` and the underlying fs must not be on a
copy-on-write filesystem that confuses overlay (ZFS subvolumes as LXC
rootfs are a known problem). The standard workaround is
`fuse-overlayfs` as a `mount_program` in `storage.conf`, OR fall back to
the `vfs` storage driver (correct but slow — full copy per layer). The
Proxmox LXC `features` key also has a `fuse=<boolean>` option (default
0): "Allow using fuse file systems in a container. Note that
interactions between fuse and the freezer cgroup can potentially cause
I/O deadlocks." **So for fuse-overlayfs inside LXC, you must also set
`--features fuse=1`** in addition to nesting/keyctl. This is a critical
finding — without `fuse=1`, fuse-overlayfs cannot mount inside the LXC.
**`--network host` inside an LXC**: `--network host` makes the podman
container share the LXC's network namespace. Since the LXC already has
its own network namespace (that's what an LXC IS), `--network host`
binds the container to the LXC's netns — which is exactly what orca
wants (traefik binds 127.0.0.1:8080/8443 on the LXC's loopback, and nft
on the Proxmox host DNATs to the LXC's IP). This works correctly and is
the simplest networking model. No `netavark`/CNI needed.
**Performance**: Podman inside LXC vs on the host has negligible
overhead for CPU/memory (LXC uses the host kernel directly, no
emulation). The main overheads are: (1) double namespace overhead (LXC
netns + podman container netns — avoided with `--network host`); (2)
storage I/O — overlay-in-LXC may be slower if using `vfs` or
fuse-overlayfs vs native overlay on the host; (3) no device passthrough
limitations beyond normal LXC. For a traefik data plane (network I/O
bound, low disk I/O), LXC+podman performance is effectively native.
### (b) Recommended approach
Create the LXC unprivileged with:
```
pct create <vmid> local:vztmpl/ubuntu-24.04-standard_<ver>_amd64.tar.zst \
--hostname ingress \
--features nesting=1,keyctl=1,fuse=1 \
--unprivileged 1 \
--net0 bridge=vmbr0,hwaddr=<mac>,ip=<floating-ip>/<prefix>,gw=<gateway> \
--onboot 1 \
--memory 2048 --swap 0
```
Inside the LXC, install rootful podman:
```bash
apt-get update && apt-get install -y podman conmon crun fuse-overlayfs
```
Run traefik rootful with `--network host`:
```bash
podman run -d --name orca-traefik --restart=always --network host \
-v /etc/traefik/dynamic:/etc/traefik/dynamic \
-v /etc/traefik/traefik.yml:/etc/traefik/traefik.yml:ro \
orca-traefik:v3.3.0
```
Configure podman storage to use fuse-overlayfs as the overlay
`mount_program` (in `/etc/containers/storage.conf`):
```ini
[storage]
driver = "overlay"
runroot = "/run/containers/storage"
graphroot = "/var/lib/containers/storage"
[storage.options.overlay]
mount_program = "/usr/bin/fuse-overlayfs"
```
If fuse-overlayfs still fails (some LXC kernel configs block it), fall
back to `driver = "vfs"` in `storage.conf` (documented as a fallback,
slow but correct).
### (c) Confidence level
**High** for `nesting=1,keyctl=1` (explicitly documented as required for
docker/podman in LXC). **High** for `--network host` binding the LXC's
netns. **Medium** for the `fuse=1` requirement (the `features.fuse`
docs say it's needed for fuse filesystems in LXC, and fuse-overlayfs is
a fuse filesystem — logical, but not explicitly tested in the docs with
podman). **Medium** for the package list (Ubuntu 24.04's podman
metapackage may already pull conmon/crun — needs verification on the
actual template version).
### (d) Pitfalls
1. **Missing `fuse=1` feature** → fuse-overlayfs cannot mount → podman
storage init fails. Must include `fuse=1` in `--features`, or fall
back to `vfs` storage driver.
2. **`vfs` storage driver is slow** (full copy per layer, no CoW) — only
use as fallback. For traefik (small image, read-mostly), the
performance hit is acceptable but not ideal.
3. **ZFS-backed LXC rootfs + overlay** is a known-bad combination. If
the Proxmox storage is ZFS, the LXC rootfs is a ZFS subvolume, and
native overlay may fail. fuse-overlayfs (with `fuse=1`) is the
workaround.
4. **`keyctl=1` breaks systemd-networkd** — the Proxmox docs warn:
"Essentially, you can choose between running systemd-networkd or
docker [keyctl]." If the LXC uses systemd-networkd for networking,
enabling `keyctl=1` can cause systemd-networkd to fail. orca's
ingress LXC uses static IP config (not systemd-networkd), so this is
not an issue — but document it.
5. **`--privileged` LXC is unnecessary and unsafe** — don't use it.
Unprivileged + nesting/keyctl is the supported path.
6. **AppArmor in LXC** may restrict podman. Proxmox LXC uses AppArmor;
podman generally works but if conmon is blocked, check
`dmesg`/audit for AppArmor denials.
7. **`--onboot 1`** starts the LXC on Proxmox boot, but does NOT
auto-start the podman container — see Topic 6.
---
## Topic 4: Traefik v3.3 container image customization
### (a) Findings
**Base image entrypoint**: The official `traefik:v3.3.0` Dockerfile
(GitHub `traefik/traefik` v3.3 branch, `Dockerfile`):
```dockerfile
FROM alpine:3.21
RUN apk add --no-cache --no-progress ca-certificates tzdata
COPY ./dist/$TARGETPLATFORM/traefik /
EXPOSE 80
VOLUME ["/tmp"]
ENTRYPOINT ["/traefik"]
```
So the binary is at `/traefik` (not `/usr/local/bin/traefik`), and
`ENTRYPOINT ["/traefik"]` is correct. There is no default `CMD`
traefik expects CLI args or `--configFile`.
**Bake static config + CMD**: Yes. You can bake
`/etc/traefik/traefik.yml` into the image and set
`CMD ["--configFile=/etc/traefik/traefik.yml"]`. Traefik reads static
config from the file specified by `--configFile` (Traefik docs
"Static Configuration: File"). The static config can also be at the
default search paths (`/traefik.yml`, `/etc/traefik/traefik.yml`, etc.)
without `--configFile`, but being explicit is safer. Example
`Dockerfile.traefik`:
```dockerfile
FROM traefik:v3.3.0
COPY traefik.yml /etc/traefik/traefik.yml
CMD ["--configFile=/etc/traefik/traefik.yml"]
```
The existing `internal/traefik/install.go:72` already uses
`--configFile=/etc/traefik/traefik.yml`, so the container CMD matches.
**CA-based TLS resolver — CRITICAL FINDING**: Traefik v3.3
`certificatesResolvers` supports ONLY `acme` and `tailscale` (confirmed
in the static-config reference: `certificatesResolvers.<name>.acme.*`
and `certificatesResolvers.<name>.tailscale.*` are the only sub-keys).
**There is NO `certificatesResolvers.orca.tls` pointing at a CA file.**
The existing `internal/emitter/traefik.go:53` defines
`traefikRouterTLSCertResolver = "orca"` and the dynamic config emits
`tls: certResolver: orca` — but this only works if a
`certificatesResolvers.orca` exists in the STATIC config, and that
resolver must be `acme` or `tailscale`. A custom CA is NOT a
"certificate resolver" in Traefik's terminology.
**How custom CA TLS actually works in Traefik**: Custom CA certificates
are provided via the DYNAMIC config, not a cert resolver:
- **Server-side TLS cert**: `tls.certificates: [{certFile: ..., keyFile:
...}]` in dynamic config (Traefik "TLS" docs, "User defined"
section).
- **Client-auth CA (mTLS)**: `tls.options.<name>.clientAuth.caFiles:
[...]` + `clientAuth.clientAuthType: RequireAndVerifyClientCert` in
dynamic config.
- **Default cert**: `tls.stores.default.defaultCertificate: {certFile,
keyFile}` in dynamic config.
So the orca "step-ca root CA as TLS resolver" model is architecturally
mismatched with Traefik. The correct approach for v0.14:
- Drop `certResolver: orca` from the dynamic-config router TLS stanza
(it references a non-existent resolver).
- Instead, emit `tls.certificates` with the step-ca-issued cert+key
(server identity) and `tls.options.default.clientAuth.caFiles` with
the step-ca root CA (for mTLS client verification).
- OR, if mTLS is not required for v0.14 and only server TLS is needed,
emit `tls.stores.default.defaultCertificate` pointing at a
step-ca-issued server cert+key.
**Graceful degradation when CA file is absent**: Traefik logs an error
and holds the last-good dynamic config if a cert file is missing or
unparseable (documented behavior: "If the new config is malformed,
Traefik logs an error and holds last-good config"). It does NOT crash
on a missing dynamic-config cert file — it skips that cert and logs.
For the STATIC config, a missing `--configFile` IS fatal (traefik won't
start). So: bake a minimal valid static config in the image (always
present), and mount dynamic config (certs) from the host — if the cert
file is absent, traefik starts but that TLS config doesn't load.
**CAP_NET_BIND_SERVICE for `--network host`**: With `--network host`,
the container shares the host (LXC) netns. Binding ports < 1024 requires
`CAP_NET_BIND_SERVICE` OR root. Since orca runs rootful podman (root
inside the LXC), the traefik process runs as root and can bind privileged
ports without `CAP_NET_BIND_SERVICE`. **But** orca binds 127.0.0.1:8080
and 127.0.0.1:8443 (both ≥ 1024), so `CAP_NET_BIND_SERVICE` is NOT
needed at all. No `--cap-add` required. (If the opt-out
`traefik-on-public-ip` mode binds :80/:443 directly, then root handles
it — still no cap needed for rootful.)
**Base image is Alpine** — note that the orca traefik install currently
downloads a Linux binary (`internal/traefik/install.go:24`) which is the
same binary. The container image uses the official Alpine-based image
with `ca-certificates` and `tzdata` pre-installed, which is an advantage
(mTLS CA verification needs `ca-certificates`).
### (b) Recommended approach
`Dockerfile.traefik`:
```dockerfile
FROM traefik:v3.3.0
LABEL org.opencontainers.image.title="orca-traefik"
LABEL org.opencontainers.image.source="https://git.cloudinit.dev/coreci/orca"
# Bake the static config (entrypoints, file provider, logging).
# The dynamic config (routers, certs) is mounted at runtime.
COPY traefik.yml /etc/traefik/traefik.yml
ENTRYPOINT ["/traefik"]
CMD ["--configFile=/etc/traefik/traefik.yml"]
```
Static config (`traefik.yml`, rendered by
`internal/emitter/traefik.go:282-300` — already correct, just bake it):
```yaml
entryPoints:
websecure:
address: "127.0.0.1:8443"
web:
address: "127.0.0.1:8080"
traefik:
address: "127.0.0.1:8081"
providers:
file:
directory: "/etc/traefik/dynamic"
watch: true
log:
level: INFO
format: json
accessLog:
format: json
```
Run command (no `--cap-add`, no SELinux flag — see Topic 7):
```bash
podman run -d --name orca-traefik --restart=always --network host \
-v /etc/traefik/dynamic:/etc/traefik/dynamic \
orca-traefik:v3.3.0
```
**TLS model change for v0.14**: The `TraefikEmitter` dynamic config must
drop `tls.certResolver: orca` and instead reference a TLS cert from the
dynamic `tls.certificates` store. This is a code change in
`internal/emitter/traefik.go:185-188` (the `tls:` stanza). Until the
step-ca integration mints server certs, emit a `tls: {}` stanza (Traefik
will use its generated default cert) OR omit `tls:` entirely (plain
HTTP). Document this as a v0.14 limitation: real mTLS lands when step-ca
mints certs into the dynamic dir.
### (c) Confidence level
**High** for the Dockerfile structure (verified against the official
v3.3 Dockerfile). **High** that `certificatesResolvers` only supports
acme/tailscale (confirmed in the static-config reference — no other
sub-keys exist). **High** that CAP_NET_BIND_SERVICE is not needed (ports
≥ 1024 + rootful). **Medium** for graceful degradation behavior on
missing dynamic cert (documented but should be tested).
### (d) Pitfalls
1. **`certificatesResolvers.orca.tls` does not exist** — the existing
emitter emits `certResolver: orca` which Traefik will reject or
ignore (router TLS with a non-existent resolver → Traefik logs a
warning and may not serve TLS). This is the biggest v0.14 finding:
the TLS model must change from "cert resolver" to "dynamic
tls.certificates".
2. **Binary path is `/traefik`** in the official image, not
`/usr/local/bin/traefik`. The `Dockerfile.traefik` extends the
official image so `ENTRYPOINT ["/traefik"]` is inherited — don't
override it unless you copy the binary elsewhere.
3. **`CMD` vs `ENTRYPOINT`**: `ENTRYPOINT ["/traefik"]` + `CMD
["--configFile=..."]` means `podman run orca-traefik` runs
`/traefik --configFile=...`. If the operator passes extra args
(`podman run orca-traefik --log.level=DEBUG`), they APPEND to CMD
(not replace) — actually they REPLACE CMD. To append, use
`podman run orca-traefik --configFile=/etc/traefik/traefik.yml
--log.level=DEBUG`. Document this.
4. **Static config baked, dynamic mounted**: the static config
(`traefik.yml`) is baked into the image (immutable, versioned). The
dynamic config (routers/services/certs) is bind-mounted from
`/etc/traefik/dynamic` on the host so orca can update it atomically
(C-10 protocol). Do NOT bake the dynamic config into the image.
5. **Alpine base + CA certs**: the official image has
`ca-certificates` installed. If you build a custom image from
`scratch` or `distroless`, you must install CA certs or TLS
verification to upstreams (step-ca, ACME) will fail.
6. **`--network host` + Alpine**: Alpine's `/etc/resolv.conf` handling
under `--network host` is fine (shares host netns). No issue.
---
## Topic 5: `pct create` with floating IP + MAC
### (a) Findings
**Correct `pct create` syntax**: Per the Proxmox VE Linux Container docs
("Managing Containers with pct", line 1792 and "CLI Usage Examples",
line 1812):
```
pct set 100 -net0 name=eth0,bridge=vmbr0,ip=192.168.15.147/24,gw=192.168.15.1
```
The `net[n]` parameter format (docs line 959):
```
net[n]: name=<string> [,bridge=<bridge>] [,firewall=<1|0>] [,gw=<GatewayIPv4>]
[,gw6=<GatewayIPv6>] [,hwaddr=<XX:XX:XX:XX:XX:XX>]
[,ip=<(IPv4/CIDR|dhcp|manual)>] [,ip6=<...>] [,type=<veth>]
```
So the correct `pct create` with static IP, custom MAC, and gateway is:
```bash
pct create <vmid> local:vztmpl/ubuntu-24.04-standard_<ver>_amd64.tar.zst \
--hostname ingress \
--net0 name=eth0,bridge=vmbr0,hwaddr=<mac>,ip=<floating-ip>/<prefix>,gw=<gateway>
```
**`ip=<ip>/<prefix>` with public IP**: Yes, the `ip` value takes
`IPv4/CIDR` format (docs: `ip=<(IPv4/CIDR|dhcp|manual)>`). A public IP
with prefix works: `ip=203.0.113.10/24`. The prefix is REQUIRED (not
`ip=<ip>` bare) — without the prefix, Proxmox may not configure the
interface correctly (it needs the netmask). Use CIDR always.
**`gw=<gateway>`**: Correct parameter name (docs: `gw=<GatewayIPv4>`,
"Default gateway for IPv4 traffic"). For IPv6, use `gw6`.
**Getting the LXC's IP after `pct start`**: The IP is configured in the
LXC config (`/etc/pve/lxc/<vmid>.conf`), so it's known before start.
But to verify the LXC actually came up with it:
- `pct config <vmid>` — prints the config including `net0: ...ip=...`
(the configured IP).
- `pct list` — lists VMIDs and status (running/stopped), NOT IPs.
- `pct inspect <vmid>` — not a real command. Use `pct config`.
- `lxc-info -n <vmid>` — works if `lxc-tools` installed; shows IP
if running.
- **Most reliable**: `pct exec <vmid> -- ip -j addr show eth0 | jq -r
'.[0].addr_info[0].local'` (runs inside the LXC). Or `pct status
<vmid>` for status.
For orca, the IP is known at create time (it's in the `pct create`
command), so discovery is only needed for verification. Use
`pct config <vmid>` and parse the `net0` line, or `pct exec <vmid> --
hostname -I`.
**Proxmox native mode (separate LXC for traefik) — discovering the
LXC's bridge IP for the nft DNAT target**: In native mode, the nft
DNAT rule on the Proxmox host must target the LXC's bridge IP (not
127.0.0.1, since traefik is in a separate LXC, not on the host's
loopback). The LXC's IP is the `ip=<floating-ip>/<prefix>` from
`pct create`. The DNAT rule becomes:
```nft
tcp dport 443 dnat to <lxc-ip>:8443
tcp dport 80 dnat to <lxc-ip>:8080
```
The `<lxc-ip>` is the floating IP assigned to the LXC (or a dedicated
bridge IP if the LXC is on a private bridge). The emitter needs the
LXC's IP as input — this is a cluster-config field (e.g.
`ingress.traefik_ip` or derived from the LXC VMID via `pct config
<vmid>`). For v0.14, the `NftClusterConfig` should carry a
`TraefikDNATTarget` field (default `127.0.0.1` for the hybrid/host
mode, set to the LXC IP for native mode).
### (b) Recommended approach
```bash
# Create the ingress LXC (unprivileged, nesting for podman)
pct create 201 local:vztmpl/ubuntu-24.04-standard_24.04-1_amd64.tar.zst \
--hostname ingress \
--unprivileged 1 \
--features nesting=1,keyctl=1,fuse=1 \
--net0 name=eth0,bridge=vmbr0,hwaddr=02:ca:fe:00:00:01,ip=10.99.0.10/24,gw=10.99.0.1 \
--onboot 1 \
--memory 2048 --swap 0 \
--rootfs local-zfs:8
pct start 201
# Verify IP
pct config 201 | grep '^net0'
# or
pct exec 201 -- hostname -I
```
For the nft emitter, add a `TraefikDNATTarget` field to
`NftClusterConfig` (default `127.0.0.1`):
```go
type NftClusterConfig struct {
TrustedProbes []string
RateLimit int
RateBurst int
TraefikDNATTarget string // "127.0.0.1" (hybrid/host) or LXC IP (native)
}
```
And render:
```nft
tcp dport 443 dnat to <TraefikDNATTarget>:8443
tcp dport 80 dnat to <TraefikDNATTarget>:8080
```
### (c) Confidence level
**High** — the `pct create` syntax is directly from the Proxmox docs
examples. The `ip=<ip>/<prefix>` CIDR format and `gw=` parameter are
documented. IP discovery via `pct config`/`pct exec` is standard.
### (d) Pitfalls
1. **`ip=<ip>` without prefix** may not configure the interface — always
use CIDR (`ip=<ip>/<prefix>`).
2. **`hwaddr` must be unique** on the bridge — Proxmox auto-generates one
if omitted; if specifying a custom MAC, ensure no collision.
3. **`pct create` requires the template to be downloaded first** via
`pveam update && pveam download local <template>`. The bootstrap
(`internal/proxmox/bootstrap.go:257`) already handles template
download.
4. **`--onboot 1`** starts the LXC when the Proxmox node boots, but does
NOT start the podman container inside it (see Topic 6).
5. **Floating IP + MAC**: if the floating IP moves between nodes (HA),
the MAC should move with it so ARP caches stay valid. Proxmox HA
handles this if the LXC is HA-managed. For a non-HA single-node
ingress, a static IP is fine.
6. **Public IP directly on LXC**: if the LXC has a public IP on
`vmbr0` (bridged to the host's public interface), the LXC is directly
exposed — no nft DNAT needed (traefik binds 0.0.0.0:443 inside the
LXC). This is the "native, traefik-on-public-ip" mode. The nft DNAT
is only needed in hybrid mode (traefik on loopback, nft redirects
public → loopback) or when the LXC is on a private bridge and the
host does DNAT from public → LXC-private-IP.
---
## Topic 6: `podman run --restart=always --network host` persistence
### (a) Findings
**`--restart=always` and reboot**: Per the podman-run man page
("restart-policy" section, line 1002-1016 of the docs):
> `always`: Restart containers when they exit, regardless of status,
> retrying indefinitely.
> `unless-stopped`: Restart containers when they exit, unless the
> container was explicitly stopped by the user. After a system reboot,
> containers with this policy will be restarted by podman-restart.service
> only if they were not explicitly stopped by the user before the reboot.
> This differs from `always`, which restarts containers after a system
> reboot regardless of whether they were user-stopped.
>
> Podman provides a systemd unit file, podman-restart.service, which
> restarts containers after a system reboot.
So `--restart=always` does NOT by itself survive a reboot — it only
handles container exit/restart while the podman daemon is running. To
survive a HOST reboot, you need `podman-restart.service` enabled. The
podman docs explicitly say: "Podman provides a systemd unit file,
podman-restart.service, which restarts containers after a system
reboot."
**Podman is NOT a single long-running daemon** (unlike dockerd). Podman
is a CLI that launches containers via `conmon` + `runc`/`crun`. There is
no "podman service" that must be running for containers to stay up —
once started, a container is monitored by `conmon` (a separate process).
But for RESTART-on-reboot, `podman-restart.service` must be enabled so
that on boot, systemd runs `podman restart <container>` for all
containers with a restart policy.
**Ubuntu 24.04**: Installing the `podman` package installs the
`podman-restart.service` systemd unit but does NOT enable it by default.
The operator must enable it:
```bash
systemctl enable --now podman-restart.service
```
(There is no `podman.socket`/`podman.service` needed for rootful podman
— those are for the podman API socket, not for container lifecycle. The
only service needed for reboot-persistence is `podman-restart.service`.)
**Inside an LXC with `--onboot 1`**: `--onboot 1` starts the LXC when
Proxmox boots. Inside the LXC, systemd starts. If
`podman-restart.service` is enabled inside the LXC, it will restart the
traefik container when the LXC's systemd reaches it. So the chain is:
Proxmox boots → `--onboot 1` starts LXC → LXC systemd starts →
`podman-restart.service` restarts the container. This works, but
requires `podman-restart.service` enabled INSIDE the LXC.
**`--restart=always` vs `unless-stopped`**: Per the docs,
`unless-stopped` is generally preferable — if the operator explicitly
stops the container (e.g. for maintenance), it won't auto-restart on
reboot. `always` restarts even if explicitly stopped, which is
surprising for maintenance. For orca's traefik (which the operator may
stop for upgrades), `unless-stopped` is the better default. However,
orca's own update protocol (C-10 atomic reload) doesn't stop the
container — it hot-reloads via dynamic-config rename. So either policy
works; `unless-stopped` is safer for operator-driven stops.
**Alternative — systemd unit**: The podman docs recommend: "When
running containers in systemd services, use the restart functionality
provided by systemd. In other words, do not use this option in a
container unit, instead set the `Restart=` systemd directive." For
production, generate a systemd unit via `podman generate systemd` (or
the newer Quadlet) and let systemd manage restart. This is more robust
than `--restart=always` + `podman-restart.service`. For v0.14, the
simpler `--restart=unless-stopped` + `podman-restart.service` is
acceptable; a future version can move to Quadlet.
### (b) Recommended approach
1. Inside the LXC, enable `podman-restart.service`:
```bash
systemctl enable --now podman-restart.service
```
2. Run traefik with `--restart=unless-stopped` (not `always`):
```bash
podman run -d --name orca-traefik --restart=unless-stopped \
--network host \
-v /etc/traefik/dynamic:/etc/traefik/dynamic \
orca-traefik:v3.3.0
```
3. The orca installer (replacing `internal/traefik/install.go`) should:
- Install podman + deps inside the LXC.
- Enable `podman-restart.service`.
- `podman pull` the orca-traefik image.
- `podman run` with the above flags.
### (c) Confidence level
**High** — the podman-run man page explicitly documents the
`podman-restart.service` requirement and the `always` vs
`unless-stopped` semantics. The LXC `--onboot 1` → systemd →
podman-restart chain is logically sound.
### (d) Pitfalls
1. **`--restart=always` alone does NOT survive reboot** — must enable
`podman-restart.service`. This is the #1 pitfall.
2. **`podman-restart.service` is not enabled by default** on Ubuntu 24.04
— the installer must enable it.
3. **`--restart=always` restarts even after explicit `podman stop`** —
use `unless-stopped` to respect operator-driven stops.
4. **No `podman.service` daemon needed** — don't confuse
`podman-restart.service` (boot re-start) with `podman.service`/`podman.socket`
(API socket). Only the former is needed.
5. **`conmon` must persist** — if `conmon` dies, the container is
orphaned. `conmon` is launched by podman at `run` time and is a
standalone process; it survives independent of the podman CLI. On
reboot, `podman-restart.service` re-launches everything.
6. **Quadlet is the modern approach** (systemd unit generator for
podman) — consider for a post-v0.14 hardening pass. For v0.14,
`--restart=unless-stopped` + `podman-restart.service` is sufficient.
7. **`--network host` + reboot**: on reboot, `podman-restart.service`
re-runs the container with the same flags including `--network host`.
The LXC's netns is the same, so traefik rebinds 127.0.0.1:8080/8443.
No issue.
---
## Topic 7: SELinux volume mount `:Z` flag
### (a) Findings
Per the podman-run man page ("volume" section, line 1540-1544):
> Labeling systems like SELinux require that proper labels be placed on
> volume content mounted into a container. Without a label, the security
> system might prevent the processes running inside the container from
> using the content. By default, Podman does not change the labels set
> by the OS.
>
> To change a label in the container context, add either of two suffixes
> `:z` or `:Z` to the volume mount. These suffixes tell Podman to relabel
> file objects on the shared volumes. The `z` option tells Podman that
> two or more containers share the volume content. As a result, Podman
> labels the content with a shared content label. Shared volume labels
> allow all containers to read/write content. The `Z` option tells Podman
> to label the content with a private unshared label. Only the current
> container can use a private volume.
**On non-SELinux systems (Ubuntu, Debian/Proxmox)**: The `:z`/`:Z`
suffixes are no-ops when SELinux is disabled. Podman detects that
SELinux is not active and skips relabeling. There is no error and no
behavioral change. (Podman source: the relabel path checks
`selinux.GetEnabled()` and returns early if false.)
**Ubuntu 24.04**: No SELinux by default (Ubuntu uses AppArmor). The
`:Z` flag is a harmless no-op.
**Proxmox VE**: Debian-based, no SELinux by default (uses AppArmor for
LXC). The `:Z` flag is a harmless no-op.
**`:z` (lowercase, shared) vs `:Z` (uppercase, private)**: For orca's
traefik dynamic-config mount (`/etc/traefik/dynamic`), only the traefik
container reads it (orca writes from outside via the SSH-push
transport, not as a container). So `:Z` (private) is semantically
correct IF SELinux were active. But since orca writes the files from
the host side (not as a container), `:Z` would relabel the host
directory to a private container label, which could then BLOCK the host
orca process from writing to it (if SELinux were active). For the
non-SELinux case it doesn't matter. For forward-compatibility (if orca
is ever deployed on RHEL/SELinux), the safest choice is to OMIT the flag
and let the operator/site-local policy handle labels, OR use `:z`
(shared) so both host-side writers and the container can access it.
**Recommendation for orca**: Omit the `:Z` flag entirely. Rationale:
- Ubuntu 24.04 and Proxmox (the v0.14 targets) have no SELinux — the
flag is a no-op.
- Omitting it avoids the footgun where `:Z` relabels a host directory
and breaks host-side writes on future SELinux deployments.
- If an SELinux deployment is added later, add a documented
site-local `chcon -Rt container_file_t /etc/traefik/dynamic` step
instead of baking `:Z` into the run command.
### (b) Recommended approach
Omit the `:Z` flag. The podman run command is:
```bash
podman run -d --name orca-traefik --restart=unless-stopped \
--network host \
-v /etc/traefik/dynamic:/etc/traefik/dynamic \
-v /etc/traefik/traefik.yml:/etc/traefik/traefik.yml:ro \
orca-traefik:v3.3.0
```
Note: the static config is mounted `:ro` (read-only — it's baked into
the image, but mounting it read-only from the host allows updates
without rebuilding). The dynamic config is mounted read-write (Traefik
doesn't write to it, but the `:ro` would prevent the atomic-rename
from the host side if the mount is a bind mount — actually bind mounts
are of the directory, and `:ro` makes the mount read-only IN the
container, not on the host. The host can still rename files in the
source directory; the container just sees them read-only. Since Traefik
only reads, `:ro` is fine and safer for the dynamic dir too. But the
C-10 atomic rename happens on the HOST filesystem, so `:ro` on the
container mount doesn't block it.)
**Use `:ro` on both mounts** (Traefik never writes to either):
```bash
-v /etc/traefik/dynamic:/etc/traefik/dynamic:ro \
-v /etc/traefik/traefik.yml:/etc/traefik/traefik.yml:ro \
```
### (c) Confidence level
**High** — the podman docs explicitly document `:z`/`:Z` as SELinux
relabel suffixes and that they are no-ops when SELinux is disabled.
Ubuntu and Proxmox do not enable SELinux by default.
### (d) Pitfalls
1. **`:Z` on a host-writable directory can block host-side writes on
SELinux systems** — the private label makes the directory owned by
the container's SELinux context. Since orca writes dynamic config
from the host, `:Z` is the wrong choice if SELinux is ever enabled.
2. **`:z` (shared) is safer than `:Z` (private)** for orca's use case,
but still unnecessary on non-SELinux.
3. **Omitting the flag is the cleanest** for Ubuntu/Proxmox. Document
that SELinux deployments need `chcon` or `--security-opt
label=disable`.
4. **`--security-opt label=disable`** disables SELinux separation for
the container entirely — a valid alternative on SELinux systems if
you don't want to relabel, but reduces isolation. Not needed on
Ubuntu/Proxmox.
5. **Bind-mount `:ro` is container-side read-only**, not host-side. The
host can still write/rename; the container sees a read-only view.
This is correct for Traefik (read-only consumer) and doesn't block
the C-10 atomic rename (which is a host-side operation).
---
## Summary: highest-impact findings for v0.14
| # | Topic | Key finding | Impact |
|---|-------|-------------|--------|
| 1 | nft postrouting | Add `ip saddr 127.0.0.0/8 oifname != "lo" masquerade` chain | **Code change** in `internal/emitter/nft.go` |
| 1 | nft flush-table | `flush table` on non-existent table errors on first apply | **Installer fix** (pre-create table) |
| 2 | pve-firewall | Same-priority base chains have undefined order | **Code change**: shift orca input/forward priority |
| 3 | podman in LXC | Must add `fuse=1` to `--features` for fuse-overlayfs | **pct create fix** |
| 4 | traefik TLS | `certificatesResolvers.orca.tls` does NOT exist — only acme/tailscale | **Architecture change**: use dynamic `tls.certificates`, drop `certResolver` |
| 4 | traefik Dockerfile | `ENTRYPOINT ["/traefik"]`, binary at `/traefik`, CMD `--configFile` | **New Dockerfile.traefik** |
| 6 | podman restart | `--restart=always` needs `podman-restart.service` enabled to survive reboot | **Installer must enable service** |
| 6 | podman restart | Use `unless-stopped`, not `always` | **Run-flag change** |
| 7 | SELinux `:Z` | No-op on Ubuntu/Proxmox; omit it | **Run-flag simplification** |