Skip to content

Host and Container Networking Internals

A technical reference covering Linux network namespaces, rootless container networking, TUN/TAP devices, and their intersection with VPN workloads.

1. The Rootless Container Networking Problem

Running containers as a non-root user is the recommended security posture: a container escape can only yield the privileges of the host user, not root. Podman supports this natively; Docker historically did not.

The problem is that many useful network operations — creating virtual interfaces, configuring routing, writing iptables rules — are gated on CAP_NET_ADMIN in the kernel's initial user namespace, i.e. real root. A rootless container runner must solve this without elevating to root.

2. Network Namespaces — the Kernel Primitive

A network namespace (netns) is a kernel isolation unit. Each netns has its own:

  • network interfaces
  • routing tables
  • iptables/nftables rules
  • sockets

Every process belongs to exactly one netns. Containers get their own netns at creation time, which is why ip link inside a container only shows that container's interfaces.

Critically, every netns is owned by a user namespace. This ownership determines which capabilities are effective for operations inside that netns:

initial user namespace  →  owns host netns      (real root required)
container user namespace  →  owns container netns  (container user sufficient)

This ownership relationship is what makes rootless networking possible at all.

3. Userspace Network Stacks

The need

A rootless container's netns is isolated — it has no interfaces connected to the outside world by default. Connecting it to the host network requires creating a path across the netns boundary. In rootful mode this is done with kernel veth pairs or macvlan, both of which need root. In rootless mode a different mechanism is needed.

slirp4netns

The original solution. "slirp" is a technique dating from the 1990s (SLIP emulation over serial lines) that implements a complete TCP/IP stack in userspace.

How it works:

  1. Creates a TAP device in the container netns
  2. Reads raw frames from that TAP in userspace
  3. Parses and re-originates the packets as unprivileged host socket calls
  4. Translates responses back into raw frames and writes them to the TAP

Each packet crosses the kernel↔userspace boundary twice, and slirp4netns terminates and re-originates TCP connections internally. Consequences:

  • Source IP addresses are lost (NAT at the slirp layer)
  • Higher CPU overhead from full protocol stack re-processing
  • Not suitable for workloads that depend on client IP visibility

pasta

The modern replacement, default in Podman 5+. pasta does not reimplement TCP/IP. Instead:

  1. Creates a TAP device in the container netns (still needs /dev/net/tun)
  2. Connects it to a host socket
  3. Uses splice() to move data between the two kernel buffers directly

The kernel handles the protocol stack on both sides. pasta is just the coordinator that sets up the path and keeps it running. Source IPs are preserved on inbound connections.

Why these are Podman-specific — Docker's daemon model

slirp4netns and pasta are only needed for rootless runtimes.

Docker runs a root daemon (dockerd) that owns all container lifecycle operations. When Docker creates a container network, it does so in the daemon process which runs as root — so it can create kernel veth pairs and bridge interfaces directly, with no userspace TCP/IP stack needed. All Docker containers, even those running as a non-root user inside the container, benefit from this root-level network setup.

Podman is daemonless: each podman run is a direct process fork by the calling user, with no privileged background service. There is nothing running as root to set up kernel networking on its behalf. This is why Podman needs slirp4netns or pasta — and why Docker does not.

Rootless Docker (added later, via rootlesskit) faces the same problem and uses the same solution space.

4. How pasta Uses splice() for Zero-Copy Forwarding

The standard way to forward data between two sockets in userspace:

kernel buffer A → copy to userspace → copy to kernel buffer B

Two copies, two context switches per direction.

splice(2) is a Linux syscall that moves data between two kernel file descriptors without passing through userspace memory:

kernel buffer A → kernel buffer B  (zero userspace involvement)

pasta uses splice() to connect the container-side TAP file descriptor to the host-side socket file descriptor. The kernel transfers data directly between the two buffers; pasta wakes up to coordinate but does not touch the payload.

This is not the same as a fully autonomous kernel path (like veth or macvlan in rootful mode, where no userspace process is involved at all). pasta still participates per-batch. But the per-packet copy overhead is eliminated.

5. TUN/TAP Devices

Virtual network device taxonomy

All of the following are network interfaces from the kernel's perspective, but they differ in topology and backing:

Device Layer Backed by Typical use
TUN L3 (IP packets) userspace fd VPN tunnel (OpenVPN, WireGuard)
TAP L2 (Ethernet frames) userspace fd VM networking, VPN bridged mode, pasta/slirp4netns
veth L2 paired kernel interface Connect two network namespaces
bridge L2 set of attached ports Virtual L2 switch, container networking
dummy L3 nothing Loopback-like placeholder, testing
loopback L3 kernel loopback 127.0.0.1 within a netns
macvlan/ipvlan L2/L3 physical NIC Direct LAN attachment (rootful only)

TUN and TAP are the same kernel driver (drivers/net/tun.c), distinguished only by the flag passed to ioctl(TUNSETIFF): IFF_TUN vs IFF_TAP. TUN delivers raw IP packets; TAP delivers full Ethernet frames including MAC headers. OpenVPN supports both modes; WireGuard is TUN-only.

veth pairs

A veth (virtual Ethernet) is always created as a pair. A packet sent into one end comes out the other, like a physical patch cable. Neither end terminates in userspace — the kernel handles both.

netns A: veth0 ══════ veth1 :netns B

The canonical use is connecting a container netns to a bridge in the host netns: one veth end lives in the container (usually named eth0), the other lives in the host and is attached to a bridge.

veth differs from TAP in that there is no userspace fd involved — it is a purely kernel-level path and carries no per-packet userspace overhead.

bridge

A bridge is a virtual L2 switch. You attach interfaces to it; it learns MAC addresses and forwards frames between ports accordingly. It does not originate or terminate traffic itself.

        bridge0
       /   |   \
    eth0  veth0  veth2
           |      |
       container1 container2

This is the standard container networking primitive: each container gets a veth pair, one end in the container netns, the other attached to a shared bridge that has a host NIC or NAT rule for outbound access.

TAP and bridge are often used together for VM networking: a TAP device (backed by QEMU's fd) is attached to a bridge alongside a physical NIC, giving the VM a presence on the LAN.

What /dev/net/tun is

/dev/net/tun is a character device (major 10, minor 200). It is the factory for creating TUN and TAP virtual network interfaces from userspace:

  • TUN — layer 3 (IP packets). Used by OpenVPN, WireGuard.
  • TAP — layer 2 (Ethernet frames). Used by pasta, slirp4netns, VM hypervisors.

The open() + ioctl(TUNSETIFF) permission model

Creating a TUN interface is a two-step kernel interaction:

// Step 1: open the factory device
int fd = open("/dev/net/tun", O_RDWR);

// Step 2: allocate an interface inside the current netns
struct ifreq ifr = { .ifr_flags = IFF_TUN };
ioctl(fd, TUNSETIFF, &ifr);

Step 1 uses standard Unix file permissions. The device node is crw-rw-rw- — world read/write — so any user can open it. The device node is intentionally not the security boundary.

Step 2 is where the kernel enforces access control. The check is:

Does the calling process hold CAP_NET_ADMIN in the user namespace that owns the target network namespace?

This is the key: the capability is checked against the owning user namespace of the netns, not the initial user namespace globally.

Why capability checks are namespace-scoped

Because a rootless container's netns is owned by the container's user namespace (not the host's), CAP_NET_ADMIN granted inside that user namespace is sufficient for all TUN/interface operations scoped to that netns. The kernel does not require initial-namespace root.

This is also why network: host breaks rootless VPN containers: host netns is owned by the initial user namespace, so the capability check requires real root.

6. Running VPN Servers Rootless

Why privileged: true was thought necessary

Two independent issues fed this misconception:

A crun device cgroup bug (fixed in crun 1.13.0, October 2024) crun's device cgroup allowlist setup was too restrictive for character devices in rootless mode. Even with CAP_NET_ADMIN and --device /dev/net/tun correctly specified, the ioctl was blocked. The fix landed in crun 1.13.0. Users on affected versions observed that privileged: true was the only workaround, and that assumption propagated into documentation and compose templates.

iptables-legacy not working in user namespaces The legacy iptables binary uses the xtables kernel interface, which requires privileges in the initial user namespace. It fails with:

can't initialize iptables table `nat': Permission denied (you must be root)

This is a separate and permanent limitation — not a bug. The fix is to use iptables-nft (backed by nftables) which operates correctly within a user-namespace-owned netns.

The actual minimal requirements

cap_add:
  - NET_ADMIN
devices:
  - /dev/net/tun:/dev/net/tun
sysctls:
  - net.ipv4.ip_forward=1

No privileged: true needed on crun ≥ 1.13.

Why iptables-legacy fails but iptables-nft works

iptables-legacy speaks to the kernel via setsockopt() on a raw socket with IPT_SO_SET_REPLACE. This path requires CAP_NET_ADMIN in the initial user namespace.

iptables-nft is a frontend over nftables, which uses the nfnetlink netlink family. nftables was designed with network namespaces in mind: operations are scoped to the calling process's netns, and CAP_NET_ADMIN in the owning user namespace is sufficient.

The kylemanna/openvpn image ships both binaries. Swapping the startup command from iptables to iptables-nft is the only change required.

iptables persistence and split rule stores

iptables-save and iptables-restore are not persistence mechanisms — they dump and reload rules to/from a flat text file. Actual persistence requires something to call iptables-restore at boot (e.g. the iptables-persistent package, or a systemd unit with ExecStart).

The two backends maintain separate, invisible-to-each-other rule stores:

iptables-legacy-save  →  dumps x_tables rules
iptables-nft-save     →  dumps nf_tables rules  (disjoint set)
nft list ruleset      →  dumps nf_tables rules  (same as above)

Rules written via iptables-legacy are not visible to iptables-nft and vice versa. On systems with mixed tooling — Docker using legacy, application using nft — this produces silent rule gaps. Always use one backend consistently, and verify with iptables-nft -L or nft list ruleset rather than iptables -L when using the nft path.

nftables rules are also not persisted by default — they live in kernel memory and are lost on reboot. The nftables.service systemd unit (shipped with the nftables package on Debian/Ubuntu and RHEL/Fedora) loads /etc/nftables.conf at boot; enabling it is all that is needed:

nft list ruleset > /etc/nftables.conf
systemctl enable nftables.service

On Alpine there is no built-in persistence unit; a local init script is required. In container contexts rules are typically re-applied on each startup via the compose command, so host persistence is rarely relevant.

7. The network: host Boundary Collapse

Using network: host in a rootless container instructs the runtime to skip creating a private netns and attach the container directly to the host network namespace.

The consequence for capability checks:

With private netns:
  container user namespace owns container netns
  → CAP_NET_ADMIN inside container user namespace is sufficient

With network: host:
  host netns is owned by initial user namespace (real root)
  → CAP_NET_ADMIN inside container user namespace is NOT sufficient
  → kernel returns EPERM

All rootless networking capabilities — TUN creation, iptables-nft rules, route management — stop working under network: host. The isolation boundary that makes rootless possible is eliminated.

Capability check matrix by network mode

Operation Netns touched Private netns (rootless) network: host (rootless)
ip tuntap add tun0 container netns ✓ ✗
ip addr add / ip route add container netns ✓ ✗
ip link set eth0 up container netns (veth end) ✓ ✗
iptables-nft -t nat … container netns ✓ ✗
iptables-legacy -t nat … any ✗ ✗
ip link set eth0 on real host NIC host netns ✗ ✗

8. Kubernetes Equivalents

securityContext

securityContext is the Kubernetes API surface for the same capability and privilege settings configured manually in compose files:

securityContext:
  capabilities:
    add: ["NET_ADMIN"]
  privileged: false

Kubernetes passes this to the CRI (containerd or CRI-O), which passes it to the OCI runtime (crun or runc), which performs the same namespace and cgroup setup as podman run --cap-add NET_ADMIN. Same kernel mechanism, different entry point.

Device plugins

Device plugins solve a scheduling problem: some nodes have a device available and some do not. Without a plugin, there is no way to express "this pod needs /dev/net/tun to be present and accessible" as a schedulable resource.

The flow:

  1. A device plugin DaemonSet runs on each node and registers a resource with kubelet: e.g. vendor.com/tun: 1
  2. A pod requests the resource:
resources:
  limits:
    vendor.com/tun: "1"
  1. The scheduler places the pod only on nodes where the resource is available
  2. Kubelet queries the plugin for the device path and mount instructions
  3. Kubelet passes those to the CRI, which passes them to crun/runc

The end result is the same kernel-level device access — the value is in making it visible to the scheduler and auditable in cluster resource accounting.

CNI vs device plugins

These are parallel, non-overlapping concerns:

CNI Device plugin
Responsibility Network namespace wiring Hardware/device resource allocation
When it runs After pod netns is created During pod scheduling and startup
Example Calico assigns pod IP and routes GPU plugin exposes /dev/nvidia0

CNI plugins themselves often use /dev/net/tun internally (Flannel VXLAN, WireGuard-based CNIs) — but that is the CNI plugin running with cluster-level privileges, entirely separate from workload pod permissions.

References

Kernel and namespaces

Rootless containers and Podman networking

iptables and nftables

VPN and rootless containers

Kubernetes