Host and Container Networking Internals
A technical reference covering Linux network namespaces, rootless container networking, TUN/TAP devices, and their intersection with VPN workloads.
1. The Rootless Container Networking Problem
Running containers as a non-root user is the recommended security posture: a container escape can only yield the privileges of the host user, not root. Podman supports this natively; Docker historically did not.
The problem is that many useful network operations — creating virtual
interfaces, configuring routing, writing iptables rules — are gated on
CAP_NET_ADMIN in the kernel's initial user namespace, i.e. real root.
A rootless container runner must solve this without elevating to root.
2. Network Namespaces — the Kernel Primitive
A network namespace (netns) is a kernel isolation unit. Each netns has
its own:
- network interfaces
- routing tables
- iptables/nftables rules
- sockets
Every process belongs to exactly one netns. Containers get their own netns at
creation time, which is why ip link inside a container only shows that
container's interfaces.
Critically, every netns is owned by a user namespace. This ownership determines which capabilities are effective for operations inside that netns:
initial user namespace → owns host netns (real root required)
container user namespace → owns container netns (container user sufficient)
This ownership relationship is what makes rootless networking possible at all.
3. Userspace Network Stacks
The need
A rootless container's netns is isolated — it has no interfaces connected to the outside world by default. Connecting it to the host network requires creating a path across the netns boundary. In rootful mode this is done with kernel veth pairs or macvlan, both of which need root. In rootless mode a different mechanism is needed.
slirp4netns
The original solution. "slirp" is a technique dating from the 1990s (SLIP emulation over serial lines) that implements a complete TCP/IP stack in userspace.
How it works:
- Creates a TAP device in the container netns
- Reads raw frames from that TAP in userspace
- Parses and re-originates the packets as unprivileged host socket calls
- Translates responses back into raw frames and writes them to the TAP
Each packet crosses the kernel↔userspace boundary twice, and slirp4netns terminates and re-originates TCP connections internally. Consequences:
- Source IP addresses are lost (NAT at the slirp layer)
- Higher CPU overhead from full protocol stack re-processing
- Not suitable for workloads that depend on client IP visibility
pasta
The modern replacement, default in Podman 5+. pasta does not reimplement TCP/IP. Instead:
- Creates a TAP device in the container netns (still needs
/dev/net/tun) - Connects it to a host socket
- Uses
splice()to move data between the two kernel buffers directly
The kernel handles the protocol stack on both sides. pasta is just the coordinator that sets up the path and keeps it running. Source IPs are preserved on inbound connections.
Why these are Podman-specific — Docker's daemon model
slirp4netns and pasta are only needed for rootless runtimes.
Docker runs a root daemon (dockerd) that owns all container lifecycle
operations. When Docker creates a container network, it does so in the daemon
process which runs as root — so it can create kernel veth pairs and bridge
interfaces directly, with no userspace TCP/IP stack needed. All Docker
containers, even those running as a non-root user inside the container, benefit
from this root-level network setup.
Podman is daemonless: each podman run is a direct process fork by the
calling user, with no privileged background service. There is nothing running
as root to set up kernel networking on its behalf. This is why Podman needs
slirp4netns or pasta — and why Docker does not.
Rootless Docker (added later, via rootlesskit) faces the same problem and
uses the same solution space.
4. How pasta Uses splice() for Zero-Copy Forwarding
The standard way to forward data between two sockets in userspace:
kernel buffer A → copy to userspace → copy to kernel buffer B
Two copies, two context switches per direction.
splice(2) is a Linux syscall that moves data between two kernel file
descriptors without passing through userspace memory:
kernel buffer A → kernel buffer B (zero userspace involvement)
pasta uses splice() to connect the container-side TAP file descriptor to the
host-side socket file descriptor. The kernel transfers data directly between
the two buffers; pasta wakes up to coordinate but does not touch the payload.
This is not the same as a fully autonomous kernel path (like veth or macvlan in rootful mode, where no userspace process is involved at all). pasta still participates per-batch. But the per-packet copy overhead is eliminated.
5. TUN/TAP Devices
Virtual network device taxonomy
All of the following are network interfaces from the kernel's perspective, but they differ in topology and backing:
| Device | Layer | Backed by | Typical use |
|---|---|---|---|
| TUN | L3 (IP packets) | userspace fd | VPN tunnel (OpenVPN, WireGuard) |
| TAP | L2 (Ethernet frames) | userspace fd | VM networking, VPN bridged mode, pasta/slirp4netns |
| veth | L2 | paired kernel interface | Connect two network namespaces |
| bridge | L2 | set of attached ports | Virtual L2 switch, container networking |
| dummy | L3 | nothing | Loopback-like placeholder, testing |
| loopback | L3 | kernel loopback | 127.0.0.1 within a netns |
| macvlan/ipvlan | L2/L3 | physical NIC | Direct LAN attachment (rootful only) |
TUN and TAP are the same kernel driver (drivers/net/tun.c), distinguished
only by the flag passed to ioctl(TUNSETIFF): IFF_TUN vs IFF_TAP. TUN
delivers raw IP packets; TAP delivers full Ethernet frames including MAC
headers. OpenVPN supports both modes; WireGuard is TUN-only.
veth pairs
A veth (virtual Ethernet) is always created as a pair. A packet sent into one end comes out the other, like a physical patch cable. Neither end terminates in userspace — the kernel handles both.
netns A: veth0 ══════ veth1 :netns B
The canonical use is connecting a container netns to a bridge in the host
netns: one veth end lives in the container (usually named eth0), the other
lives in the host and is attached to a bridge.
veth differs from TAP in that there is no userspace fd involved — it is a purely kernel-level path and carries no per-packet userspace overhead.
bridge
A bridge is a virtual L2 switch. You attach interfaces to it; it learns MAC addresses and forwards frames between ports accordingly. It does not originate or terminate traffic itself.
bridge0
/ | \
eth0 veth0 veth2
| |
container1 container2
This is the standard container networking primitive: each container gets a veth pair, one end in the container netns, the other attached to a shared bridge that has a host NIC or NAT rule for outbound access.
TAP and bridge are often used together for VM networking: a TAP device (backed by QEMU's fd) is attached to a bridge alongside a physical NIC, giving the VM a presence on the LAN.
What /dev/net/tun is
/dev/net/tun is a character device (major 10, minor 200). It is the
factory for creating TUN and TAP virtual network interfaces from userspace:
- TUN — layer 3 (IP packets). Used by OpenVPN, WireGuard.
- TAP — layer 2 (Ethernet frames). Used by pasta, slirp4netns, VM hypervisors.
The open() + ioctl(TUNSETIFF) permission model
Creating a TUN interface is a two-step kernel interaction:
// Step 1: open the factory device
int fd = open("/dev/net/tun", O_RDWR);
// Step 2: allocate an interface inside the current netns
struct ifreq ifr = { .ifr_flags = IFF_TUN };
ioctl(fd, TUNSETIFF, &ifr);
Step 1 uses standard Unix file permissions. The device node is
crw-rw-rw- — world read/write — so any user can open it. The device node
is intentionally not the security boundary.
Step 2 is where the kernel enforces access control. The check is:
Does the calling process hold
CAP_NET_ADMINin the user namespace that owns the target network namespace?
This is the key: the capability is checked against the owning user namespace of the netns, not the initial user namespace globally.
Why capability checks are namespace-scoped
Because a rootless container's netns is owned by the container's user
namespace (not the host's), CAP_NET_ADMIN granted inside that user namespace
is sufficient for all TUN/interface operations scoped to that netns. The
kernel does not require initial-namespace root.
This is also why network: host breaks rootless VPN containers: host netns is
owned by the initial user namespace, so the capability check requires real
root.
6. Running VPN Servers Rootless
Why privileged: true was thought necessary
Two independent issues fed this misconception:
A crun device cgroup bug (fixed in crun 1.13.0, October 2024)
crun's device cgroup allowlist setup was too restrictive for character devices
in rootless mode. Even with CAP_NET_ADMIN and --device /dev/net/tun
correctly specified, the ioctl was blocked. The fix landed in crun 1.13.0.
Users on affected versions observed that privileged: true was the only
workaround, and that assumption propagated into documentation and compose
templates.
iptables-legacy not working in user namespaces
The legacy iptables binary uses the xtables kernel interface, which requires
privileges in the initial user namespace. It fails with:
can't initialize iptables table `nat': Permission denied (you must be root)
This is a separate and permanent limitation — not a bug. The fix is to use
iptables-nft (backed by nftables) which operates correctly within a
user-namespace-owned netns.
The actual minimal requirements
cap_add:
- NET_ADMIN
devices:
- /dev/net/tun:/dev/net/tun
sysctls:
- net.ipv4.ip_forward=1
No privileged: true needed on crun ≥ 1.13.
Why iptables-legacy fails but iptables-nft works
iptables-legacy speaks to the kernel via setsockopt() on a raw socket with
IPT_SO_SET_REPLACE. This path requires CAP_NET_ADMIN in the initial user
namespace.
iptables-nft is a frontend over nftables, which uses the nfnetlink netlink
family. nftables was designed with network namespaces in mind: operations are
scoped to the calling process's netns, and CAP_NET_ADMIN in the owning user
namespace is sufficient.
The kylemanna/openvpn image ships both binaries. Swapping the startup
command from iptables to iptables-nft is the only change required.
iptables persistence and split rule stores
iptables-save and iptables-restore are not persistence mechanisms —
they dump and reload rules to/from a flat text file. Actual persistence
requires something to call iptables-restore at boot (e.g. the
iptables-persistent package, or a systemd unit with ExecStart).
The two backends maintain separate, invisible-to-each-other rule stores:
iptables-legacy-save → dumps x_tables rules
iptables-nft-save → dumps nf_tables rules (disjoint set)
nft list ruleset → dumps nf_tables rules (same as above)
Rules written via iptables-legacy are not visible to iptables-nft and vice
versa. On systems with mixed tooling — Docker using legacy, application using
nft — this produces silent rule gaps. Always use one backend consistently, and
verify with iptables-nft -L or nft list ruleset rather than iptables -L
when using the nft path.
nftables rules are also not persisted by default — they live in kernel
memory and are lost on reboot. The nftables.service systemd unit (shipped
with the nftables package on Debian/Ubuntu and RHEL/Fedora) loads
/etc/nftables.conf at boot; enabling it is all that is needed:
nft list ruleset > /etc/nftables.conf
systemctl enable nftables.service
On Alpine there is no built-in persistence unit; a local init script is required. In container contexts rules are typically re-applied on each startup via the compose command, so host persistence is rarely relevant.
7. The network: host Boundary Collapse
Using network: host in a rootless container instructs the runtime to skip
creating a private netns and attach the container directly to the host network
namespace.
The consequence for capability checks:
With private netns:
container user namespace owns container netns
→ CAP_NET_ADMIN inside container user namespace is sufficient
With network: host:
host netns is owned by initial user namespace (real root)
→ CAP_NET_ADMIN inside container user namespace is NOT sufficient
→ kernel returns EPERM
All rootless networking capabilities — TUN creation, iptables-nft rules, route
management — stop working under network: host. The isolation boundary that
makes rootless possible is eliminated.
Capability check matrix by network mode
| Operation | Netns touched | Private netns (rootless) | network: host (rootless) |
|---|---|---|---|
ip tuntap add tun0 |
container netns | ✓ | ✗ |
ip addr add / ip route add |
container netns | ✓ | ✗ |
ip link set eth0 up |
container netns (veth end) | ✓ | ✗ |
iptables-nft -t nat … |
container netns | ✓ | ✗ |
iptables-legacy -t nat … |
any | ✗ | ✗ |
ip link set eth0 on real host NIC |
host netns | ✗ | ✗ |
8. Kubernetes Equivalents
securityContext
securityContext is the Kubernetes API surface for the same capability and
privilege settings configured manually in compose files:
securityContext:
capabilities:
add: ["NET_ADMIN"]
privileged: false
Kubernetes passes this to the CRI (containerd or CRI-O), which passes it to
the OCI runtime (crun or runc), which performs the same namespace and cgroup
setup as podman run --cap-add NET_ADMIN. Same kernel mechanism, different
entry point.
Device plugins
Device plugins solve a scheduling problem: some nodes have a device available
and some do not. Without a plugin, there is no way to express "this pod needs
/dev/net/tun to be present and accessible" as a schedulable resource.
The flow:
- A device plugin DaemonSet runs on each node and registers a resource with
kubelet: e.g.
vendor.com/tun: 1 - A pod requests the resource:
resources:
limits:
vendor.com/tun: "1"
- The scheduler places the pod only on nodes where the resource is available
- Kubelet queries the plugin for the device path and mount instructions
- Kubelet passes those to the CRI, which passes them to crun/runc
The end result is the same kernel-level device access — the value is in making it visible to the scheduler and auditable in cluster resource accounting.
CNI vs device plugins
These are parallel, non-overlapping concerns:
| CNI | Device plugin | |
|---|---|---|
| Responsibility | Network namespace wiring | Hardware/device resource allocation |
| When it runs | After pod netns is created | During pod scheduling and startup |
| Example | Calico assigns pod IP and routes | GPU plugin exposes /dev/nvidia0 |
CNI plugins themselves often use /dev/net/tun internally (Flannel VXLAN,
WireGuard-based CNIs) — but that is the CNI plugin running with cluster-level
privileges, entirely separate from workload pod permissions.
References
Kernel and namespaces
- Linux kernel docs — Network namespaces
- Linux kernel docs — TUN/TAP driver
- user_namespaces(7) — Linux man page
- network_namespaces(7) — Linux man page
- capabilities(7) — Linux man page
- splice(2) — Linux man page
Rootless containers and Podman networking
- Podman rootless tutorial
- Podman basic networking guide
- Podman rootless limitations
- pasta — fast, zero-copy network namespace tool
- slirp4netns
iptables and nftables
- nftables wiki — Moving from iptables to nftables
- nftables wiki — Netfilter and user namespaces
- iptables-nft(8) — Linux man page
- nft(8) — Linux man page
VPN and rootless containers
- crun 1.13.0 release notes — device cgroup fix
- OpenVPN unprivileged user wiki
- Fedora discussion — OpenVPN /dev/net/tun Operation not permitted