From 3886d956ccd3d8d611fd190f44a25f1406b3961a Mon Sep 17 00:00:00 2001 From: KoloMl Date: Thu, 30 Jul 2026 12:34:35 +0400 Subject: [PATCH] Add docker-tailscale-route: fix Docker losing internet under Tailscale exit-node Tailscale's exit-node "route all traffic" mode captures Docker's forwarded container traffic into its tunnel and drops it as an unrecognized flow, causing every outbound request from inside a container to hang until timeout. This installs a systemd timer that keeps an ip rule in place routing Docker's bridge subnet around Tailscale's catch-all route instead. Co-Authored-By: Claude Code --- README.md | 90 ++++++++++++++++++++++++++++++++++ docker-tailscale-route.service | 18 +++++++ docker-tailscale-route.sh | 57 +++++++++++++++++++++ docker-tailscale-route.timer | 22 +++++++++ install.sh | 63 ++++++++++++++++++++++++ uninstall.sh | 33 +++++++++++++ 6 files changed, 283 insertions(+) create mode 100644 README.md create mode 100644 docker-tailscale-route.service create mode 100755 docker-tailscale-route.sh create mode 100644 docker-tailscale-route.timer create mode 100755 install.sh create mode 100755 uninstall.sh diff --git a/README.md b/README.md new file mode 100644 index 0000000..e3ea529 --- /dev/null +++ b/README.md @@ -0,0 +1,90 @@ +# ddev-tailscale-docker-fix + +> **Note:** these scripts were written by an LLM (Claude Code), diagnosing +> and fixing an issue live on one specific machine (Fedora + NetworkManager +> + Tailscale + Docker). They worked there, but review them yourself before +> running as root on another machine — network/firewall behavior varies +> across distros, Tailscale versions, and configurations. + +Fixes Docker containers (e.g. ddev) losing internet access when Tailscale +is running with an exit node ("route all traffic") enabled. + +## The problem + +With an exit node and "route all traffic" enabled, Tailscale installs a +catch-all `ip rule` that captures every packet not explicitly exempted +and routes it out via the `tailscale0` interface. Tailscale exempts its +own daemon traffic from that capture with a firewall mark, but has no +way to exempt traffic that only passes *through* the host (forwarded) +rather than originating from it. + +Docker containers' outbound traffic is exactly that: forwarded, not +locally-originated. It gets no exemption, so it's swept into +Tailscale's catch-all route, sent out via `tailscale0`, and dropped by +Tailscale's own anti-spoofing firewall rule (`ts-forward`) because it +isn't a recognized Tailscale-tunneled flow. + +Symptom: every DNS lookup / outbound HTTP request from inside a +container hangs until timeout whenever Tailscale is connected — WP core +update checks, plugin license pings, `composer`/`npm` installs, +anything. The host's own applications aren't affected, since their +traffic is locally-originated and gets Tailscale's exemption mark. + +## The fix + +A higher-priority `ip rule` that routes Docker's private bridge subnet +(default: `172.16.0.0/12`, covering Docker's default address pool) +through the normal routing table instead of Tailscale's catch-all. This +doesn't weaken the VPN for anything else on the host — it only affects +packets forwarded from Docker's own bridge networks. + +The rule is kept in place by a systemd timer that re-checks/re-adds it +every 15s (and 5s after boot), rather than hooked to a network event: +`tailscale0` is a tun device created directly by `tailscaled`, not +through a NetworkManager connection profile, so NetworkManager never +fires dispatcher/udev events we could reliably hook into for it +(confirmed by testing — don't waste time re-trying that route). + +## Install + +``` +sudo ./install.sh +``` + +Optional overrides if Docker's default address pool doesn't apply on +your machine (e.g. you have 15+ concurrent Docker networks and Docker +has fallen back to a different private range): + +``` +sudo ./install.sh --subnet 172.16.0.0/12 --priority 5200 --table main +``` + +## Verify + +``` +ip rule list | grep 172.16.0.0/12 +ddev exec curl -sS -o /dev/null -w '%{time_total}\n' https://api.wordpress.org +``` + +## Uninstall + +``` +sudo ./uninstall.sh +``` + +## Caveats + +- Only covers Docker networks within `172.16.0.0/12`. If you have more + than ~15 concurrent Docker networks on one machine, Docker falls back + to `192.168.0.0/20` blocks and then `10.0.0.0/8` — those aren't + covered by default. Don't just widen the subnet to `192.168.0.0/16` + or `10.0.0.0/8` to compensate: those ranges are also real LAN/VPN + subnet-route space, and blanket-exempting them could route traffic + around the tunnel that's actually supposed to go through it. If you + hit this, either raise `--subnet` to the specific overflow range + Docker actually assigned, or extend the script to enumerate live + `docker network inspect` subnets instead of using a fixed range. +- Assumes `tailscaled`'s own ip rules stay in the 5210+ priority range + observed on the machine this was built on; if a future Tailscale + version changes that, priority 5200 should still safely sit above it + as long as it's below whatever Tailscale uses for its catch-all rule. diff --git a/docker-tailscale-route.service b/docker-tailscale-route.service new file mode 100644 index 0000000..0c7b579 --- /dev/null +++ b/docker-tailscale-route.service @@ -0,0 +1,18 @@ +# Written by an LLM (Claude Code) — review before trusting on a machine +# other than the one this was diagnosed on. See README.md. +# +# Ensures Docker's bridge traffic bypasses Tailscale's exit-node route +# (see docker-tailscale-route.sh for the full explanation of why this is +# needed). Triggered on a schedule by the companion +# docker-tailscale-route.timer unit rather than a network event: on +# Linux, tailscale0 is a tun device created directly by tailscaled, not +# through a NetworkManager connection profile, so NetworkManager never +# fires dispatcher events for it going up/down. + +[Unit] +Description=Ensure Docker bridge traffic bypasses Tailscale's exit-node route +After=network.target + +[Service] +Type=oneshot +ExecStart=/usr/local/sbin/docker-tailscale-route.sh diff --git a/docker-tailscale-route.sh b/docker-tailscale-route.sh new file mode 100755 index 0000000..f659e47 --- /dev/null +++ b/docker-tailscale-route.sh @@ -0,0 +1,57 @@ +#!/bin/bash +# +# Written by an LLM (Claude Code) — review before trusting on a machine +# other than the one this was diagnosed on. See README.md. +# +# Ensures Docker's own bridge traffic bypasses Tailscale's exit-node +# catch-all route. +# +# Background: a machine running Tailscale with an exit node and "route +# all traffic" enabled installs a catch-all `ip rule` that captures +# every packet not explicitly exempted and sends it out via the +# tailscale0 interface. Tailscale exempts its own daemon traffic from +# that capture with a firewall mark, but it has no way to exempt +# traffic that only passes THROUGH the host (forwarded) rather than +# originating from it. +# +# Docker containers' outbound traffic (dev environments, package +# managers, anything containers do that talks to the internet) is +# exactly that: forwarded traffic, not locally-originated. With no +# exemption, it gets swept into Tailscale's catch-all route, sent out +# via tailscale0, and dropped by Tailscale's own anti-spoofing firewall +# rule because it isn't a recognized Tailscale-tunneled flow. Symptom: +# every DNS lookup / outbound HTTP request from inside a container +# hangs until timeout whenever Tailscale is connected. +# +# The fix: keep a higher-priority `ip rule` in place that routes +# Docker's private bridge subnet through the normal routing table (the +# real gateway/NIC) instead of Tailscale's catch-all. This does NOT +# weaken or bypass the VPN for anything else on the host — it only +# affects packets forwarded from Docker's own bridge networks, which +# were never meant to be tunneled through the exit node in the first +# place. +# +# Defaults assume Docker's default address-pool behavior (bridge +# networks allocated from 172.17.0.0/16 through 172.31.0.0/16, all +# within 172.16.0.0/12). If you have more than ~15 concurrent Docker +# networks on a machine, Docker falls back to other private ranges +# (192.168.0.0/20 blocks, then 10.0.0.0/8) that this default won't +# cover — override DOCKER_SUBNET in /etc/default/docker-tailscale-route +# if that applies to you. Widening the range casually isn't advisable: +# 192.168.0.0/16 and 10.0.0.0/8 are also real LAN/VPN-subnet-route +# space, and blanket-exempting them could route traffic around the +# tunnel that's actually supposed to go through it. + +set -euo pipefail + +DOCKER_SUBNET="172.16.0.0/12" +RULE_PRIORITY="5200" +TARGET_TABLE="main" + +# Optional overrides, e.g. if Docker's address pool differs on this +# machine or the priority collides with something else. +[ -f /etc/default/docker-tailscale-route ] && . /etc/default/docker-tailscale-route + +if ! ip rule list | grep -q "from ${DOCKER_SUBNET} lookup ${TARGET_TABLE}"; then + ip rule add from "$DOCKER_SUBNET" lookup "$TARGET_TABLE" priority "$RULE_PRIORITY" +fi diff --git a/docker-tailscale-route.timer b/docker-tailscale-route.timer new file mode 100644 index 0000000..c115257 --- /dev/null +++ b/docker-tailscale-route.timer @@ -0,0 +1,22 @@ +# Written by an LLM (Claude Code) — review before trusting on a machine +# other than the one this was diagnosed on. See README.md. +# +# Runs docker-tailscale-route.service shortly after boot and then every +# 15s afterward, so the ip rule that exempts Docker's bridge traffic from +# Tailscale's exit-node catch-all route gets reinstated quickly after a +# reboot or a Tailscale reconnect (both can wipe/recreate Tailscale's own +# ip rules and leave ours missing in the process). Polling is used +# instead of a network-event hook because NetworkManager doesn't manage +# the tailscale0 interface and never fires dispatcher/udev events we +# could reliably hook into for it (confirmed by testing). + +[Unit] +Description=Periodically ensure Docker bridge traffic bypasses Tailscale's exit-node route + +[Timer] +OnBootSec=5s +OnUnitActiveSec=15s +Unit=docker-tailscale-route.service + +[Install] +WantedBy=timers.target diff --git a/install.sh b/install.sh new file mode 100755 index 0000000..b158e53 --- /dev/null +++ b/install.sh @@ -0,0 +1,63 @@ +#!/bin/bash +# +# Written by an LLM (Claude Code) — review before trusting on a machine +# other than the one this was diagnosed on. See README.md. +# +# Installs the docker-tailscale-route fix: a systemd timer that keeps an +# ip rule in place routing Docker's bridge traffic around Tailscale's +# exit-node catch-all route. See docker-tailscale-route.sh for the full +# explanation of the underlying problem. +# +# Usage: +# sudo ./install.sh [--subnet CIDR] [--priority N] [--table NAME] +# +# The optional flags only need to be passed if Docker's default address +# pool doesn't apply on this machine (see the comment in +# docker-tailscale-route.sh); they're written to +# /etc/default/docker-tailscale-route and sourced by the script at +# runtime. + +set -euo pipefail + +if [ "$(id -u)" -ne 0 ]; then + echo "Must be run as root (sudo ./install.sh)" >&2 + exit 1 +fi + +SCRIPT_DIR="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)" + +SUBNET="" +PRIORITY="" +TABLE="" + +while [ $# -gt 0 ]; do + case "$1" in + --subnet) SUBNET="$2"; shift 2 ;; + --priority) PRIORITY="$2"; shift 2 ;; + --table) TABLE="$2"; shift 2 ;; + *) echo "Unknown argument: $1" >&2; exit 1 ;; + esac +done + +install -o root -g root -m 0755 "$SCRIPT_DIR/docker-tailscale-route.sh" /usr/local/sbin/docker-tailscale-route.sh +install -o root -g root -m 0644 "$SCRIPT_DIR/docker-tailscale-route.service" /etc/systemd/system/docker-tailscale-route.service +install -o root -g root -m 0644 "$SCRIPT_DIR/docker-tailscale-route.timer" /etc/systemd/system/docker-tailscale-route.timer + +if [ -n "$SUBNET" ] || [ -n "$PRIORITY" ] || [ -n "$TABLE" ]; then + { + [ -n "$SUBNET" ] && echo "DOCKER_SUBNET=\"$SUBNET\"" + [ -n "$PRIORITY" ] && echo "RULE_PRIORITY=\"$PRIORITY\"" + [ -n "$TABLE" ] && echo "TARGET_TABLE=\"$TABLE\"" + } > /etc/default/docker-tailscale-route + chown root:root /etc/default/docker-tailscale-route + chmod 0644 /etc/default/docker-tailscale-route +fi + +systemctl daemon-reload +systemctl enable --now docker-tailscale-route.timer + +echo "Installed. Current status:" +systemctl status docker-tailscale-route.timer --no-pager +echo +echo "ip rule list:" +ip rule list diff --git a/uninstall.sh b/uninstall.sh new file mode 100755 index 0000000..2ca8ab6 --- /dev/null +++ b/uninstall.sh @@ -0,0 +1,33 @@ +#!/bin/bash +# +# Written by an LLM (Claude Code) — review before trusting on a machine +# other than the one this was diagnosed on. See README.md. +# +# Removes everything install.sh set up: the timer, service, script, +# optional config override, and the ip rule itself. +# +# Usage: sudo ./uninstall.sh + +set -euo pipefail + +if [ "$(id -u)" -ne 0 ]; then + echo "Must be run as root (sudo ./uninstall.sh)" >&2 + exit 1 +fi + +systemctl disable --now docker-tailscale-route.timer 2>/dev/null || true + +rm -f /etc/systemd/system/docker-tailscale-route.timer +rm -f /etc/systemd/system/docker-tailscale-route.service +rm -f /usr/local/sbin/docker-tailscale-route.sh +rm -f /etc/default/docker-tailscale-route + +systemctl daemon-reload + +DOCKER_SUBNET="172.16.0.0/12" +TARGET_TABLE="main" +if ip rule list | grep -q "from ${DOCKER_SUBNET} lookup ${TARGET_TABLE}"; then + ip rule del from "$DOCKER_SUBNET" lookup "$TARGET_TABLE" +fi + +echo "Uninstalled."