Resolve BuilderNet endpoints at runtime in flashbox-l1 - #216
MoeMahhouk wants to merge 10 commits into
Conversation
Replace the hardcoded BuilderNet IPs (firewall-config, toggle-config and
the container's /etc/hosts) with runtime resolution of the public names
rpc.buildernet.org and direct-{us,eu,ap}.buildernet.org, so BuilderNet
can move nodes without an image release while the firewall stays a
static, auditable iptables ruleset.
- Build egress-resolver v0.1.1 (flashbots/egress-resolver) and run it as a
systemd oneshot every 60s: the resolve step queries the names over
DNS-over-TLS to 1.1.1.1/1.0.0.1 as an unprivileged uid and accepts only
DNSSEC-validated global unicast IPv4 answers; the apply step (root,
CAP_NET_ADMIN only) refills two iptables chains atomically, renders the
container hosts file from the same addresses, sweeps conntrack and
writes status.json plus a Prometheus textfile.
- firewall-config: create DYN_BNET_PRODUCTION_OUT (accept tcp/443 to the
current addresses, jumped from PRODUCTION_OUT) and
DYN_BNET_MAINTENANCE_OUT (drop every address seen since boot, jumped
from MAINTENANCE_OUT before the accept-all rules); allow only the
egress-resolver uid to reach the two resolvers on 853, as the first
rules of ALWAYS_OUT; drop the searcher uid's loopback DNS in production
so systemd-resolved's cache cannot answer the container either.
- toggle: add image hook seams; toggle-config refuses production unless
egress-resolver completed a run within 5 minutes of this boot and
rpc.buildernet.org has an address, and kills flows to the resolved
addresses when leaving production. PRODUCTION_ENDPOINTS keeps only the
static Flashbots Protect IPs.
- Container: start podman with --no-hosts and bind-mount
/run/flashbox-endpoints read-only; /etc/hosts is a symlink into it, so
atomic replacements on the host are visible immediately. Per-node
fb-*.nodes.buildernet.org names are no longer resolvable in production.
- Observability: enable node-exporter's textfile collector for
/run/egress-resolver/metrics, keep egress_resolver_* in the node scrape
job, add flashbox:egress_resolver_* recording rules.
- Document the mechanism in the flashbox-l1 readme and add integration
tests for the hosts file and the production egress cycle.
- Split the resolver into two oneshot units so the sandboxes differ: egress-resolver.service runs the DNS step as the egress-resolver uid with no capabilities and RestrictAddressFamilies=AF_INET AF_UNIX; egress-resolver-apply.service runs the firewall step with CAP_NET_ADMIN and RestrictAddressFamilies=AF_NETLINK AF_UNIX, so its lack of network access is enforced rather than assumed. Both use TimeoutStartSec (RuntimeMaxSec has no effect on oneshot units) and are PartOf=searcher-firewall.service, so a firewall re-initialisation, which recreates the dynamic chains empty, re-runs them immediately. The timer and the container unit now reference the apply unit. - firewall-config derives the resolver allowlist from egress-resolver.toml via `egress-resolver print-servers`; mkosi.postinst runs `check-config` so a bad configuration fails the build rather than the boot. - toggle-config: leaving production kills flows to every address in the maintenance chain (a superset of the production chain, covering a retired address whose earlier sweep failed); entering production also requires the last run to have installed its policy (apply_ok, hosts_ok) and warns when the addresses are last-known-good. - Observability: match node-exporter's full-path file label for the textfile mtime, default absent series to 0, and export apply/hosts/ conntrack health, required-endpoint freshness and per-endpoint address presence as flashbox:* recording rules. - Readme: document that rpc.buildernet.org now maps to its own DNS answers (regional pinning via direct-*), that a firewall restart resets the dynamic chains, and the two-unit privilege split. - Integration test: assert that a name cached by the host resolver during maintenance does not resolve from the container in production. - Pin egress-resolver v0.1.2.
…tely The production image has no guest-OS access, so nothing can restart searcher-firewall.service on a running box. State what the image does instead of describing a manual restart.
…ness age in toggle - pin egress-resolver v0.1.3 (conntrack exit-code semantics, answers kept when the resolve deadline fires mid-batch, dirty-chain rewrite, answers applied once, per-endpoint last_fresh_uptime_secs) - recording rules: `or vector(0)` kept an unlabelled 0 next to the labelled scraped value on a healthy VM, since the two sides never share a label set; use `or on() vector(0)` so the fallback only appears when nothing was scraped. Add flashbox:egress_resolver_max_fresh_age_seconds and promtool unit tests (healthy, unhealthy, missing) under observability/tests - toggle: the last-known-good warning now says how long ago the last fresh answer for each required endpoint was - readme: state the maintenance guarantee precisely (addresses observed this boot), the container start dependency on the first apply run, and what the image needs from BuilderNet's DNS (publish before serving, DNSSEC on every CNAME hop)
- pin egress-resolver v0.1.4: the already-applied marker for answers.json was lost after one rejected run, so an unchanged file under the age limit counted as fresh again on the third apply and reset the freshness bookkeeping; the marker is now carried forward in status.json - add .github/workflows/observability-rules.yaml: promtool check rules and test rules for modules/flashbox/observability on pull requests and pushes to main that touch the module, pinned to prom/prometheus:v2.53.4 (the Prometheus generation Debian trixie ships)
- pin egress-resolver v0.1.5: the applied-answers marker advanced as soon as the answers file was usable, so a failed firewall update followed by a silent resolve step lost its retry; the marker now advances only when the update succeeded and the same file is retried on the next run - observability-rules workflow: note that it is path-filtered and must not become a required check without an unconditional no-op job - readme: the pickup time for a new BuilderNet address is stated as the expected delay with resolution healthy; a failed cycle adds an interval
| [resolver] | ||
| # Recursive resolvers, tried in order. Must match EGRESS_RESOLVER_DNS in | ||
| # firewall-config, which is the only egress the egress-resolver uid may use. | ||
| servers = ["1.1.1.1", "1.0.0.1"] |
There was a problem hiding this comment.
We could consider running our own
There was a problem hiding this comment.
yes, it is configurable here such that we can do such a thing in the future
…d lock - pin egress-resolver v0.1.6: apply takes the toggle lock before reading the clock, the previous status and the answers file, and rejects any answers generation not newer than the last installed one. Closes the interleavings found by fuzzing concurrent apply invocations; not reachable through the image's single systemd unit, but the binary no longer depends on that. - init-firewall.sh takes /etc/searcher-network.lock for its whole run, the lock toggle and egress-resolver already hold while they touch the runtime-populated chains, so a flush can never interleave with an update. A re-run after boot still resets those chains; nothing in the image does that. - readme: a firewall re-initialisation restarts the observed-address history, and all writers of the dynamic chains hold the same lock.
ameba23
left a comment
There was a problem hiding this comment.
Not had a really deep look but i don't see any red flags, and +1 for the high level idea.
This effectively hands more control to the buildernet.org domain owner to define what the allowed buildernet endpoints are rather than having them auditable as part of the image. I'd say this trade-off is worth it because these IP addresses anyway don't tell us anything about what the builder runs, and an auditor can still independently resolve these names at a given time. Requiring an image update every time they change is impractical.
I think the most important thing is that last known good addresses are available when DNS fails.
alexhulbert
left a comment
There was a problem hiding this comment.
I like the idea of using DoT/DNSSEC here and writing to the hosts file, but this feels like a lot of moving parts and custom rust code for setting a few hostnames, especially since all of this is in the critical path in terms of security. Specifically, the iptables comment stuff feels pretty hacky and I think we could probably accomplish the same or more security with significantly less code.
I'm happy to reimplement this by writing a script that just calls systemd-resolved from a service on a systemd timer. Systemd-resolved already takes care of DoT and DNSSEC and we wouldn't have to trust Cloudflare as much as the current impl. The replacement egress resolver would just need to check the result of systemd-resolved to make sure it's authenticated and then write the hosts file/firewall.
Is there any particular reason you don't want to just put the state in a file? I can't think of any scenario where an attacker could hack that but not this implementation, assuming the file is owned by root and written atomically. I think I could probably get something just as robust with the same security guarantees in a couple hundred lines total.
Also, if I'm not mistaken, because the container can start even when the dynamic firewall rules haven't been setup, the maintenance blacklist could be empty and a searcher would be able to get orderflow in maintenance mode. If cloudflare DNS was down for a few minutes while the image booted up, then the orderflow would be super easy to leak out of the image.
The custom rust code (egress-resolver) is an attempt to replace implementing these parts in several long bash scripts that are more brittle. See #214, it is the ground work of this PR and sets the idea of DoT/DNSSEC, but to do so it has several places I commented there that would either break or increase the design complexity. Most of the code is not the DNS part (~370 lines), it is the apply side: atomic rewrite of the two chains, reading them back to verify, last-known-good handling, killing flows that are no longer allowed, locking against toggle, and the status/metrics that toggle and monitoring gate on. Those are the parts where the review rounds and fuzzy testing found the subtle bugs, and they come with ~1.7k lines of tests now. In regards the iptables comment stuff: the comments are not just informative. Each rule carries the hostname it was resolved from, and that is how the tool rebuilds its per-name view from the live ruleset every run without keeping its own database. We could move that into a file instead, but then there are two sources of truth (the file and the kernel) that can disagree after a partial restore or a flush, and we would need the same read-back/reconcile logic anyway. I don't consider it a blocker either way, happy to iterate on it if you see a concrete problem with it
Agreed on the trust point: today we trust Cloudflare's AD bit, and validating DNSSEC locally would turn Cloudflare into just a transport. That is already on the follow-up list. But note that systemd-resolved would replace only the resolve step (the DoT client), not the apply step that writes the firewall and the hosts file and everything that comes with it. So the shape I would go for is to swap the resolve step for resolvectl (systemd 257 in trixie has Things to keep in mind if you take a stab at it: Debian ships resolved with DNSSEC=no and DNSOverTLS=no, so it needs configuring, resolved also serves all the host's DNS (apt etc. in maintenance), so it has to run in allow-downgrade mode and the per-query "authenticated" check has to be enforced for the BuilderNet names (fail closed for those, everything else keeps working). And to your question about the container: that split already exists today, the container resolves through resolved's stub in maintenance and is blocked from it in production by a uid rule, and only the resolver's uid may talk to the DoT servers in production. That would stay the same. We are also trying to keep the security guarantees as close as possible to before while providing a way to dynamically resolve those production endpoints without rebuilding the image every time, so whatever replaces it needs to keep the toggle gate, the last-known-good behaviour, the flow teardown and the status/metrics.
If you mean the resolved addresses per name: security-wise you are right, a root-owned file in /run written atomically is not easier to attack than the kernel. The reason is consistency, not security. The kernel is what actually enforces the policy, so the tool reads the policy back from the kernel and derives the hosts file from it, that way the hosts file can never contain an address the firewall does not allow, and
If you think you could achieve the full idea e2e in a couple hundred lines, please go ahead, but see above in regards the requirements and the guarantees that need to stay intact to avoid security regressions caused by missing the product's security requirements. #214 was the smaller version and the review of it is where most of these requirements came from.
Correct! The container waits for the first apply run to finish (After=), but if both DoT resolvers are unreachable at boot that run finishes "successfully" with empty chains, because there is nothing in the kernel to fall back on yet, and the container starts with an empty maintenance blocklist until the first successful resolution (retried every 60s). The readme states this as the known window for not-yet-observed addresses. For context, the old static list had the inverse window: it never depended on DNS, but it also never learned a new address, so a BuilderNet node added after the image build was reachable in maintenance for as long as it took to rebuild the image, provided BuilderNet had allowlisted the box (which is the normal process for an onboarded searcher). We already had a regression of that kind, where the BuilderNet endpoints were updated while the static list in the image was out of date. So the boot window needs closing in this PR. I want to think about the options before picking one (a static fallback for the build-time addresses, gating the container start on the first successful resolution, or something else), since each has a different availability trade-off. Suggestions welcome. Edit: |
…is resolved The container waited for the first egress-resolver apply run to finish, but not for it to succeed. With DNS-over-TLS unreachable at boot that run exits 0 with empty chains (there is no last-known-good state yet), so the container started with an empty maintenance blocklist and could reach BuilderNet over the general HTTPS rule until the first successful resolution. A searcher can reboot at will, so a DoT outage made this window easy to hit. Add egress-resolver-wait-ready as ExecStartPre= of the container unit. It waits until the last apply run of this boot installed its policy and published the hosts file, no endpoint is in state "missing" (every name has had an authenticated answer, addresses or an authenticated NXDOMAIN/NODATA), and the kernel's maintenance chain holds every address the status lists, which catches a status file left over from before a firewall re-initialisation. It has no timeout of its own and the oneshot container unit has none by default; production would be refused by toggle in that state anyway. A stop while waiting exits cleanly. Tested against real status output of egress-resolver v0.1.6 with real iptables (DoT down at boot, one name failing, authenticated empty answer, last-known-good later in the boot, firewall re-init with and without an apply run, corrupted or foreign status) and under systemd 257 (no start timeout with the manager default lowered to 10 s, release within one poll interval, clean stop and restart during the wait).
Resolve BuilderNet endpoints at runtime in flashbox-l1
Replaces the hardcoded BuilderNet IPs in
firewall-config,toggle-configand the container/etc/hostswith runtime resolution ofrpc.buildernet.organddirect-{us,eu,ap}.buildernet.org. Flashbots Protect endpoints stay static. Supersedes #214.What the image does now
egress-resolver(flashbots/egress-resolver v0.1.5, built from tag inmkosi.build) runs every minute as two oneshot units:egress-resolver.service(resolve): uidegress-resolver, no capabilities,RestrictAddressFamilies=AF_INET AF_UNIX. Queries the names in/etc/bob/egress-resolver.tomlover DNS-over-TLS to 1.1.1.1 / 1.0.0.1, accepts only DNSSEC-validated (AD) answers whose A records belong to the queried name or its CNAME chain, rejects non-global addresses, 30 s deadline with every answer recorded as it arrives. Writes/run/egress-resolver/resolve/answers.json.egress-resolver-apply.service(apply): root withCapabilityBoundingSet=CAP_NET_ADMINandRestrictAddressFamilies=AF_NETLINK AF_UNIX(no IP sockets). Takes the toggle lock, reads the two dynamic chains back from the kernel (rule comments carry the hostname, so the kernel is the only policy state), computes the new policy, rewritesDYN_BNET_PRODUCTION_OUT(ACCEPT tcp/443 to current addresses) andDYN_BNET_MAINTENANCE_OUT(DROP every address seen since boot) atomically withiptables-restore -n, reads them back and requires an exact match, rewrites a chain that holds anything but one rule per address, renders/run/flashbox-endpoints/hosts(ro bind mount, container/etc/hostssymlink), sweeps conntrack by destination for flows not allowed in the current mode, and writesstatus.jsonplus a node-exporter textfile.answers.jsonis opened without following symlinks, ignored when missing, corrupt, older than 180 s or already installed (its generation time equals applied_answers_uptime_secs, which advances only when a firewall update succeeds and is carried forward in status.json, so a failed update is retried from the same file), andlast_fresh_uptime_secsper endpoint records when each name last had a fresh answer.Firewall: the only DNS path in production is the resolver's own uid-scoped DoT rule (first in
ALWAYS_OUT, targets derived from the TOML viaegress-resolver print-servers). A uid-scoped loopback:53 DROP forsearcherinPRODUCTION_OUTcloses the resolved-stub cache leak found during testing. No native nft rules, no ipset, no new packages.Toggle: refuses production unless the timer is active,
status.jsonis from this boot and under 5 min old,apply_ok && hosts_ok,rpc.buildernet.orghas an address, and the production chain is non-empty. Warns (does not refuse) when addresses are last-known-good, and says how long ago the last fresh answer was. Leaving production sweeps every address in the maintenance chain.Observability: node-exporter textfile collector,
egress_resolver_*kept by the scrape config, recording rulesflashbox:egress_resolver_{is_running,apply_ok,hosts_ok,conntrack_ok,required_endpoints_resolved,required_endpoints_fresh,production_addresses,max_fresh_age_seconds,endpoint_has_addresses}withor on() vector(0)fallbacks. Their promtool unit tests live inmodules/flashbox/observability/tests/and run in CI (.github/workflows/observability-rules.yaml, promtool pinned to the trixie Prometheus generation).Behaviour changes for searchers
rpc.buildernet.orgmaps to what BuilderNet's DNS returns for it (today the two US nodes), not to every node. Usedirect-<region>.buildernet.orgfor regional pinning.fb-*.nodes.buildernet.org) are no longer resolvable from the container in production.Guarantees and known limits
required_fresh,last_fresh_uptime_secsandflashbox:egress_resolver_max_fresh_age_secondssay so and for how long.BuilderNet contract (to announce)
The four names are the contract; publish a node's A record before it serves
/bob; keepbuildernet.orgDNSSEC-signed and, if a name ever becomes a CNAME, every zone on the chain signed (an unsigned hop means no validated answer and the image keeps last-known-good addresses for that name until reboot); announce changes.Verification
-D warnings; e2e against real iptables 1.8.11 / conntrack 1.4.8 in a trixie container (fresh apply, unchanged, retirement in production, stale and already-applied answers, dirty chain rewrite, symlinked answers refused, firewall re-init repopulation, toggle gate). The v0.1.4 test runs three applies on an unchanged answers file and fails against the v0.1.3 logic; the v0.1.5 test re-runs a failed firewall update from the same file and fails against the v0.1.4 logic.iptables-saveclean; steady-state runs report unchanged with no dirty rewrites; forced last-known-good cycle (DoT rules removed) keeps the addresses, exposes the age in status, metrics and the toggle warning, and recovers on the next fresh answer; three applies on an unchanged answers file all stay rejected with the marker intact; toggle cycle with gate and teardown; in-container production probe: BuilderNet reachable by name, 1.1.1.1:443/80 and libc DNS blocked, cached-stub line blocked, only the resolver's DoT flows in conntrack; live Prometheus against the real scrape yields exactly one labelled series per fallback rule.promtool check rulesandpromtool test rules(healthy, unhealthy and missing input shapes) pass locally with promtool 2.53.4 and in CI.test_egress_resolver_hosts,test_egress_resolver_production_cycle(GCP run pending).