Skip to content

Resolve BuilderNet endpoints at runtime in flashbox-l1 - #216

Open
MoeMahhouk wants to merge 10 commits into
mainfrom
moe/flashbox-l1-dynamic-egress
Open

MoeMahhouk wants to merge 10 commits into
mainfrom
moe/flashbox-l1-dynamic-egress

Conversation

@MoeMahhouk

@MoeMahhouk MoeMahhouk commented Sep 16, 2026 •

Copy link
Copy Markdown
Member

Resolve BuilderNet endpoints at runtime in flashbox-l1

Replaces the hardcoded BuilderNet IPs in firewall-config, toggle-config and the container /etc/hosts with runtime resolution of rpc.buildernet.org and direct-{us,eu,ap}.buildernet.org. Flashbots Protect endpoints stay static. Supersedes #214.

What the image does now

egress-resolver (flashbots/egress-resolver v0.1.5, built from tag in mkosi.build) runs every minute as two oneshot units:

  • egress-resolver.service (resolve): uid egress-resolver, no capabilities, RestrictAddressFamilies=AF_INET AF_UNIX. Queries the names in /etc/bob/egress-resolver.toml over DNS-over-TLS to 1.1.1.1 / 1.0.0.1, accepts only DNSSEC-validated (AD) answers whose A records belong to the queried name or its CNAME chain, rejects non-global addresses, 30 s deadline with every answer recorded as it arrives. Writes /run/egress-resolver/resolve/answers.json.
  • egress-resolver-apply.service (apply): root with CapabilityBoundingSet=CAP_NET_ADMIN and RestrictAddressFamilies=AF_NETLINK AF_UNIX (no IP sockets). Takes the toggle lock, reads the two dynamic chains back from the kernel (rule comments carry the hostname, so the kernel is the only policy state), computes the new policy, rewrites DYN_BNET_PRODUCTION_OUT (ACCEPT tcp/443 to current addresses) and DYN_BNET_MAINTENANCE_OUT (DROP every address seen since boot) atomically with iptables-restore -n, reads them back and requires an exact match, rewrites a chain that holds anything but one rule per address, renders /run/flashbox-endpoints/hosts (ro bind mount, container /etc/hosts symlink), sweeps conntrack by destination for flows not allowed in the current mode, and writes status.json plus a node-exporter textfile. answers.json is opened without following symlinks, ignored when missing, corrupt, older than 180 s or already installed (its generation time equals applied_answers_uptime_secs, which advances only when a firewall update succeeds and is carried forward in status.json, so a failed update is retried from the same file), and last_fresh_uptime_secs per endpoint records when each name last had a fresh answer.

Firewall: the only DNS path in production is the resolver's own uid-scoped DoT rule (first in ALWAYS_OUT, targets derived from the TOML via egress-resolver print-servers). A uid-scoped loopback:53 DROP for searcher in PRODUCTION_OUT closes the resolved-stub cache leak found during testing. No native nft rules, no ipset, no new packages.

Toggle: refuses production unless the timer is active, status.json is from this boot and under 5 min old, apply_ok && hosts_ok, rpc.buildernet.org has an address, and the production chain is non-empty. Warns (does not refuse) when addresses are last-known-good, and says how long ago the last fresh answer was. Leaving production sweeps every address in the maintenance chain.

Observability: node-exporter textfile collector, egress_resolver_* kept by the scrape config, recording rules flashbox:egress_resolver_{is_running,apply_ok,hosts_ok,conntrack_ok,required_endpoints_resolved,required_endpoints_fresh,production_addresses,max_fresh_age_seconds,endpoint_has_addresses} with or on() vector(0) fallbacks. Their promtool unit tests live in modules/flashbox/observability/tests/ and run in CI (.github/workflows/observability-rules.yaml, promtool pinned to the trixie Prometheus generation).

Behaviour changes for searchers

  • rpc.buildernet.org maps to what BuilderNet's DNS returns for it (today the two US nodes), not to every node. Use direct-<region>.buildernet.org for regional pinning.
  • Per-node names (fb-*.nodes.buildernet.org) are no longer resolvable from the container in production.
  • The container starts only once every BuilderNet name has an authenticated answer this boot and the firewall holds the maintenance blocklist, during a DNS-over-TLS outage at boot it waits until resolution succeeds.

Guarantees and known limits

  • Production: only addresses currently in DNS for the four names are reachable on 443. With resolution healthy, a retired address is removed, and its flows killed, within one resolver interval after the resolver's cache expires, so typically the record TTL of 300 s plus 60 s. A failed resolve or apply cycle adds one interval.
  • Maintenance: blocks the BuilderNet addresses observed during this boot. An address not yet resolved (first run pending, or a newly published node until the next run) is reachable through the general HTTPS rule until the resolver sees it. Closing this needs BuilderNet-side authorisation of the searcher's mode.
  • DNS outage: last-known-good addresses stay installed (a stale allowlist cannot gain addresses); required_fresh, last_fresh_uptime_secs and flashbox:egress_resolver_max_fresh_age_seconds say so and for how long.

BuilderNet contract (to announce)

The four names are the contract; publish a node's A record before it serves /bob; keep buildernet.org DNSSEC-signed and, if a name ever becomes a CNAME, every zone on the chain signed (an unsigned hop means no validated answer and the image keeps last-known-good addresses for that name until reboot); announce changes.

Verification

  • egress-resolver: 62 unit/e2e tests, clippy -D warnings; e2e against real iptables 1.8.11 / conntrack 1.4.8 in a trixie container (fresh apply, unchanged, retirement in production, stale and already-applied answers, dirty chain rewrite, symlinked answers refused, firewall re-init repopulation, toggle gate). The v0.1.4 test runs three applies on an unchanged answers file and fails against the v0.1.3 logic; the v0.1.5 test re-runs a failed firewall update from the same file and fails against the v0.1.4 logic.
  • Image on QEMU (dev build, v0.1.3 and v0.1.4): both units succeed each cycle with the stated sandboxes; chains and hosts populated; iptables-save clean; steady-state runs report unchanged with no dirty rewrites; forced last-known-good cycle (DoT rules removed) keeps the addresses, exposes the age in status, metrics and the toggle warning, and recovers on the next fresh answer; three applies on an unchanged answers file all stay rejected with the marker intact; toggle cycle with gate and teardown; in-container production probe: BuilderNet reachable by name, 1.1.1.1:443/80 and libc DNS blocked, cached-stub line blocked, only the resolver's DoT flows in conntrack; live Prometheus against the real scrape yields exactly one labelled series per fallback rule.
  • Recording rules: promtool check rules and promtool test rules (healthy, unhealthy and missing input shapes) pass locally with promtool 2.53.4 and in CI.
  • Integration tests added: test_egress_resolver_hosts, test_egress_resolver_production_cycle (GCP run pending).

Replace the hardcoded BuilderNet IPs (firewall-config, toggle-config and
the container's /etc/hosts) with runtime resolution of the public names
rpc.buildernet.org and direct-{us,eu,ap}.buildernet.org, so BuilderNet
can move nodes without an image release while the firewall stays a
static, auditable iptables ruleset.

- Build egress-resolver v0.1.1 (flashbots/egress-resolver) and run it as a
  systemd oneshot every 60s: the resolve step queries the names over
  DNS-over-TLS to 1.1.1.1/1.0.0.1 as an unprivileged uid and accepts only
  DNSSEC-validated global unicast IPv4 answers; the apply step (root,
  CAP_NET_ADMIN only) refills two iptables chains atomically, renders the
  container hosts file from the same addresses, sweeps conntrack and
  writes status.json plus a Prometheus textfile.
- firewall-config: create DYN_BNET_PRODUCTION_OUT (accept tcp/443 to the
  current addresses, jumped from PRODUCTION_OUT) and
  DYN_BNET_MAINTENANCE_OUT (drop every address seen since boot, jumped
  from MAINTENANCE_OUT before the accept-all rules); allow only the
  egress-resolver uid to reach the two resolvers on 853, as the first
  rules of ALWAYS_OUT; drop the searcher uid's loopback DNS in production
  so systemd-resolved's cache cannot answer the container either.
- toggle: add image hook seams; toggle-config refuses production unless
  egress-resolver completed a run within 5 minutes of this boot and
  rpc.buildernet.org has an address, and kills flows to the resolved
  addresses when leaving production. PRODUCTION_ENDPOINTS keeps only the
  static Flashbots Protect IPs.
- Container: start podman with --no-hosts and bind-mount
  /run/flashbox-endpoints read-only; /etc/hosts is a symlink into it, so
  atomic replacements on the host are visible immediately. Per-node
  fb-*.nodes.buildernet.org names are no longer resolvable in production.
- Observability: enable node-exporter's textfile collector for
  /run/egress-resolver/metrics, keep egress_resolver_* in the node scrape
  job, add flashbox:egress_resolver_* recording rules.
- Document the mechanism in the flashbox-l1 readme and add integration
  tests for the hosts file and the production egress cycle.
- Split the resolver into two oneshot units so the sandboxes differ:
  egress-resolver.service runs the DNS step as the egress-resolver uid
  with no capabilities and RestrictAddressFamilies=AF_INET AF_UNIX;
  egress-resolver-apply.service runs the firewall step with
  CAP_NET_ADMIN and RestrictAddressFamilies=AF_NETLINK AF_UNIX, so its
  lack of network access is enforced rather than assumed. Both use
  TimeoutStartSec (RuntimeMaxSec has no effect on oneshot units) and are
  PartOf=searcher-firewall.service, so a firewall re-initialisation,
  which recreates the dynamic chains empty, re-runs them immediately.
  The timer and the container unit now reference the apply unit.
- firewall-config derives the resolver allowlist from
  egress-resolver.toml via `egress-resolver print-servers`;
  mkosi.postinst runs `check-config` so a bad configuration fails the
  build rather than the boot.
- toggle-config: leaving production kills flows to every address in the
  maintenance chain (a superset of the production chain, covering a
  retired address whose earlier sweep failed); entering production also
  requires the last run to have installed its policy (apply_ok,
  hosts_ok) and warns when the addresses are last-known-good.
- Observability: match node-exporter's full-path file label for the
  textfile mtime, default absent series to 0, and export apply/hosts/
  conntrack health, required-endpoint freshness and per-endpoint address
  presence as flashbox:* recording rules.
- Readme: document that rpc.buildernet.org now maps to its own DNS
  answers (regional pinning via direct-*), that a firewall restart
  resets the dynamic chains, and the two-unit privilege split.
- Integration test: assert that a name cached by the host resolver
  during maintenance does not resolve from the container in production.
- Pin egress-resolver v0.1.2.
…tely

The production image has no guest-OS access, so nothing can restart
searcher-firewall.service on a running box. State what the image does
instead of describing a manual restart.
…ness age in toggle

- pin egress-resolver v0.1.3 (conntrack exit-code semantics, answers kept
  when the resolve deadline fires mid-batch, dirty-chain rewrite, answers
  applied once, per-endpoint last_fresh_uptime_secs)
- recording rules: `or vector(0)` kept an unlabelled 0 next to the labelled
  scraped value on a healthy VM, since the two sides never share a label
  set; use `or on() vector(0)` so the fallback only appears when nothing was
  scraped. Add flashbox:egress_resolver_max_fresh_age_seconds and promtool
  unit tests (healthy, unhealthy, missing) under observability/tests
- toggle: the last-known-good warning now says how long ago the last fresh
  answer for each required endpoint was
- readme: state the maintenance guarantee precisely (addresses observed this
  boot), the container start dependency on the first apply run, and what
  the image needs from BuilderNet's DNS (publish before serving, DNSSEC on
  every CNAME hop)
- pin egress-resolver v0.1.4: the already-applied marker for answers.json
  was lost after one rejected run, so an unchanged file under the age limit
  counted as fresh again on the third apply and reset the freshness
  bookkeeping; the marker is now carried forward in status.json
- add .github/workflows/observability-rules.yaml: promtool check rules and
  test rules for modules/flashbox/observability on pull requests and pushes
  to main that touch the module, pinned to prom/prometheus:v2.53.4 (the
  Prometheus generation Debian trixie ships)
- pin egress-resolver v0.1.5: the applied-answers marker advanced as soon
  as the answers file was usable, so a failed firewall update followed by a
  silent resolve step lost its retry; the marker now advances only when the
  update succeeded and the same file is retried on the next run
- observability-rules workflow: note that it is path-filtered and must not
  become a required check without an unconditional no-op job
- readme: the pickup time for a new BuilderNet address is stated as the
  expected delay with resolution healthy; a failed cycle adds an interval
[resolver]
# Recursive resolvers, tried in order. Must match EGRESS_RESOLVER_DNS in
# firewall-config, which is the only egress the egress-resolver uid may use.
servers = ["1.1.1.1", "1.0.0.1"]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We could consider running our own

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yes, it is configurable here such that we can do such a thing in the future

…d lock

- pin egress-resolver v0.1.6: apply takes the toggle lock before reading
  the clock, the previous status and the answers file, and rejects any
  answers generation not newer than the last installed one. Closes the
  interleavings found by fuzzing concurrent apply invocations; not reachable
  through the image's single systemd unit, but the binary no longer depends
  on that.
- init-firewall.sh takes /etc/searcher-network.lock for its whole run, the
  lock toggle and egress-resolver already hold while they touch the
  runtime-populated chains, so a flush can never interleave with an update.
  A re-run after boot still resets those chains; nothing in the image does
  that.
- readme: a firewall re-initialisation restarts the observed-address
  history, and all writers of the dynamic chains hold the same lock.
@MoeMahhouk
MoeMahhouk marked this pull request as ready for review September 21, 2026 14:24
@MoeMahhouk
MoeMahhouk requested review from a team as code owners September 21, 2026 14:24

@ameba23 ameba23 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not had a really deep look but i don't see any red flags, and +1 for the high level idea.

This effectively hands more control to the buildernet.org domain owner to define what the allowed buildernet endpoints are rather than having them auditable as part of the image. I'd say this trade-off is worth it because these IP addresses anyway don't tell us anything about what the builder runs, and an auditor can still independently resolve these names at a given time. Requiring an image update every time they change is impractical.

I think the most important thing is that last known good addresses are available when DNS fails.

@alexhulbert alexhulbert left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I like the idea of using DoT/DNSSEC here and writing to the hosts file, but this feels like a lot of moving parts and custom rust code for setting a few hostnames, especially since all of this is in the critical path in terms of security. Specifically, the iptables comment stuff feels pretty hacky and I think we could probably accomplish the same or more security with significantly less code.

I'm happy to reimplement this by writing a script that just calls systemd-resolved from a service on a systemd timer. Systemd-resolved already takes care of DoT and DNSSEC and we wouldn't have to trust Cloudflare as much as the current impl. The replacement egress resolver would just need to check the result of systemd-resolved to make sure it's authenticated and then write the hosts file/firewall.

Is there any particular reason you don't want to just put the state in a file? I can't think of any scenario where an attacker could hack that but not this implementation, assuming the file is owned by root and written atomically. I think I could probably get something just as robust with the same security guarantees in a couple hundred lines total.

Also, if I'm not mistaken, because the container can start even when the dynamic firewall rules haven't been setup, the maintenance blacklist could be empty and a searcher would be able to get orderflow in maintenance mode. If cloudflare DNS was down for a few minutes while the image booted up, then the orderflow would be super easy to leak out of the image.

@MoeMahhouk

MoeMahhouk commented Sep 29, 2026 •

Copy link
Copy Markdown
Member Author

I like the idea of using DoT/DNSSEC here and writing to the hosts file, but this feels like a lot of moving parts and custom rust code for setting a few hostnames, especially since all of this is in the critical path in terms of security. Specifically, the iptables comment stuff feels pretty hacky and I think we could probably accomplish the same or more security with significantly less code.

The custom rust code (egress-resolver) is an attempt to replace implementing these parts in several long bash scripts that are more brittle. See #214, it is the ground work of this PR and sets the idea of DoT/DNSSEC, but to do so it has several places I commented there that would either break or increase the design complexity. Most of the code is not the DNS part (~370 lines), it is the apply side: atomic rewrite of the two chains, reading them back to verify, last-known-good handling, killing flows that are no longer allowed, locking against toggle, and the status/metrics that toggle and monitoring gate on. Those are the parts where the review rounds and fuzzy testing found the subtle bugs, and they come with ~1.7k lines of tests now.

In regards the iptables comment stuff: the comments are not just informative. Each rule carries the hostname it was resolved from, and that is how the tool rebuilds its per-name view from the live ruleset every run without keeping its own database. We could move that into a file instead, but then there are two sources of truth (the file and the kernel) that can disagree after a partial restore or a flush, and we would need the same read-back/reconcile logic anyway. I don't consider it a blocker either way, happy to iterate on it if you see a concrete problem with it

I'm happy to reimplement this by writing a script that just calls systemd-resolved from a service on a systemd timer. Systemd-resolved already takes care of DoT and DNSSEC and we wouldn't have to trust Cloudflare as much as the current impl. The replacement egress resolver would just need to check the result of systemd-resolved to make sure it's authenticated and then write the hosts file/firewall.

Agreed on the trust point: today we trust Cloudflare's AD bit, and validating DNSSEC locally would turn Cloudflare into just a transport. That is already on the follow-up list. But note that systemd-resolved would replace only the resolve step (the DoT client), not the apply step that writes the firewall and the hosts file and everything that comes with it. So the shape I would go for is to swap the resolve step for resolvectl (systemd 257 in trixie has resolvectl query --json, so the authenticated flag is easy to check) and keep the rest as is, rather than a rewrite.

Things to keep in mind if you take a stab at it: Debian ships resolved with DNSSEC=no and DNSOverTLS=no, so it needs configuring, resolved also serves all the host's DNS (apt etc. in maintenance), so it has to run in allow-downgrade mode and the per-query "authenticated" check has to be enforced for the BuilderNet names (fail closed for those, everything else keeps working). And to your question about the container: that split already exists today, the container resolves through resolved's stub in maintenance and is blocked from it in production by a uid rule, and only the resolver's uid may talk to the DoT servers in production. That would stay the same. We are also trying to keep the security guarantees as close as possible to before while providing a way to dynamically resolve those production endpoints without rebuilding the image every time, so whatever replaces it needs to keep the toggle gate, the last-known-good behaviour, the flow teardown and the status/metrics.

Is there any particular reason you don't want to just put the state in a file? I can't think of any scenario where an attacker could hack that but not this implementation, assuming the file is owned by root and written atomically.

If you mean the resolved addresses per name: security-wise you are right, a root-owned file in /run written atomically is not easier to attack than the kernel. The reason is consistency, not security. The kernel is what actually enforces the policy, so the tool reads the policy back from the kernel and derives the hosts file from it, that way the hosts file can never contain an address the firewall does not allow, and iptables-save stays a complete audit of what is enforced. Since the review rounds there is also a small bookkeeping file (/run/egress-resolver/status.json, root-owned, atomic) that carries the freshness markers between runs. The mode itself still lives in /etc/searcher-network.state with the flock from toggle, which egress-resolver and now init-firewall.sh share. Not pretty, but until the full architecture clean-up I used what already exists.

I think I could probably get something just as robust with the same security guarantees in a couple hundred lines total.

If you think you could achieve the full idea e2e in a couple hundred lines, please go ahead, but see above in regards the requirements and the guarantees that need to stay intact to avoid security regressions caused by missing the product's security requirements. #214 was the smaller version and the review of it is where most of these requirements came from.

Also, if I'm not mistaken, because the container can start even when the dynamic firewall rules haven't been setup, the maintenance blacklist could be empty and a searcher would be able to get orderflow in maintenance mode. If cloudflare DNS was down for a few minutes while the image booted up, then the orderflow would be super easy to leak out of the image.

Correct! The container waits for the first apply run to finish (After=), but if both DoT resolvers are unreachable at boot that run finishes "successfully" with empty chains, because there is nothing in the kernel to fall back on yet, and the container starts with an empty maintenance blocklist until the first successful resolution (retried every 60s). The readme states this as the known window for not-yet-observed addresses.

For context, the old static list had the inverse window: it never depended on DNS, but it also never learned a new address, so a BuilderNet node added after the image build was reachable in maintenance for as long as it took to rebuild the image, provided BuilderNet had allowlisted the box (which is the normal process for an onboarded searcher). We already had a regression of that kind, where the BuilderNet endpoints were updated while the static list in the image was out of date.

So the boot window needs closing in this PR. I want to think about the options before picking one (a static fallback for the build-time addresses, gating the container start on the first successful resolution, or something else), since each has a different availability trade-off. Suggestions welcome.

Edit:
Following up on the boot window: went with gating the container start rather than a static fallback, since a static list would again miss endpoints added after the build. The container unit now has an ExecStartPre (egress-resolver-wait-ready) that waits until every BuilderNet name has had an authenticated answer this boot (addresses, or an authenticated NXDOMAIN) and the kernel's maintenance chain holds the result. During a DoT outage at boot the container simply waits; production would be refused in that state anyway. What remains is a node published after the container started, which is blocked once the next resolver run sees it.

…is resolved

The container waited for the first egress-resolver apply run to finish, but
not for it to succeed. With DNS-over-TLS unreachable at boot that run exits 0
with empty chains (there is no last-known-good state yet), so the container
started with an empty maintenance blocklist and could reach BuilderNet over
the general HTTPS rule until the first successful resolution. A searcher can
reboot at will, so a DoT outage made this window easy to hit.

Add egress-resolver-wait-ready as ExecStartPre= of the container unit. It
waits until the last apply run of this boot installed its policy and
published the hosts file, no endpoint is in state "missing" (every name has
had an authenticated answer, addresses or an authenticated NXDOMAIN/NODATA),
and the kernel's maintenance chain holds every address the status lists,
which catches a status file left over from before a firewall
re-initialisation. It has no timeout of its own and the oneshot container
unit has none by default; production would be refused by toggle in that
state anyway. A stop while waiting exits cleanly.

Tested against real status output of egress-resolver v0.1.6 with real
iptables (DoT down at boot, one name failing, authenticated empty answer,
last-known-good later in the boot, firewall re-init with and without an
apply run, corrupted or foreign status) and under systemd 257 (no start
timeout with the manager default lowered to 10 s, release within one poll
interval, clean stop and restart during the wait).
@MoeMahhouk
MoeMahhouk requested a review from alexhulbert October 1, 2026 12:16
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants