Skip to content

fix(ci): point LG_ENV at -dual-tftp target when staging kernel + ramdisk - #1277

Merged
francoriba merged 4 commits into
libremesh:masterfrom
fcefyn-testbed:fix/librerouter-dual-tftp-env-override
Sep 26, 2026
Merged

francoriba merged 4 commits into
libremesh:masterfrom
fcefyn-testbed:fix/librerouter-dual-tftp-env-override

Conversation

@francoriba

Copy link
Copy Markdown
Contributor

Context

Follow-up to libremesh/libremesh-tests#5 which splits the LibreRouter v1 target YAML into two variants:

  • targets/librerouter_librerouter-v1.yaml - single-image (default)
  • targets/librerouter_librerouter-v1-dual-tftp.yaml - kernel + rootfs uImage

Why the split: the single-image variant is what mesh tests, LibreMesh releases and openwrt-tests healthcheck actually use (initramfs-kernel.bin, only LG_IMAGE set). The previous librerouter YAML required both LG_IMAGE and LG_IMAGE_INITRD unconditionally, breaking every non-dual flow with:

InvalidConfigError: configuration file '.../librerouter_librerouter-v1.yaml'
refers to unknown variable 'LG_IMAGE_INITRD'

Problem in this CI

.github/workflows/build-firmware.yml pre-sets:

LG_ENV=targets/${{ matrix.device }}.yaml

After the libremesh-tests split, that path points at the single-image variant.

When tools/ci/lab_stage_firmware.sh then detects the dual-tftp payload (bin + uimage pair produced by build_image.sh for devices like the LibreRouter v1 on ath79), it exports LG_IMAGE_INITRD — but the single-image YAML does not reference that variable, so labgrid loads only the kernel and boot fails without a rootfs.

Fix

When lab_stage_firmware.sh enters the dual-TFTP branch, override LG_ENV with the -dual-tftp sibling target file so labgrid resolves LG_IMAGE_INITRD and issues both TFTP loads:

DUAL_ENV="targets/${DEVICE}-dual-tftp.yaml"
echo "LG_ENV=$DUAL_ENV" >> "$GITHUB_ENV"

No change for single-image devices (openwrt_one, linksys_e8450, bananapi_bpi-r4, qemu_*): the else branch keeps the pre-set LG_ENV.

Depends on

Follow-up to libremesh/libremesh-tests#5 which split the LibreRouter v1
target YAML into two variants:

- targets/librerouter_librerouter-v1.yaml           (single-image, default)
- targets/librerouter_librerouter-v1-dual-tftp.yaml (kernel + rootfs uImage)

The single-image variant is what mesh tests, LibreMesh releases and
openwrt-tests healthcheck actually use (initramfs-kernel.bin, only
LG_IMAGE set). The previous librerouter YAML required both LG_IMAGE and
LG_IMAGE_INITRD unconditionally, breaking every non-dual flow.

.github/workflows/build-firmware.yml pre-sets LG_ENV to
targets/${device}.yaml, which after the libremesh-tests split points at
the single-image variant. When lab_stage_firmware.sh detects the
dual-tftp payload (bin + uimage pair produced by build_image.sh for
devices like the LibreRouter v1 on ath79), it exports LG_IMAGE_INITRD
but the single-image YAML does not reference that variable, so labgrid
loads only the kernel and boot fails without a rootfs.

Override LG_ENV with the -dual-tftp sibling target file in the dual-TFTP
branch so labgrid resolves LG_IMAGE_INITRD and issues both TFTP loads.
No change for single-image devices (openwrt_one, linksys_e8450,
bananapi_bpi-r4, qemu_*): the else branch keeps the pre-set LG_ENV.
lab_stage_firmware.sh may override LG_ENV to a -dual-tftp variant
(targets/${DEVICE}-dual-tftp.yaml) when staging a dual-TFTP payload.
Those target YAMLs contain:

  RemotePlace:
    name: !template "$LG_PLACE"

When labgrid-client parses the YAML at lock time, it expands the
template — but LG_PLACE is written to GITHUB_ENV *after* the lock call
succeeds (echo "LG_PLACE=$PLACE" >> "$GITHUB_ENV"), so the variable
is undefined and labgrid raises:

  InvalidConfigError: configuration file 'targets/librerouter_v1-dual-tftp.yaml'
  refers to unknown variable 'LG_PLACE'

Fix: declare LG_PLACE as a step-level env var using the matrix value so
it is present in the process environment when labgrid-client runs.
@francoriba
francoriba deployed to physical-lab September 9, 2026 13:27 — with GitHub Actions Active
@francoriba
francoriba deployed to physical-lab September 9, 2026 13:27 — with GitHub Actions Active
@francoriba
francoriba deployed to physical-lab September 9, 2026 13:27 — with GitHub Actions Active
…ches

v7 fetches raw.githubusercontent.com/astral-sh/versions/main/v1/uv.ndjson
exactly once with no retry. Run 34356587946 died on that fetch for two
physical-DUT jobs (librerouter_1, belkin_rt3200_1) with:

    ##[error]Failed to fetch version data: 503 Backend.max_conn reached

before any stage/lock/test step could run. v10.0.1 wraps the manifest
fetch in up to 3 attempts with progressive backoff
(astral-sh/setup-uv#1016), covering exactly this class of failure.

v10's only breaking change is enable-cache: auto, which this workflow
does not set, so the bump is behavior-preserving.
…raw ndjson 503)

setup-uv fetches raw.githubusercontent.com/astral-sh/versions/main/v1/uv.ndjson
to resolve versions and download URLs. That endpoint has returned HTTP 503
(Varnish/Backend.max_conn) throughout runs 34356587946 and 34368726582,
killing every physical-DUT job at 'Install uv' before any stage/lock/test
ran. setup-uv@v10 added retries only for thrown network errors, not for
HTTP error statuses (astral-sh/setup-uv#1016), so 503 is not covered.

The standalone installer (astral.sh/uv/$UV_VERSION/install.sh) resolves
the version straight from releases.astral.sh (mirror) and falls back to
github.com/astral-sh/uv/releases, both of which bypass the ndjson manifest
entirely and are served by Cloudflare, not Fastly/Varnish.

UV_VERSION is pinned to 0.12.11, the same version used by successful jobs
in the affected runs (found in tool-cache log line 'Found uv in tool-cache
for 0.12.11'). Bump deliberately.

@ilario ilario left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You are the master of the continuous deployment, so I approve any change you consider that make sense to do on those workflows :) Feel free to merge yourself when you think think that the code is ready.

@francoriba
francoriba merged commit 6d9f26f into libremesh:master Sep 26, 2026
144 of 154 checks passed

This branch had an error being deployed

1 failed deployment
physical-lab — 00631e71 Deployed Sep 23, 2026 by francoriba via test-firmware (librerouter_1 / 24.10.6) #124
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants