Repository navigation
feat(nightly): add QEMU integration tests for baremetal-sci_usi - #356
Draft
saubhikdattagithub wants to merge 39 commits into
Draft
saubhikdattagithub wants to merge 39 commits into
saubhikdattagithub wants to merge 39 commits into
Conversation
Add the gardenlinux QEMU integration test suite to the nightly workflow, running against the baremetal-sci_usi (amd64) flavor after the build step. Signed-off-by: Saubhik Datta <50126327+saubhikdattagithub@users.noreply.github.com>
The cross-repo call to test_flavor_qemu.yml failed because gl-features-parse cannot resolve the CNAME from gardenlinux-sci's feature tree. Inline the steps and compute CNAME directly from the VERSION and COMMIT files restored from the build cache. Signed-off-by: Saubhik Datta <50126327+saubhikdattagithub@users.noreply.github.com>
The artifact name uses the short 8-character commit hash, not the full SHA stored in the COMMIT file. Signed-off-by: Saubhik Datta <50126327+saubhikdattagithub@users.noreply.github.com>
- Run QEMU integration tests for all 5 testable flavors via matrix - Fix CNAME computation: use short 8-char commit hash (cut -c1-8) - Use fail-fast: false so one flavor failure doesn't cancel others Signed-off-by: Saubhik Datta <50126327+saubhikdattagithub@users.noreply.github.com>
Signed-off-by: Saubhik Datta <50126327+saubhikdattagithub@users.noreply.github.com>
baremetal-sci_pxe produces a .pxe.tar.gz artifact, not a plain .tar.gz. Use conditional logic: untar if .tar.gz, else mv all matching files.
VERSION/COMMIT cache files store the raw inputs ('today', workflow
commit), not the resolved build values. Use flavor-version-data.json
which has the actual version (e.g. 2380.0.0) and commit_id used by
the build system to name the artifacts.
Groups the 4 QEMU integration test jobs under a collapsible "Test" section in the GitHub Actions UI, matching the build/publish pattern. Excludes baremetal-sci_pxe (UKI format not yet supported by test framework — tracked as follow-up).
…EMU tests - _usi/image.cc.tar: hardlink uki→boot.efi and squashfs→root.squashfs so the gardenlinux test framework can locate boot files without breaking existing convert.*~cc.tar scripts that reference the original names - _pxe/image.pxe.tar.gz: include cmdline in the pxe archive (was already created by ukify but missing from the tar output) - test_qemu.yml: re-add baremetal-sci_pxe to the QEMU test matrix
- test_qemu.yml: convert .cc.tar to .pxe.tar.gz when no .pxe.tar.gz exists, so the gardenlinux run_qemu.sh handles USI images as UKI PXE boot (it detects boot.efi inside and chains via iPXE) - _usi/requirements.mod: set uefi=true so the test framework starts QEMU with UEFI firmware (required for UKI boot)
…pecific test gaps - _pxe/image.pxe.tar.gz: read cmdline from /etc/kernel/cmdline instead of hardcoding it, so all cmdline.d fragments (ip=dhcp from _pxe, security=apparmor from gardener) are included in the PXE boot cmdline - test_qemu.yml: switch to matrix include with per-flavor expected_users and deselect_tests fields; sci flavors declare openstack system users as expected and deselect tests that reflect intentional SCI behavior differences: - test_metal_ipmievd_service_enabled: sci preset disables ipmievd - test_ssh_sshguard_iptables_configured: sci has nftables, uses nftables backend - test_ssh_sshguard_iptables_backend_configured: same as above
- _usi/image.cc.tar: write cmdline file to dest_dir so that when cc.tar is repacked as pxe.tar.gz the test framework finds the required vmlinuz/initrd/cmdline/root.squashfs files - _usi/requirements.mod: remove uefi=true — OVMF does not support -boot order=nc for network boot on x86_64; SeaBIOS handles traditional iPXE network boot correctly - _pxe/file.include/etc/kernel/cmdline.d/80-pxe.cfg: add PXE-specific kernel cmdline params (ip=dhcp gl.live=1 gl.ovl=/:tmpfs) so the live initrd switch-root succeeds when reading from /etc/kernel/cmdline
- _usi/requirements.mod: restore uefi=true so the gardenlinux test
framework boots USI images with OVMF, enabling iPXE to chainload
boot.efi (UKI); without UEFI, SeaBIOS-based iPXE cannot execute
EFI applications
- test_qemu.yml: deselect three metal kernel cmdline tests for all sci
flavors (sci_usi, sci_usidev, sci_pxe); sci/00-default.cfg and
sci/10-console.cfg intentionally override metal defaults — they drop
earlyprintk=ttyS0,115200 and console=ttyS0,115200 in favour of the
SCI-specific console configuration:
- test_metal_kernel_cmdline_default: expects earlyprintk=ttyS0,115200
- test_console_configuration_in_cmdline_metal: expects console=ttyS0
- test_console_configuration_in_cmdline_metal_bautrates: expects
console=ttyS0,115200
OVMF/EDK2 on x86_64 cannot network-boot via virtio-net-pci because the device has no UEFI PXE ROM; -boot once=n,order=c only works with SeaBIOS. UKI boot (chain-loading boot.efi) therefore fails with "BdsDxe: No bootable option or device was found" and the VM hangs at the UEFI menu until the 40-minute job timeout. Fix: when repacking cc.tar to pxe.tar.gz, strip boot.efi and uki so run_qemu.sh selects traditional vmlinuz/initrd/cmdline iPXE boot (is_uki=0). Also patch the .requirements file to set uefi=false so SeaBIOS is used instead of OVMF; SeaBIOS iPXE boots from network via the virtio-net option ROM without needing a UEFI NIC boot entry. The cc.tar already contains vmlinuz, initrd, cmdline, and root.squashfs (added in 9ed7723) in addition to boot.efi/uki, so all files needed for traditional PXE boot are present.
…r.gz The _usi build produces cmdline from /etc/kernel/cmdline which does not include the live-boot parameters added by the _pxe feature (ip=dhcp gl.live=1 gl.ovl=/:tmpfs). Without gl.live=1 the gardenlinux initrd does not switch to live mode and the system fails to boot in QEMU (no /dev/vda root device is available). When converting cc.tar to pxe.tar.gz for QEMU testing, append these live-boot parameters to the cmdline file so the traditional iPXE boot path works the same way it does for _pxe flavors.
_usi flavors (sci_usi, sci_usidev, controlplane_usi) produce a .cc.tar artifact containing UKI components (vmlinuz, initrd, EROFS squashfs). The embedded initrd is built for EROFS/systemd-repart boot on real hardware and does not include the gardenlinux-live dracut module. Attempting to live-boot via QEMU (pxe.tar.gz + ip=dhcp gl.live=1) fails silently because the initrd cannot switch root from squashfs: it expects EROFS, and the gardenlinux-live module (which handles gl.live=1 and gl.url) is absent. The build artifact also includes a .tar (full rootfs) which is suitable for chroot-based integration testing. Switch _usi flavors to chroot testing via ./test .build/$CNAME.tar, skipping the now-unnecessary cc.tar → pxe.tar.gz conversion step. - Add use_chroot: true/false to each matrix entry - Skip QEMU install and KVM setup for chroot flavors - Remove the broken cc.tar → pxe.tar.gz conversion block - Copy chroot.test.log/xml in addition to qemu.test.log/xml
Both test failures are expected given the sci feature set: 1. test_usi_no_udev_rules_image_dissect: upstream _usi/file.exclude removes /usr/lib/udev/rules.d/90-image-dissect.rules, but the sci-local _usi override does not carry this exclusion. File is benign on real hardware. 2. test_only_expected_capabilities_are_set: sci flavors install sssd (scibase) and cloud-hypervisor (vhost) which set file capabilities intentionally. The test's allowlist only covers arping. These are feature-correct configurations, not defects. Deselect both tests for all _usi flavors (sci_usi, sci_usidev, controlplane_usi).
The previous approach used a systemd-boot Type 1 entry with 'linux /EFI/BOOT/gardenlinux.efi', but the 'linux' keyword expects a bzImage Linux kernel, not a UKI (PE32+ EFI binary). Systemd-boot's boot.c call_image_start() rejects it with 'Load error'. Fix: place the UKI directly as EFI/BOOT/BOOTX64.EFI so OVMF loads it natively as an EFI binary, bypassing the systemd-boot kernel loader path. Since the UKI's embedded cmdline does not include ignition parameters (the kvm feature's 50-ignition.cfg is not part of baremetal-sci_usi), patch the .cmdline PE section using objcopy before writing to the ESP. Also removes the now-unnecessary tar download and systemd-boot extraction.
The previous two-step approach stored jq output in a variable: LAYERS=$(oras manifest fetch ... | jq -r '.layers[]') UKI_SHA=$(echo $LAYERS | jq 'select(.mediaType==...) | .digest') jq '.layers[]' outputs one JSON object per line (multiple documents). echo $LAYERS (unquoted) collapses whitespace, producing a single-line stream that jq cannot parse as separate objects, so select() returns nothing and UKI_SHA is empty. The subsequent oras blob fetch then downloads an empty file, causing objcopy to fail with "input file is empty". Fix: use a single pipeline so jq receives a proper stream: UKI_SHA=$(oras manifest fetch ... | jq -r '.layers[] | select(.mediaType=="application/io.gardenlinux.uki") | .digest')
GHCR requires an explicit Accept header for OCI image manifests. Without --media-type, oras manifest fetch defaults to requesting application/vnd.oci.image.index.v1+json, which GHCR returns HTTP 404 for tags that point to image manifests (not index manifests). The 404 results in empty output, so jq returns SHA256 of empty string, and oras blob fetch downloads a 0-byte file. Fix: pass --media-type 'application/vnd.oci.image.manifest.v1+json' to ensure GHCR serves the image manifest with its layers array.
oras manifest fetch v1.2.2 has content negotiation issues with GHCR when requesting application/vnd.oci.image.manifest.v1+json: it returns empty output despite the tag existing. curl with explicit Accept and anonymous GHCR token reliably returns the image manifest with layers. Still use oras blob fetch to download the UKI blob itself (works fine).
…arly on empty UKI PR builds with version=today use the 'today' apt suite which has an older systemd-ukify (261.2) that produces empty UKI files. Switching to version=now uses the current numbered suite (e.g. 2383.0.0) with systemd-ukify 262 which produces correct non-empty UKI binaries. Also add an explicit validation check in the build action to fail early with a clear error message if the downloaded UKI blob is empty, instead of propagating the empty file into objcopy which produces a confusing error.
Passing the literal string "now" as VERSION caused the upload_oci.yml to construct CNAME with "now" (e.g. baremetal-sci_usi-amd64-now-COMMIT) while the build job had already resolved it to the actual version (e.g. 2383.0.0-COMMIT), resulting in a tar.gz filename mismatch. Resolve "now" via gardenlinux/bin/garden-version in the set_version job so all downstream jobs (build, upload, test) use the same concrete version.
Hardlinks in a tar archive cause alphabetically-later entries (uki, squashfs) to be stored as zero-byte hardlink references when the alphabetically-earlier entry (boot.efi, root.squashfs) holds the actual data. tar --extract --to-stdout on the later entry then produces 0 bytes, resulting in an empty UKI/squashfs in the OCI artifact. Use cp instead of ln so both entries are stored independently.
tar stores hardlinks with data only in the first (alphabetically) entry. boot.efi < uki and root.squashfs < squashfs, so boot.efi/root.squashfs hold the actual data while uki/squashfs are zero-byte hardlink refs. Fix convert.uki~cc.tar to extract boot.efi and convert.squashfs~cc.tar to extract root.squashfs. The ln aliases in image.cc.tar are preserved to avoid running out of disk space on the build runner.
… UEFI With firmware='efi', libvirt's <boot dev='hd'/> in <os> doesn't control OVMF's boot order — OVMF was falling through to Boot0001 which pointed to the virtio-net NIC (PXE) instead of the virtio disk's ESP. Set <boot order='1'/> on the disk device, which maps to OVMF's BootXXXX NVRAM variables via QEMU's -boot strict mechanism, ensuring OVMF boots from the disk (and finds EFI/BOOT/BOOTX64.EFI) first.
The VM does a two-boot cycle on firstboot (persist downloads UKI from OCI
then reboots). Increase MAX_ITER from 30 to 60 (5 min -> 10 min) to allow
time for the persist script to download the UKI and complete the second boot.
Also capture /var/log/HV{1,2}.log (QEMU serial console output with OVMF +
kernel boot messages) in debug artifacts for easier failure diagnosis.
The Ubuntu OVMF_VARS_4M.fd template has a pre-populated Boot0001 entry pointing to a raw PCI device path, not an EFI file path. This causes OVMF to attempt booting from a raw block device rather than scanning the FAT32 ESP partition for \EFI\BOOT\BOOTX64.EFI. Fix: create a zeroed (empty) VARS file per VM via dd, and reference it explicitly in the domain XML without firmware='efi' auto-detection. With no valid boot entries in NVRAM, OVMF runs its default boot path discovery which scans all block devices for \EFI\BOOT\BOOTX64.EFI.
…TX64.EFI Root cause: libvirt always injects bootindex:1 on the first disk when pflash UEFI is configured. With virtio-blk, OVMF creates a ShortForm boot entry (just PciRoot/Pci(addr)) with no partition/file path — it tries to execute the raw block device as an EFI app, silently fails, and the VM never boots. With SATA (AHCI), OVMF's AhciDxe driver properly enumerates the GPT partition table and builds a FullPath boot entry including the partition descriptor and \EFI\BOOT\BOOTX64.EFI. The disk is already correctly partitioned by the build action with a FAT32 ESP containing the UKI.
With libvirt's pflash UEFI setup, the SATA disk always gets bootindex=1 injected, which causes OVMF to create a device-level Boot0001 entry (Sata(0x0,0xFFFF,0x0)) and try to execute the raw disk as an EFI app — failing silently. Add -global ide-hd.bootindex=0 to the QEMU commandline to suppress bootindex on the SATA device. OVMF then falls back to scanning block devices for \EFI\BOOT\BOOTX64.EFI, which the persist script places there on firstboot. Also replace the zeroed VARS file with a copy of the stock OVMF_VARS to ensure OVMF's boot variable infrastructure is intact.
libvirt's bootindex=1 causes OVMF to create a device-level Sata() boot entry (Boot0001) that points to the raw SATA controller rather than a file. OVMF tries to execute the disk as an EFI app and hangs forever — no fallback scanning occurs. Fix: use virt-fw-vars to inject a File(\EFI\BOOT\BOOTX64.EFI) entry into the OVMF VARS before starting the VMs. This entry becomes Boot0000 and is first in BootOrder, so OVMF finds and executes the UKI on the ESP before even attempting the bad device-path entry. Also add python3-virt-firmware to apt dependencies.
The UKI's embedded cmdline only has console=tty0 (VGA). Without console=ttyS0, the kernel produces no output on the libvirt serial device, making it impossible to debug boot failures. Also needed for the persist firstboot script's output to be visible.
The kernel crashes silently (zero serial output) after OVMF hands off to the UKI. earlycon=uart8250,io,0x3f8 and earlyprintk=serial,ttyS0 enable serial output before the kernel serial driver initializes, which will reveal the crash/panic message.
The default pc-i440fx machine type is legacy PC chipset. Modern UKIs use EFI protocols (e.g. EFI_LOADED_IMAGE, GOP) that are better supported under q35 (PCIe-based machine). Switch to q35 to match what the Garden Linux UKI expects for its EFI stub.
After 'persist' firstboot downloads and installs the production UKI, the second boot uses the freshly-downloaded UKI whose original cmdline has no console=ttyS0 and no ignition.platform.id=qemu. This means: - no serial console output on the second boot - ignition doesn't run -> SSH authorized_keys are not set up - sshd may not be enabled Fix by injecting an /opt/persist/network_up.sh hook (via ignition) that persist calls (chrooted into the new sysroot) between the UKI download and systemd-repart. The hook: - patches the downloaded UKI's .cmdline with console=ttyS0 and ignition.platform.id=qemu so the second boot mirrors the first - runs 'systemctl enable ssh.service' in the new sysroot Also enable ssh.service via the ignition systemd unit config for the first boot, and increase the HV SSH wait timeout from 60 to 90 iterations (15 min) to give persist + second boot time to complete.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
nightly.yamlworkflow onmaingardenlinux/.github/workflows/test_flavor_qemu.ymlat the same pinned commit already used forbuild.yml(c15e9789)baremetal-sci_usi(amd64) flavor — the primary SCI hypervisor imageBackports needed
Once validated on
main, the same change needs to be backported to:rel-2150-dev— same flavorbaremetal-sci_usi, pinned commit53320d17rel-1877-dev— flavormetal-sci_usi(different target name), pinned commit6f022335Test plan
workflow_dispatchand verifytest_qemujob passesbaremetal-sci_usi-amd64boots successfully in QEMU on GitHub Actions runner