My NAS runs on a Supermicro X10SDV-TLN4F — Xeon D-1541, 128 GB of ECC, a 22 TB ZFS pool, and a BIOS dated 2017. It had been up since March. I did some network configuration work one morning, rebooted, and it didn’t come back.
What it showed instead was the Supermicro splash screen with A2 in the bottom
right corner, and a keyboard that did nothing. DEL wouldn’t get me into Setup,
F11 wouldn’t open the boot menu, F12 wouldn’t PXE. The machine sat there
indefinitely.
A2 is an AMI Aptio progress code meaning “IDE Detect” — the firmware enumerating SATA devices. That framing sent me down the wrong path for the first two hours, because “stuck detecting storage” reads like “a disk is dying.” It isn’t necessarily. A2 is the last code the firmware latched, and the firmware stops updating it once it hands off. A machine displaying A2 forever might be stuck enumerating storage, or it might be several stages further on and no longer writing new codes.
Getting into a machine whose keyboard is dead
The BMC was reachable even though the host wasn’t, which turned out to be the
single most useful fact of the day. ipmitool gave me sensors, the event log,
and — critically — the ability to force the next boot into Setup without
pressing anything:
ipmitool -I lanplus -H <bmc> -U ADMIN -E chassis bootdev bios
ipmitool -I lanplus -H <bmc> -U ADMIN -E chassis power reset
That sets an IPMI boot flag the firmware reads at the start of BDS. It bypasses
the keypress entirely. I used it maybe fifteen times over the course of the day,
and without it I’d have been stuck at “the keyboard doesn’t work” with no way
in. If you run Supermicro hardware and haven’t used chassis bootdev, it’s
worth knowing before you need it.
The sensors were all boring, which is itself evidence:
12V 12.128 CPU Temp 50 C
5VCC 5.026 System Temp 41 C
3.3VCC 3.367 VBAT 3.226
Rails nominal, CMOS battery fine, no power overload, no drive fault, no fan fault. The System Event Log was full — 512 entries, 100% used, overflow set — and every single entry from February 2025 to the night before was a correctable ECC error, mostly on DIMMA1. That matters later, but nothing in it pointed at a boot failure.
Five theories the evidence killed
A failing drive. The obvious reading of “IDE Detect.” I got into Setup via the IPMI flag and every one of the six SATA devices enumerated with the correct model and capacity. Two 6 TB HGSTs, two 22 TB Seagates, two 1 TB Samsung SSDs. Nothing missing, nothing mis-sized.
A full /boot. The morning’s work had rebuilt the initramfs, and /boot
was a 387 MB partition. This was my favourite theory for hours because the
timing was perfect. When I finally got a live USB booted and looked, /boot was
61% used — 213 MB of 373 MB, all three kernels intact with matching initrds.
Completely healthy. I’d been confidently wrong about it in three separate
messages.
A stale NVRAM boot entry. The boot menu had four entries: UEFI OS,
ubuntu, Ubuntu, and a Windows Boot Manager left over from a transplanted
disk. UEFI OS was first, and it pointed at the fallback path
\EFI\BOOT\BOOTX64.EFI, which chained to a 2021 copy of grubx64.efi sitting
next to a 2026 shim. That’s a genuine landmine and I was pleased to find it. It
also wasn’t the problem: promoting ubuntu to first changed nothing.
Boot order generally. The killer here was launching the bootloader by explicit path from the EFI Shell, which bypasses NVRAM completely:
fs0:\EFI\ubuntu\shimx64.efi
Same hang, no output. So it wasn’t the entries, the ordering, or the firmware’s choice of target. Whatever was failing was failing after the firmware handed control over.
A video artefact. By this point I wondered whether GRUB was running fine and
the iKVM just wasn’t capturing the mode it selected. Cheap test: launch it, then
type reboot blind. If the machine reboots, GRUB is alive and I’m staring at a
stale framebuffer. Nothing happened. Genuinely hung.
The EFI Shell is a real debugger
The X10SDV has a built-in EFI Shell as a boot option, and I’d never used it for anything. It turns out to be the best tool available when you’re below the OS.
map -r lists every filesystem and block device with its full device path:
fs0: PciRoot(0x0)/Pci(0x1F,0x2)/Sata(0x4,0xFFFF,0x0)/HD(3,MBR,0x52796A31,0x800,0x32000)
fs1: PciRoot(0x0)/Pci(0x1F,0x2)/Sata(0x5,0xFFFF,0x0)/HD(2,GPT,090296AA-...,0xE1800,0x32000)
That’s readable once you know the grammar. Sata(0x4,...) is SATA port 4.
HD(3,MBR,...) is partition 3 on an MBR-labelled disk. So fs0 is the ESP on
the disk in port 4, and it’s partition three of an MBR disk — which told me
immediately that my “modern UEFI system” had an MBR boot disk with the ESP
mistyped as 0x0b (W95 FAT32) instead of 0xEF. The firmware mounted it
anyway, because EDK2’s partition driver attaches to any FAT it finds, but it was
wrong and had been wrong for years.
From there ls, type to read files, and the ability to execute any EFI binary
by path let me bisect the boot chain properly: shim hangs, GRUB hangs, a
completely different 2021 GRUB build hangs, the EFI Shell itself runs fine,
Windows Boot Manager boots to a spinner. Three different loaders failing
identically while two others work is a strong signal, and I still couldn’t turn
it into a diagnosis.
Not root-caused
I never found it. I want to be plain about that, because the genre convention is to end with the answer, and the ones that don’t end with an answer mostly don’t get written — which quietly teaches people that not finding a root cause means they failed at debugging.
What I did instead was replace the entire boot stack: repartition to GPT, a
correctly typed 2 GiB ESP, mdadm RAID1, LUKS2, LVM, and systemd-boot rather
than GRUB. The machine boots. It’s a better configuration than what it replaced
on every axis I care about. But I can’t tell you the original fault, and there’s
a real chance I’ve buried it rather than fixed it.
The one lead I’m still uneasy about is memory. /etc/default/grub on that
machine contained this:
GRUB_BADRAM="0x1f07817000,0xfffffffffffff000,0x1a664ed000,0xfffffffffffff000"
Two bad 4 KB pages, masked out deliberately by someone — me, years ago —
rather than replaced. Combined with 512 correctable ECC events concentrated on
one DIMM and accelerating over the preceding two months, that’s a machine with
known-bad memory. GRUB_BADRAM only takes effect once grub.cfg executes, so
GRUB’s own load and relocation were never protected by it. A loader landing on a
bad page hangs with no output, deterministically, which fits everything I saw.
It also fits nothing I could prove, and the live USB booted through shim and
GRUB on the same firmware all day without trouble. So I carried the mitigation
forward — systemd-boot has no GRUB_BADRAM equivalent, so the two pages are now
reserved on the kernel command line instead:
memmap=4K$0x1f07817000 memmap=4K$0x1a664ed000
and confirmed after boot that the kernel honoured it:
kernel: user-defined physical RAM map:
kernel: e820: reserve RAM buffer [mem 0x1a664ed000-0x1a67ffffff]
kernel: e820: reserve RAM buffer [mem 0x1f07817000-0x1f07ffffff]
That’s the honest end state. A working machine, a better boot stack, a mitigation preserved across a bootloader change, and an unexplained failure I’ve routed around rather than understood.
Notes
- POST code A2 on AMI Aptio is “IDE Detect”, but a latched code tells you where the firmware last was, not where it’s stuck.
ipmitool chassis bootdev biosis the way into Setup on a machine with no working keyboard.- The built-in EFI Shell reads device paths, lists FAT filesystems, and executes arbitrary EFI binaries. It is the right tool for bisecting a boot chain.
GRUB_BADRAMdoes not protect GRUB itself. The mask applies only aftergrub.cfgruns.