Skip to content

Latest commit

 

History

History
373 lines (302 loc) · 16.4 KB

File metadata and controls

373 lines (302 loc) · 16.4 KB

Boot and image loading

This document describes how the machine goes from power-on to the first userspace instruction: the hand-rolled first-stage loader (loader/), the layout of the kernel8.img boot image, how that image is assembled at build time (the loader binary concatenated with a TAR-format initial ramdisk), and how the loader parses and maps the kernel ELF before handing control to kernel_main.

The loader is a small, freestanding AArch64 program that shares the kernel's libkernel, cxx, elf, and tar libraries but runs before virtual memory exists. Its whole job is to set up enough of the machine (an MMU, a physical page allocator, a mapped kernel image and stack) that the real kernel can take over in the high half of the address space.

Table of contents

The boot image

The Raspberry Pi's own firmware (the VideoCore GPU) loads a file named kernel8.img from the boot medium to physical address 0x80000 and starts the first ARM core executing there. There is no GRUB, no U-Boot, and no second stage; the loader is what the firmware jumps into directly.

kernel8.img is not an ELF or a filesystem image. It is a raw concatenation of two blobs:

kernel8.img
├─ rtloader.img   raw binary of the loader, linked to run at 0x80000
└─ initrd.tar     USTAR archive: the kernel ELF + all userspace programs

Because the loader is linked at 0x80000 and placed at the front, the firmware's entry point lands exactly on _start. The loader knows where the initrd begins because the linker emits a loader_end_ symbol at the end of the loader's .bss; the TAR archive starts at that address (see loader/linker.ld and the loader_end_ reference in loader/loader.h).

Building the image

The final image is produced by the kernel8.img custom target in CMakeLists.txt. Three things feed into it:

  1. rtloader.img — the loader is linked into rtloader.elf, then llvm-objcopy -O binary flattens it to a raw binary. The .bss section is forced to alloc,load so that objcopy lays it out at its file offset rather than dropping it; the loader's linker.ld carries a prominent warning that no loader section may be aligned, because objcopy positions sections by file offset (not VMA) and any padding would desynchronize the two.

    COMMAND llvm-objcopy rtloader.elf -O binary
            --set-section-flags=.bss=alloc,load rtloader.img
  2. The sysroot — every add_user_program target (see docs/writing-a-program.md) emits its binary into build/sysroot/ at its archive path (e.g. sbin/initd, bin/sh). Just before packing, the kernel8.img target also copies two files in:

    • userspace/services/init.lua → sysroot/etc/init.lua (the Lua init script)
    • build/rtkernel.elf → sysroot/boot/rtkernel.elf (the kernel itself, shipped as a normal file inside the ramdisk)
  3. initrd.tar — scripts/make_initrd.py walks the sysroot and packs everything into a USTAR archive.

The three steps then run in order and the result is concatenated:

COMMAND python3 scripts/make_initrd.py build/initrd.tar build/sysroot/
COMMAND cat rtloader.img initrd.tar > kernel8.img

The kernel ELF therefore travels inside the same ramdisk as userspace: the loader extracts it at boot (see Loading the kernel ELF), and the kernel later re-reads the same archive to find the init program.

The initial ramdisk

The ramdisk is a plain USTAR (POSIX UNIX-standard tar) archive, chosen because the format is trivial to parse without a filesystem: a sequence of 512-byte headers, each followed by the file's contents padded to a 512-byte boundary. make_initrd.py produces it with Python's tarfile in USTAR_FORMAT, storing each file under an absolute-normalized archive path:

with tarfile.open(out, mode="w", format=tarfile.USTAR_FORMAT) as tar:
    for root, dirs, files in os.walk(sysroot):
        for file in files:
            arc = os.path.normpath("/" + os.path.relpath(path, sysroot))
            tar.add(path, arc)

On the device side the archive is walked by tar::Parser (lib/tar/), which is a forward iterator over entries — the loader and the kernel both iterate for (auto& file : tar::Parser{base}) and compare file.header->name. A representative initrd looks like:

boot/rtkernel.elf      the kernel ELF image
etc/init.lua           Lua startup script
sbin/initd             userspace init (the first task)
sbin/clockd  ...       system services
bin/sh  bin/dmesg ...  tools

First instructions: start.S

Execution begins at _start in loader/start.S, placed in the .text.boot section so the linker keeps it at the very front of the image. It runs on all four cores simultaneously, so its first act is to park every core except core 0:

mrs x1, mpidr_el1
and x1, x1, #3        // core index
cbnz x1, exit         // non-zero core -> wfe loop forever

Core 0 then normalizes the exception level. The GPU firmware may start execution in EL2 (hypervisor) or EL1; the loader requires EL1. On finding itself in EL2 it configures the EL2 traps and performs a fake exception return down to EL1:

  • cptr_el2 / cpacr_el1 — clear the FP/SIMD/SVE traps so those registers are usable.
  • hcr_el2 — set RW (EL1 is AArch64) and clear E2H, putting the CPU in the Non-secure EL1&0 translation regime.
  • spsr_el2 / elr_el2 — stage a return to el1_entry with all interrupts masked, then eret.

At el1_entry it programs sctlr_el1 (reserved bits + WFE/WFI), sets the stack pointer to stack_top_ (a 4 KiB stack carved out of .bootloader.stack in the BSS), and branches to loader_main.

The firmware passes the physical address of the device tree blob in x0. start.S never clobbers x0, so it arrives intact as the first argument to loader_main(void* dtb); the same discipline lets jump_to_kernel forward the loader-data pointer through to kernel_main.

start.S also exposes the trampoline used at the very end of the loader:

jump_to_kernel:      // (arg0 in x0, entry in x1, stack in x2)
   mov sp, x2
   br  x1

The loader: loader_main

loader/main.cc is the core of the loader. It runs identity-mapped with the MMU still off and performs, in order: early UART setup, boot-arg parsing, an initrd scan, physical-memory bootstrap, paging setup, kernel ELF loading, kernel-stack allocation, MMU enable, and finally the jump.

Boot arguments and the device tree

The loader first brings up the UARTs so it can log, then reads boot arguments from two sources, later overriding earlier:

  1. DEFAULT_BOOTARGS — baked in at build time from CMake (default init=sbin/initd logmask=7; see the DEFAULT_BOOTARGS option in the README).
  2. The bootargs property from the device tree blob whose address arrived in x0. loader/dtb.cc is a minimal flattened-device-tree walker: it validates the 0xd00dfeed magic, then scans the struct block for the bootargs string property.

loader/bootargs.cc tokenizes the argument string into key=value pairs and fills in gLoaderData. Two keys are understood:

  • init=<path> → LoaderData::init_image, the archive path of the first task.
  • logmask=<n> → LoaderData::kernel_log_mask, the kernel's log verbosity.

Unknown keys are skipped with a warning; a bare key (no =) is treated as a boolean flag set to "1".

Scanning the initrd

The loader iterates the TAR archive at &loader_end_ to do two things: locate the kernel ELF and find where the whole loader+initrd complex ends in physical memory.

for (auto& file : tar::Parser{&loader_end_}) {
   if (file.size == 0) continue;
   if (!cxx::strcmp("boot/rtkernel.elf", file.header->name))
      kernel_image = file.file;   // remember the kernel's location
   initrd_end = file.end;          // track the running end of the archive
}
initrd_end += 512 * 2;             // account for the two zero terminator blocks

initrd_end becomes the low-water mark for everything the loader allocates after this point — the physical page allocator is told this is where free memory begins.

Bootstrapping physical memory

loader/mem.cc sets up a page-frame database (PFN DB): one core::mem::Page struct per 4 KiB physical frame, from address 0 up to the end of usable RAM (0x3b400000, taken from the Pi 4 Linux memory map). The DB lives immediately after the initrd, and the physical layout ends up as:

0x00000000  ┌───────────────────┐
            │  free             │
0x00080000  ├───────────────────┤  &loader_begin_
            │  loader binary    │
            ├───────────────────┤  &loader_end_
            │  initrd (TAR)     │
            ├───────────────────┤  initrd_end  == PFN DB start
            │  PFN database     │  (~4 MiB, one Page per frame)
            ├───────────────────┤
            │  free             │  <- alloc_page_phys() hands these out
0x3b400000  └───────────────────┘  phys_mem_end

Frames covering the loader, initrd, and the PFN DB itself are marked allocated (flags = 1); everything else is pushed onto a free-list. alloc_page_phys() pops a frame off that list, zeroes it, and records it on a PageList so the kernel can later account for (and, for loader-only pages, reclaim) it. The three reserved ranges are recorded in gLoaderData (loader_range, initrd_range, pfndb_range) for handoff.

Bootstrapping paging

bootstrap_paging() configures the AArch64 MMU and builds the initial 4-level page tables:

  • TCR_EL1 — 4 KiB granule for both TTBR0/TTBR1, 48-bit virtual addresses (T0SZ = T1SZ = 16), Normal write-back cacheable walks.
  • MAIR_EL1 — attribute index 0 = Device nGnRnE, index 1 = Normal write-back read/write-allocate.
  • Two L0 tables are allocated — one for the low half (TTBR0, user) and one for the high half (TTBR1, kernel) — each with a recursive self-map installed in entry 511 so page tables can be edited through their own virtual window.

It then installs the mappings the kernel needs to survive the transition. Note everything the kernel touches is also mapped at KernelOffset (0xffff'0000'0000'0000), the high-half kernel window:

Mapping Where Purpose
loader → loader identity (low) keep executing after the MMU turns on
loader → loader + KernelOffset high kernel can reach loader data
initrd → initrd + KernelOffset high, read-only kernel re-reads the ramdisk
PFN DB → PFN DB + KernelOffset high kernel inherits the frame database
device MMIO → both identity + high UART/timer/GIC (see DevicePages)

The map() routine walks (and allocates as needed) the L0→L1→L2→L3 tables by slicing the virtual address into 9-bit indices, choosing the low or high L0 table from the top 16 bits of the address, and panicking on a non-canonical address. TTBR0_EL1 and TTBR1_EL1 are written at the end, but the MMU is not enabled yet.

Loading the kernel ELF

With paging tables ready, load_kernel_elf() parses the kernel image found during the initrd scan using elf::Parser (lib/elf/) and maps each loadable segment into the high half:

for (auto& ph : krn.get_pheaders()) {
   if (ph.p_type != PT_LOAD) continue;
   // Permissions from the ELF program-header flags:
   uint8_t ap = 0b00, xn = 0;
   if (!(ph.p_flags & PF_X)) xn = 1;   // no execute
   if (!(ph.p_flags & PF_W)) ap = 0b10; // read-only
   Range v(ph.p_vaddr, align_up(ph.p_vaddr + ph.p_memsz, 4_KiB));
   Range p(file_base + ph.p_offset, file_base + ph.p_offset + ph.p_filesz);
   // Allocate fresh frames and copy the segment's file bytes into them:
   map_create_range(v, p, PageDescriptor{...}, KernelPages);
}
return krn.img->e_entry;   // the kernel's virtual entry point

map_create_range() allocates a fresh physical frame for each virtual page, copies the segment's file bytes in (p_filesz worth; the tail beyond filesz up to memsz stays zeroed, which is how .bss gets zero-initialized), and maps it with the derived R/W/X permissions. Every segment must be page-aligned in its virtual address — the loader asserts this rather than handling misaligned segments. All kernel virtual addresses come straight from the ELF, which is linked to run at 0xffff'8000'0000'0000 (KernelStart).

Kernel stack and enabling the MMU

The loader then allocates the kernel's initial stack just below KernelStart, with a guard page beneath it (mapped with access_flag = 0 so any overflow faults immediately) and a spare zero page whose physical address is stashed in LoaderData::zero_page_addr — the kernel points TTBR0 at it to cheaply unmap the entire low half after boot.

Because no further allocations happen after this, every loader-only page is freed (page->flags = 0), returning that memory to the kernel. Finally the MMU comes on:

uint64_t sctlr = read_sctlr_el1();
sctlr |= (1 << 12);  // I-cache (unless DISABLE_I_CACHE)
sctlr |= (1 << 2);   // D-cache (unless DISABLE_D_CACHE)
sctlr |= (1 << 0);   // MMU enable
write_sctlr_el1(sctlr);

Because the loader is identity-mapped, execution continues seamlessly with paging live. The loader stamps LoaderData::magic and calls the assembly trampoline:

gLoaderData.magic = LoaderData::MAGIC_VALUE;
jump_to_kernel(&gLoaderData, entry, stack_va);

The handoff to the kernel

The contract between the two stages is the LoaderData struct (kernel/include/core/loader.h):

struct LoaderData final {
   static constexpr uint64_t MAGIC_VALUE = 0xDEADBEEF'CAFEBABEULL;
   Range    loader_range;    // physical range of the loader binary
   Range    initrd_range;    // physical range of the initrd (TAR)
   Range    pfndb_range;     // physical range of the PFN database
   uint64_t magic;           // integrity check, == MAGIC_VALUE
   uint64_t zero_page_addr;  // a spare zeroed frame (used to drop TTBR0)
   uint8_t  kernel_log_mask; // from `logmask=` boot arg
   char     init_image[64];  // from `init=` boot arg, e.g. "sbin/initd"
};

jump_to_kernel sets the stack pointer and branches to the ELF entry point, which lands in kernel_main(LoaderData&) (kernel/core/main.cc). The kernel then:

  1. Verifies magic and copy-constructs its own gLoaderData from the loader's (the loader's copy is about to be reclaimed).
  2. Reclaims the loader. core::mem::init() rebuilds the frame database from the handed-off ranges and reclaims loader-only frames; then TTBR0 is pointed at zero_page_addr and the TLB is flushed, unmapping the entire identity-mapped low half. From here the kernel runs purely in the high half.
  3. Initializes itself — runs C++ static constructors (.init_array), then the exception vectors, slab allocator, thread pool, and event subsystem.
  4. Finds the init binary. It walks the same initrd (now at initrd_range.start + KernelOffset) looking for the file named by init_image, creates it as the first task at priority 1, and erets down to EL0 to begin userspace.

From that point the boot story continues in docs/kernel.md (task creation, scheduling, and the ELF loader that backs every subsequent task) and in docs/runtime-abi.md (what happens once EL0 code starts running).