Skip to content

[libcpu][aarch64] fix boot failure when dropping from EL2 in entry_point.S - #11671

Open
Xenithya wants to merge 3 commits into
RT-Thread:masterfrom
Xenithya:fix/aarch64-el2-boot
Open

[libcpu][aarch64] fix boot failure when dropping from EL2 in entry_point.S#11671
Xenithya wants to merge 3 commits into
RT-Thread:masterfrom
Xenithya:fix/aarch64-el2-boot

Conversation

@Xenithya

Copy link
Copy Markdown

拉取/合并请求描述:(PR description)

[

为什么提交这份PR (why to submit this PR)

Fix AArch64 cold boot failure (HPFAR_EL2 Stage-2 fault / panic) when booted directly from U-Boot in EL2 mode.

你的解决方案是什么 (what is your solution)

  1. Reordered startup sequence: Executed init_cpu_el before accessing any _el1 system registers (e.g. tpidr_el1) in _start.
  2. Fixed CurrentEL logic: Re-evaluated CurrentEL after dropping from EL3 to EL2 in .init_cpu_hyp.
  3. Cleaned up EL2 hardware context:
    • Disabled EL2 MMU and Caches (sctlr_el2).
    • Cleared FP/SIMD traps (cptr_el2).
    • Invalidated stale EL2/Stage-2 TLBs (tlbi alle2is, tlbi vmalle1is) to resolve hardware fault traps.

如何测试 / How to test?

  • Hardware Platform: AArch64 board (Cortex-A CPU)
  • Bootloader: U-Boot (running in EL2 mode, without prior ATF cleanup)
  • Test Result: Cold boot successfully enters the msh prompt every time without hanging or throwing HPFAR_EL2 exceptions.

请提供验证的bsp和config (provide the config and bsp)

  • BSP:
  • .config:
  • action:

]

当前拉取/合并请求的状态 Intent for your PR

必须选择一项 Choose one (Mandatory):

  • 本拉取/合并请求是一个草稿版本 This PR is for a code-review and is intended to get feedback
  • 本拉取/合并请求是一个成熟版本 This PR is mature, and ready to be integrated into the repo

代码质量 Code Quality:

我在这个拉取/合并请求中已经考虑了 As part of this pull request, I've considered the following:

  • 已经仔细查看过代码改动的对比 Already check the difference between PR and old code
  • 代码风格正确,包括缩进空格,命名及其他风格 Style guide is adhered to, including spacing, naming and other styles
  • 没有垃圾代码,代码尽量精简,不包含#if 0代码,不包含已经被注释了的代码 All redundant code is removed and cleaned up
  • 所有变更均有原因及合理的,并且不会影响到其他软件组件代码或BSP All modifications are justified and not affect other components or BSP
  • 对难懂代码均提供对应的注释 I've commented appropriately where code is tricky
  • 代码是高质量的 Code in this PR is of high quality
  • 已经使用clang-format 源码格式化工具确保格式符合RT-Thread代码规范 This PR has been formatted with clang-format and complies with RT-Thread code specification
  • 如果是新增bsp, 已经添加ci检查到.github/ALL_BSP_COMPILE.json 详细请参考链接BSP自查

…int.S

When booting RT-Thread directly from U-Boot at EL2 mode (without prior ATF context sanitization), cold boot fails due to invalid EL1 register access and residual EL2 hardware states.

Key fixes:
1. Reorder init_cpu_el in _start: Ensure CPU drops to EL1 before executing any _el1 system register instructions (e.g., msr tpidr_el1, xzr).
2. Re-read CurrentEL in .init_cpu_hyp: Fix logic issue when dropping from EL3 to EL2 where x0 held stale register values.
3. Clean up EL2 hardware context in .init_cpu_hyp:
   - Disable EL2 MMU and Caches (sctlr_el2).
   - Untrap FP/SIMD instructions by clearing cptr_el2.
   - Invalidate stale EL2 and Stage-2 TLBs (tlbi alle2is, tlbi vmalle1is) to prevent HPFAR_EL2 translation faults.
@github-actions

Copy link
Copy Markdown

👋 感谢您对 RT-Thread 的贡献!Thank you for your contribution to RT-Thread!

为确保代码符合 RT-Thread 的编码规范,请在你的仓库中执行以下步骤运行代码格式化工作流(如果格式化CI运行失败)。
To ensure your code complies with RT-Thread's coding style, please run the code formatting workflow by following the steps below (If the formatting of CI fails to run).


🛠 操作步骤 | Steps

  1. 前往 Actions 页面 | Go to the Actions page
    点击进入工作流 → | Click to open workflow →

  2. 点击 Run workflow | Click Run workflow

  • 设置需排除的文件/目录(目录请以"/"结尾)
    Set files/directories to exclude (directories should end with "/")
  • 将目标分支设置为 \ Set the target branch to:fix/aarch64-el2-boot
  • 设置PR number为 \ Set the PR number to:11671
  1. 等待工作流完成 | Wait for the workflow to complete
    格式化后的代码将自动推送至你的分支。
    The formatted code will be automatically pushed to your branch.

完成后,提交将自动更新至 fix/aarch64-el2-boot 分支,关联的 Pull Request 也会同步更新。
Once completed, commits will be pushed to the fix/aarch64-el2-boot branch automatically, and the related Pull Request will be updated.

如有问题欢迎联系我们,再次感谢您的贡献!💐
If you have any questions, feel free to reach out. Thanks again for your contribution!

@github-actions github-actions Bot added Arch: ARM/AArch64 BSP related with arm libcpu labels Jul 31, 2026
@github-actions

github-actions Bot commented Aug 2, 2026

Copy link
Copy Markdown

MemBrowse Memory Report

gd32105r-start

  • CODE: .text -104 B (-0.1%, 69,260 B / 262,144 B, total: 26% used)

hc32f334

  • FLASH: .text -104 B (-0.1%, 99,460 B / 131,072 B, total: 76% used)

infineon-psoc6

  • flash: .text -104 B (-0.1%, 116,104 B / 262,144 B, total: 44% used)

k230

  • SRAM: .eh_frame -56 B, .text -32 B (-0.0%, 1,122,643 B / 268,300,288 B, total: 0% used)

loongson-ls1cdev

  • Code: .text -168 B (-0.0%, 362,600 B)

nuvoton-m487

  • CODE: .text -132 B (-0.0%, 355,680 B / 524,288 B, total: 68% used)

qemu-virt64-aarch64

  • Code: .eh_frame +56 B, .text +384 B (+0.0%, 1,056,924 B)
  • Data: .data +8 B (+0.0%, 517,384 B)

raspberry-pico-rp2040

  • FLASH: .rodata -24 B, .text -112 B (-0.1%, 114,064 B / 2,097,152 B, total: 5% used)

simulator

  • Code: .eh_frame +32 B, .eh_frame_hdr +8 B, .rela.dyn -432 B, .rodata -64 B, .text -83 B (-0.0%, 2,633,564 B)
  • Data: .data -512 B (-0.0%, 1,582,724 B)

stm32f407-rt-spark

  • CODE: .text -104 B (-0.1%, 85,728 B / 1,048,576 B, total: 8% used)

stm32l475-atk-pandora-llvm

@Xenithya

Xenithya commented Aug 7, 2026

Copy link
Copy Markdown
Author

关于 QEMU -smp 4 环境下 CI 测试失败的原因说明:
经过对日志和 entry_point.S 改动的排查,发现了如下情况:

  1. 在原代码中,从 EL3 eret 到 EL2 后,x0 寄存器仍保留着原来的 CurrentEL 旧值(3),导致 cmp x0, #2 条件判断失效,原代码实际上彻底跳过了 EL2 的初始化逻辑
  2. 本次 PR 修正了这一逻辑(重新读取 CurrentEL),使得 EL2 的清理工作(sctlr_el2, cptr_el2, tlbi)得以正常执行。
  3. 但由于 CI 的工作流参数硬编码了 -smp 4 启动 QEMU,在跑非 SMP 配置(如 default.cfg)时,原代码因 EL2 初始化被跳过,次核(Secondary CPUs)未能顺畅引导;而修复后 4 个 CPU 核心均被成功引导至 EL1,导致非 SMP 配置下次核过早冲入 rt_system_scheduler_start() 并触发了 to_thread 断言。
    entry_point.S 架构层面的修改本身是正确且必要的。想请教一下维护者老师,您是否赞同这个根因分析? entry_point.S 的修改逻辑是否正确?感谢!

@Rbb666

Rbb666 commented Aug 17, 2026

Copy link
Copy Markdown
Member

关于 QEMU -smp 4 环境下 CI 测试失败的原因说明: 经过对日志和 entry_point.S 改动的排查,发现了如下情况:

  1. 在原代码中,从 EL3 eret 到 EL2 后,x0 寄存器仍保留着原来的 CurrentEL 旧值(3),导致 cmp x0, #2 条件判断失效,原代码实际上彻底跳过了 EL2 的初始化逻辑
  2. 本次 PR 修正了这一逻辑(重新读取 CurrentEL),使得 EL2 的清理工作(sctlr_el2, cptr_el2, tlbi)得以正常执行。
  3. 但由于 CI 的工作流参数硬编码了 -smp 4 启动 QEMU,在跑非 SMP 配置(如 default.cfg)时,原代码因 EL2 初始化被跳过,次核(Secondary CPUs)未能顺畅引导;而修复后 4 个 CPU 核心均被成功引导至 EL1,导致非 SMP 配置下次核过早冲入 rt_system_scheduler_start() 并触发了 to_thread 断言。
    entry_point.S 架构层面的修改本身是正确且必要的。想请教一下维护者老师,您是否赞同这个根因分析? entry_point.S 的修改逻辑是否正确?感谢!

ci的问题是需要也顺便修复下的,谢谢

@rcitach rcitach left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hello,一些小的意见


/* Now we are in the end of boot cpu process */
ldr x19, =rtthread_startup
ldr x8, =rtthread_startup

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

此处为何要修改为x8,在 AArch64 PCS 中x8属于caller-saved register,在 init_mmu_early -> enable_mmu_early 中会执行 bl 指令,因此无法保证 x8 在函数调用后仍保持原有值

bic x0, x0, #(1 << 1) /* Disable Alignment check */
msr sctlr_el1, x0

mrs x0, cntkctl_el1

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

此处为何需要删除?该配置用于支持 VDSO 功能,使 EL0 可以直接访问虚拟计数器。如果不需要 VDSO 功能,建议通过宏进行条件控制,而不是直接删除

@rcitach

rcitach commented Aug 18, 2026

Copy link
Copy Markdown
Contributor

关于 QEMU -smp 4 环境下 CI 测试失败的原因说明: 经过对日志和 entry_point.S 改动的排查,发现了如下情况:

  1. 在原代码中,从 EL3 eret 到 EL2 后,x0 寄存器仍保留着原来的 CurrentEL 旧值(3),导致 cmp x0, #2 条件判断失效,原代码实际上彻底跳过了 EL2 的初始化逻辑
  2. 本次 PR 修正了这一逻辑(重新读取 CurrentEL),使得 EL2 的清理工作(sctlr_el2, cptr_el2, tlbi)得以正常执行。
  3. 但由于 CI 的工作流参数硬编码了 -smp 4 启动 QEMU,在跑非 SMP 配置(如 default.cfg)时,原代码因 EL2 初始化被跳过,次核(Secondary CPUs)未能顺畅引导;而修复后 4 个 CPU 核心均被成功引导至 EL1,导致非 SMP 配置下次核过早冲入 rt_system_scheduler_start() 并触发了 to_thread 断言。
    entry_point.S 架构层面的修改本身是正确且必要的。想请教一下维护者老师,您是否赞同这个根因分析? entry_point.S 的修改逻辑是否正确?感谢!

hello,以上
1.EL2 的初始化逻辑并不会跳过,在EL3eret 到 EL2,返回地址是 init_cpu_hyp,并不会执行 cmp x0, #2 条件判断
3.同上,可在qemu环境中运行一下,通过 backtrace 分析具体原因

/* Drop to EL1 and clean up hypervisor state */
bl init_cpu_el

/* Save cpu id temp */

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

by the way,此处为何也要删除

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

非常抱歉,这个PR是基于v5.1.0开发的,正在重新整理

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Arch: ARM/AArch64 BSP related with arm libcpu

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants