Re: [PATCH v2 10/13] mm/sparse: remove SECTION_MARKED_PRESENT
by David Carlier
On Mon, Sep 21, 2026 at 09:59:02PM +0200, David Hildenbrand (Arm) wrote:
> enum {
> - SECTION_MARKED_PRESENT_BIT,
> SECTION_HAS_MEM_MAP_BIT,
> SECTION_IS_ONLINE_BIT,
> SECTION_IS_EARLY_BIT,
Hi David,
It seems dropping SECTION_MARKED_PRESENT_BIT shifts the other section flags down
one bit. crash hardcodes SECTION_HAS_MEM_MAP as (1UL << 1) in
pfn_to_map(), so it now tests SECTION_IS_ONLINE: offline and
ZONE_DEVICE sections no longer resolve to a struct page. makedumpfile
only checks bit 0 and is fine.
Keep bit 0 reserved, or export the bits via vmcoreinfo?
Cheers.
1 week
[PATCH] arm64: move irq/overflow/stackframe init to POST_VM phase
by chenhaixiang (A)
When percpu data is in vmalloc area (pcpu_page_first_chunk fallback on
large NUMA systems), arm64_kvtop() cannot translate vmalloc addresses
during POST_GDB because vt->vmalloc_start is not yet initialized by
vm_init(). Moving the stack init functions to POST_VM ensures
vt->vmalloc_start and vt->kernel_pgd are ready before readmem calls.
Signed-off-by: Chen Haixiang <chenhaixiang3(a)h-partners.com>
---
arm64.c | 2 ++
1 file changed, 2 insertions(+)
diff --git a/arm64.c b/arm64.c
index d177092..654b972 100644
--- a/arm64.c
+++ b/arm64.c
@@ -787,7 +787,9 @@ arm64_init(int when)
*/
if(!machdep->machspec->CONFIG_ARM64_KERNELPACMASK)
arm64_recalc_KERNELPACMASK();
+ break;
+ case POST_VM:
arm64_irq_stack_init();
arm64_overflow_stack_init();
arm64_stackframe_init();
--
2.41.0
2 weeks, 2 days
Re: Subject: [BUG/RFC] arm64: stack init functions fail in POST_GDB when percpu data is in vmalloc area
by chenhaixiang (A)
Hi Dave,
Sorry for the late reply.
On 9/15/26 3:32 PM, Dave Young wrote:
> in POST_GDB when percpu data is in vmalloc area
>
> Hi Haixiang,
>
> Thanks for your follow up and details.
>
> On 9/14/26 5:10 PM, chenhaixiang (A) wrote:
> > Hi Dave,
> > Thanks for the feedback.
> >
> > On Fri, Sep 11, 2026 at 18:53 +0000, Dave Young wrote:
> >> HI Haixiang,
> >>
> >> Thanks for reporting the issue.
> >>
> >> On Thu, Sep 10, 2026 at 09:01:42AM +0000, chenhaixiang (A) wrote:
> >>> Subject: [BUG/RFC] arm64: stack init functions fail in POST_GDB when
> >> percpu data is in vmalloc area.
> >>> Hi,
> >>> While running crash on a large arm64 NUMA system, crash fails to
> >>> initialize the arm64 stack subsystems during arm64_init(POST_GDB)
> >>> and aborts startup with SEEK_ERROR on every percpu IRQ stack pointer.
> >>> Environment
> >>> -----------
> >>> - architecture : aarch64 (64K pages, CONFIG_ARM64_64K_PAGES=y)
> >>> - kernel : 6.x, large NUMA, many CPUs
> >>> - crash : current master (arm64.c)
> >>> Background on percpu allocation on this machine
> >>> -----------------------------------------------
> >>> On this arm64 machine the kernel's first percpu chunk does NOT live
> >>> in the linear map, it lives in the vmalloc area. The arm64 percpu
> >>> setup path in drivers/base/arch_numa.c:setup_per_cpu_areas() first
> >>> tries
> >>> pcpu_embed_first_chunk() and, when that fails, falls back to
> >>> pcpu_page_first_chunk()
> >>> Observed symptom / reproduction log
> >>> -----------------------------------
> >>> With crash started on the vmcore, the percpu IRQ stack pointers are
> >>> all in
> >> the
> >>> ffff800100xxxxxx range, which is the vmalloc area on this kernel.
> >>> Because vt->vmalloc_start is still 0 at POST_GDB, arm64_kvtop()
> >>> takes the early "VTOP()" shortcut and returns a bogus physical
> >>> address far beyond max_mapnr, so read_diskdump() reports SEEK_ERROR:
> >>> ...
> >>> <readmem: ffff800100000050, KVADDR, "IRQ stack pointer", 8, (ROE),
> >> 1bfe2ef0>
> >>> <read_diskdump: addr: ffff800100000050 paddr: 800100000050 cnt: 8>
> >>> read_diskdump: SEEK_ERROR: paddr/pfn: 800100000050/800100000
> >> max_mapnr: 60c000000
> >>> crash: seek error: kernel virtual address: ffff800100000050 type:
> >>> "IRQ
> >> stack pointer"
> >>> <readmem: ffff800100021050, KVADDR, "IRQ stack pointer", 8, (ROE),
> >> 1bfe2ef8>
> >>> <read_diskdump: addr: ffff800100021050 paddr: 800100021050 cnt: 8>
> >>> read_diskdump: SEEK_ERROR: paddr/pfn: 800100021050/800100021
> >> max_mapnr: 60c000000
> >>> crash: seek error: kernel virtual address: ffff800100021050 type:
> >>> "IRQ
> >> stack pointer"
> >>> <readmem: ffff800100042050, KVADDR, "IRQ stack pointer", 8, (ROE),
> >> 1bfe2f00>
> >>> <read_diskdump: addr: ffff800100042050 paddr: 800100042050 cnt: 8>
> >>> read_diskdump: SEEK_ERROR: paddr/pfn: 800100042050/800100042
> >> max_mapnr: 60c000000
> >>> crash: seek error: kernel virtual address: ffff800100042050 type:
> >>> "IRQ
> >> stack pointer"
> >>> <readmem: ffff800100063050, KVADDR, "IRQ stack pointer", 8, (ROE),
> >> 1bfe2ef08>
> >>> <read_diskdump: addr: ffff800100063050 paddr: 800100063050 cnt: 8>
> >>> read_diskdump: SEEK_ERROR: paddr/pfn: 800100063050/800100063
> >> max_mapnr: 60c000000
> >>> crash: seek error: kernel virtual address: ffff800100063050 type:
> >>> "IRQ
> >> stack pointer"
> >>> ... (one SEEK_ERROR per CPU)
> >>> Note the wrong physical addresses (800100xxxxxx) come from a direct
> >>> linear-map style translation (addr - PAGE_OFFSET) which is only
> >>> valid for the linear map, not for vmalloc-backed percpu memory.
> >>> Root cause
> >>> ----------
> >>> In arm64_init() the three stack initializers are invoked in the
> >>> POST_GDB phase:
> >>> case POST_GDB:
> >>> ...
> >>> if (!machdep->machspec->CONFIG_ARM64_KERNELPACMASK)
> >>> arm64_recalc_KERNELPACMASK();
> >>> arm64_irq_stack_init(); /* <-- readmem on percpu var
> */
> >>> arm64_overflow_stack_init(); /* <-- readmem on percpu var
> */
> >>> arm64_stackframe_init();
> >>> break;
> >>> These helpers end up calling readmem() on percpu addresses that, on
> >>> the machine described above, live in the vmalloc range. The
> >>> translation goes through arm64_kvtop() (arm64.c:1963), which for any
> >>> vmalloc address requires:
> >>> 1. vt->vmalloc_start != 0 -> to reach the IS_VMALLOC_ADDR()
> branch
> >>> 2. vt->kernel_pgd[0] -> to walk the kernel page tables
> >>> Both are populated by vm_init(), which runs in the POST_VM phase --
> >>> i.e. *after* POST_GDB. As a result, in arm64_kvtop() we hit:
> >>> if (!vt->vmalloc_start) {
> >>> *paddr = VTOP(kvaddr); /* wrong for vmalloc addresses */
> >>> return TRUE;
> >>> }
> >>> and the returned physical address is bogus, so the subsequent
> >>> readmem() fails with SEEK_ERROR (as shown in the log above: paddr
> >>> 800100xxxxxx is the linear-map translation of a vmalloc address, and
> >>> it falls well beyond max_mapnr 0x60c000000).
> >>> Proposed fix
> >>> ------------
> >>> Move the three stack initializers from POST_GDB to POST_VM, so that
> >>> vt->vmalloc_start and vt->kernel_pgd are guaranteed to be ready
> >>> before any readmem() that may touch vmalloc-backed percpu data.
> >>> diff --git a/arm64.c b/arm64.c
> >>> index d177092..654b972 100644
> >>> --- a/arm64.c
> >>> +++ b/arm64.c
> >>> @@ -787,7 +787,9 @@ arm64_init(int when)
> >>> */
> >>> if(!machdep->machspec-
> >>> CONFIG_ARM64_KERNELPACMASK)
> >>> arm64_recalc_KERNELPACMASK();
> >>> + break;
> >>>
> >>> + case POST_VM:
> >>> arm64_irq_stack_init();
> >>> arm64_overflow_stack_init();
> >>> arm64_stackframe_init();
> >>> Testing
> >>> -------
> >>> Verified on:
> >>> - the failing large-NUMA arm64 system whose percpu first chunk was
> >>> created by pcpu_page_first_chunk() (i.e. percpu lives in vmalloc):
> >>> with the patch, the percpu IRQ stack pointers are translated
> >>> through the kernel page tables and crash starts cleanly. "bt" on
> >>> the panic task produces correct irq/overflow/standby stacks;
> >>> - a regular arm64 QEMU VM whose percpu first chunk was created by
> >>> pcpu_embed_first_chunk() (i.e. percpu lives in the linear map):
> >>> no regression, stack initialization still happens before any
> >>> "bt" command.
> >>> I'd like to get feedback on whether POST_VM is the preferred phase,
> >>> or whether the maintainers would rather guard the readmem() calls
> >>> inside the three helpers with a fallback (e.g. defer percpu
> >>> translation until later). I can send a formal patch with
> >>> Signed-off-by once the approach is agreed.
> >>
> >> I'm not confident if move it to POST_VM is safe enough so I'd prefer
> >> the second approach, still keep them in POST_GDB, and defer the
> >> callbasks later in a special case.
> >>
> >> Also please add more info in patch log, eg. kernel version, error
> >> msg, etc.
> >
> > Below are the kernel version, reproduction steps, and error log:
> >
> > Kernel: arm64 linux-6.6.0 with CONFIG_GENERIC_ARCH_NUMA=y,
> > CONFIG_NEED_PER_CPU_PAGE_FIRST_CHUNK=y.
> > The same setup_per_cpu_areas() fallback path still exists in linux
> > mainline drivers/base/arch_numa.c, so this should reproduce on current
> > mainline as well.
> > On a QEMU arm64 VM, add "percpu_alloc=page" to the kernel cmdline to
> > force pcpu_page_first_chunk(), so percpu lives in vmalloc.
> > Then:
> > # makedumpfile -d 31 /proc/kcore vmcore
> > # ./crash vmlinux vmcore
> > crash log:
> > crash: seek error: kernel virtual address: ffff800100000050 type: "IRQ
> stack pointer"
> > crash: seek error: kernel virtual address: ffff800100021050 type: "IRQ
> stack pointer"
> > crash: seek error: kernel virtual address: ffff800100042050 type: "IRQ
> stack pointer"
> > crash: seek error: kernel virtual address: ffff800100063050 type: "IRQ
> stack pointer"
> > ...
> > KERNEL: vmlinux [TAINTED]
> > DUMPFILE: vmcore [PARTIAL DUMP]
> > CPUS: 4
> > RELEASE: 6.6.0.aarch64
> > MACHINE: aarch64 (unknown Mhz)
> > MEMORY: 4 GB
> > ...
> >
> > I tried the deferred approach you suggested, but it requires a new
> > flag, a new field, and a new helper, and only fixes the irq_stack_ptr
> > branch. So I looked at how x86_64 handles this.
> > x86_64_init() (x86_64.c:655-661) does not defer anything: it just
> > pre-fills vt->vmalloc_start and vt->kernel_pgd[] before the percpu
> > readmem in POST_GDB:
> > machdep->vmalloc_start = x86_64_vmalloc_start;
> > vt->vmalloc_start = machdep->vmalloc_start();
> > machdep->init_kernel_pgd();
> > ...
> > x86_64_per_cpu_init(); /* readmem percpu, translation works */
> >
> > arm64 registers the same callbacks but never invokes them eagerly,
> > leaving vt->vmalloc_start == 0 until vm_init() runs in POST_VM.
> > So I'd suggest we follow the x86_64 approach and pre-fill on arm64 as well.
> > functions in POST_GDB:
> > case POST_GDB:
> > ...
> > if (!machdep->machspec->CONFIG_ARM64_KERNELPACMASK)
> > arm64_recalc_KERNELPACMASK();
> > vt->vmalloc_start = machdep->vmalloc_start(); /* NEW */
> > machdep->init_kernel_pgd(); /* NEW */
> > arm64_irq_stack_init();
> > arm64_overflow_stack_init();
> > arm64_stackframe_init();
> > break;
> > One caveat: arm64_init_kernel_pgd() uses OFFSET(mm_struct_pgd), which
> > is only initialised by vm_init() in POST_VM. Calling it at POST_GDB
> > triggers "crash: invalid structure member offset:
> > mm_struct_pgd" because OFFSET() FATALs on INVALID_OFFSET. Adding an
> > INVALID_MEMBER() guard lets it fall through to the swapper_pg_dir
> > fallback; the later POST_VM call goes through the init_mm.pgd path as
> > usual:
> > if (!kernel_symbol_exists("init_mm") ||
> > INVALID_MEMBER(mm_struct_pgd) || /* NEW */
> > !readmem(symbol_value("init_mm") +
> OFFSET(mm_struct_pgd), ...)) {
> > if (kernel_symbol_exists("swapper_pg_dir"))
> > value = symbol_value("swapper_pg_dir");
> > ...
> > }
> >
> > Patch
> > -----
> > diff --git a/arm64.c b/arm64.c
> > --- a/arm64.c
> > +++ b/arm64.c
> > @@ -788,6 +788,9 @@ arm64_init(int when)
> > if(!machdep->machspec->CONFIG_ARM64_KERNELPACMASK)
> > arm64_recalc_KERNELPACMASK();
> >
> > + vt->vmalloc_start = machdep->vmalloc_start();
> > + machdep->init_kernel_pgd();
> > +
> > arm64_irq_stack_init();
> > arm64_overflow_stack_init();
> > arm64_stackframe_init();
> > @@ -1894,7 +1897,8 @@ arm64_init_kernel_pgd(void)
> > ulong value;
> >
> > if (!kernel_symbol_exists("init_mm") ||
> > + INVALID_MEMBER(mm_struct_pgd) ||
> > !readmem(symbol_value("init_mm") + OFFSET(mm_struct_pgd),
> KVADDR,
> > &value, sizeof(void *), "init_mm.pgd", RETURN_ON_ERROR)) {
> > if (kernel_symbol_exists("swapper_pg_dir"))
>
>
> The above proposal copies some vm_init code and looks hacky, looking again
> at the code, probably moving the irq stack init to POST_VM looks better.
> I do not have arm64 hardware to test. To ensure no regressions, could you do
> more tests with live debugging (/proc/kcore) and the crash --minimal?
>
I've done the tests you asked for with the patch applied (stack init
moved to POST_VM), on the QEMU arm64 VM with "percpu_alloc=page"
(percpu in vmalloc):
1. live debugging: crash vmlinux /proc/kcore
2. crash --minimal on the vmcore
Both before and after the patch, these two modes show no difference
in behavior and no errors. This matches the code as well -- the move
does not seem to affect either mode.
I'm not sure what else is worth testing -- please let me know if
you need more.
Thanks,
2 weeks, 3 days
Re: Subject: [BUG/RFC] arm64: stack init functions fail in POST_GDB when percpu data is in vmalloc area
by chenhaixiang (A)
Hi Dave,
Thanks for the feedback.
On Fri, Sep 11, 2026 at 18:53 +0000, Dave Young wrote:
> HI Haixiang,
>
> Thanks for reporting the issue.
>
> On Thu, Sep 10, 2026 at 09:01:42AM +0000, chenhaixiang (A) wrote:
> > Subject: [BUG/RFC] arm64: stack init functions fail in POST_GDB when
> percpu data is in vmalloc area.
> > Hi,
> > While running crash on a large arm64 NUMA system, crash fails to
> > initialize the arm64 stack subsystems during arm64_init(POST_GDB) and
> > aborts startup with SEEK_ERROR on every percpu IRQ stack pointer.
> > Environment
> > -----------
> > - architecture : aarch64 (64K pages, CONFIG_ARM64_64K_PAGES=y)
> > - kernel : 6.x, large NUMA, many CPUs
> > - crash : current master (arm64.c)
> > Background on percpu allocation on this machine
> > -----------------------------------------------
> > On this arm64 machine the kernel's first percpu chunk does NOT live in
> > the linear map, it lives in the vmalloc area. The arm64 percpu setup
> > path in drivers/base/arch_numa.c:setup_per_cpu_areas() first tries
> > pcpu_embed_first_chunk() and, when that fails, falls back to
> > pcpu_page_first_chunk()
> > Observed symptom / reproduction log
> > -----------------------------------
> > With crash started on the vmcore, the percpu IRQ stack pointers are all in
> the
> > ffff800100xxxxxx range, which is the vmalloc area on this kernel.
> > Because vt->vmalloc_start is still 0 at POST_GDB, arm64_kvtop() takes
> > the early "VTOP()" shortcut and returns a bogus physical address far
> > beyond max_mapnr, so read_diskdump() reports SEEK_ERROR:
> > ...
> > <readmem: ffff800100000050, KVADDR, "IRQ stack pointer", 8, (ROE),
> 1bfe2ef0>
> > <read_diskdump: addr: ffff800100000050 paddr: 800100000050 cnt: 8>
> > read_diskdump: SEEK_ERROR: paddr/pfn: 800100000050/800100000
> max_mapnr: 60c000000
> > crash: seek error: kernel virtual address: ffff800100000050 type: "IRQ
> stack pointer"
> > <readmem: ffff800100021050, KVADDR, "IRQ stack pointer", 8, (ROE),
> 1bfe2ef8>
> > <read_diskdump: addr: ffff800100021050 paddr: 800100021050 cnt: 8>
> > read_diskdump: SEEK_ERROR: paddr/pfn: 800100021050/800100021
> max_mapnr: 60c000000
> > crash: seek error: kernel virtual address: ffff800100021050 type: "IRQ
> stack pointer"
> > <readmem: ffff800100042050, KVADDR, "IRQ stack pointer", 8, (ROE),
> 1bfe2f00>
> > <read_diskdump: addr: ffff800100042050 paddr: 800100042050 cnt: 8>
> > read_diskdump: SEEK_ERROR: paddr/pfn: 800100042050/800100042
> max_mapnr: 60c000000
> > crash: seek error: kernel virtual address: ffff800100042050 type: "IRQ
> stack pointer"
> > <readmem: ffff800100063050, KVADDR, "IRQ stack pointer", 8, (ROE),
> 1bfe2ef08>
> > <read_diskdump: addr: ffff800100063050 paddr: 800100063050 cnt: 8>
> > read_diskdump: SEEK_ERROR: paddr/pfn: 800100063050/800100063
> max_mapnr: 60c000000
> > crash: seek error: kernel virtual address: ffff800100063050 type: "IRQ
> stack pointer"
> > ... (one SEEK_ERROR per CPU)
> > Note the wrong physical addresses (800100xxxxxx) come from a direct
> > linear-map style translation (addr - PAGE_OFFSET) which is only valid
> > for the linear map, not for vmalloc-backed percpu memory.
> > Root cause
> > ----------
> > In arm64_init() the three stack initializers are invoked in the
> > POST_GDB phase:
> > case POST_GDB:
> > ...
> > if (!machdep->machspec->CONFIG_ARM64_KERNELPACMASK)
> > arm64_recalc_KERNELPACMASK();
> > arm64_irq_stack_init(); /* <-- readmem on percpu var */
> > arm64_overflow_stack_init(); /* <-- readmem on percpu var */
> > arm64_stackframe_init();
> > break;
> > These helpers end up calling readmem() on percpu addresses that, on
> > the machine described above, live in the vmalloc range. The
> > translation goes through arm64_kvtop() (arm64.c:1963), which for any
> > vmalloc address requires:
> > 1. vt->vmalloc_start != 0 -> to reach the IS_VMALLOC_ADDR() branch
> > 2. vt->kernel_pgd[0] -> to walk the kernel page tables
> > Both are populated by vm_init(), which runs in the POST_VM phase --
> > i.e. *after* POST_GDB. As a result, in arm64_kvtop() we hit:
> > if (!vt->vmalloc_start) {
> > *paddr = VTOP(kvaddr); /* wrong for vmalloc addresses */
> > return TRUE;
> > }
> > and the returned physical address is bogus, so the subsequent
> > readmem() fails with SEEK_ERROR (as shown in the log above: paddr
> > 800100xxxxxx is the linear-map translation of a vmalloc address, and
> > it falls well beyond max_mapnr 0x60c000000).
> > Proposed fix
> > ------------
> > Move the three stack initializers from POST_GDB to POST_VM, so that
> > vt->vmalloc_start and vt->kernel_pgd are guaranteed to be ready
> > before any readmem() that may touch vmalloc-backed percpu data.
> > diff --git a/arm64.c b/arm64.c
> > index d177092..654b972 100644
> > --- a/arm64.c
> > +++ b/arm64.c
> > @@ -787,7 +787,9 @@ arm64_init(int when)
> > */
> > if(!machdep->machspec-
> >CONFIG_ARM64_KERNELPACMASK)
> > arm64_recalc_KERNELPACMASK();
> > + break;
> >
> > + case POST_VM:
> > arm64_irq_stack_init();
> > arm64_overflow_stack_init();
> > arm64_stackframe_init();
> > Testing
> > -------
> > Verified on:
> > - the failing large-NUMA arm64 system whose percpu first chunk was
> > created by pcpu_page_first_chunk() (i.e. percpu lives in vmalloc):
> > with the patch, the percpu IRQ stack pointers are translated
> > through the kernel page tables and crash starts cleanly. "bt" on
> > the panic task produces correct irq/overflow/standby stacks;
> > - a regular arm64 QEMU VM whose percpu first chunk was created by
> > pcpu_embed_first_chunk() (i.e. percpu lives in the linear map):
> > no regression, stack initialization still happens before any
> > "bt" command.
> > I'd like to get feedback on whether POST_VM is the preferred phase,
> > or whether the maintainers would rather guard the readmem() calls
> > inside the three helpers with a fallback (e.g. defer percpu
> > translation until later). I can send a formal patch with
> > Signed-off-by once the approach is agreed.
>
> I'm not confident if move it to POST_VM is safe enough so I'd prefer
> the second approach, still keep them in POST_GDB, and defer the
> callbasks later in a special case.
>
> Also please add more info in patch log, eg. kernel version, error msg,
> etc.
Below are the kernel version, reproduction steps, and error log:
Kernel: arm64 linux-6.6.0 with CONFIG_GENERIC_ARCH_NUMA=y,
CONFIG_NEED_PER_CPU_PAGE_FIRST_CHUNK=y.
The same setup_per_cpu_areas() fallback path still exists in
linux mainline drivers/base/arch_numa.c, so
this should reproduce on current mainline as well.
On a QEMU arm64 VM, add "percpu_alloc=page" to the kernel cmdline
to force pcpu_page_first_chunk(), so percpu lives in vmalloc.
Then:
# makedumpfile -d 31 /proc/kcore vmcore
# ./crash vmlinux vmcore
crash log:
crash: seek error: kernel virtual address: ffff800100000050 type: "IRQ stack pointer"
crash: seek error: kernel virtual address: ffff800100021050 type: "IRQ stack pointer"
crash: seek error: kernel virtual address: ffff800100042050 type: "IRQ stack pointer"
crash: seek error: kernel virtual address: ffff800100063050 type: "IRQ stack pointer"
...
KERNEL: vmlinux [TAINTED]
DUMPFILE: vmcore [PARTIAL DUMP]
CPUS: 4
RELEASE: 6.6.0.aarch64
MACHINE: aarch64 (unknown Mhz)
MEMORY: 4 GB
...
I tried the deferred approach you suggested, but it requires a
new flag, a new field, and a new helper, and only fixes the
irq_stack_ptr branch. So I looked at how x86_64 handles this.
x86_64_init() (x86_64.c:655-661) does not defer anything: it just
pre-fills vt->vmalloc_start and vt->kernel_pgd[] before the
percpu readmem in POST_GDB:
machdep->vmalloc_start = x86_64_vmalloc_start;
vt->vmalloc_start = machdep->vmalloc_start();
machdep->init_kernel_pgd();
...
x86_64_per_cpu_init(); /* readmem percpu, translation works */
arm64 registers the same callbacks but never invokes them eagerly,
leaving vt->vmalloc_start == 0 until vm_init() runs in POST_VM.
So I'd suggest we follow the x86_64 approach and pre-fill on arm64 as well.
functions in POST_GDB:
case POST_GDB:
...
if (!machdep->machspec->CONFIG_ARM64_KERNELPACMASK)
arm64_recalc_KERNELPACMASK();
vt->vmalloc_start = machdep->vmalloc_start(); /* NEW */
machdep->init_kernel_pgd(); /* NEW */
arm64_irq_stack_init();
arm64_overflow_stack_init();
arm64_stackframe_init();
break;
One caveat: arm64_init_kernel_pgd() uses OFFSET(mm_struct_pgd),
which is only initialised by vm_init() in POST_VM. Calling it at
POST_GDB triggers "crash: invalid structure member offset:
mm_struct_pgd" because OFFSET() FATALs on INVALID_OFFSET. Adding
an INVALID_MEMBER() guard lets it fall through to the
swapper_pg_dir fallback; the later POST_VM call goes through the
init_mm.pgd path as usual:
if (!kernel_symbol_exists("init_mm") ||
INVALID_MEMBER(mm_struct_pgd) || /* NEW */
!readmem(symbol_value("init_mm") + OFFSET(mm_struct_pgd), ...)) {
if (kernel_symbol_exists("swapper_pg_dir"))
value = symbol_value("swapper_pg_dir");
...
}
Patch
-----
diff --git a/arm64.c b/arm64.c
--- a/arm64.c
+++ b/arm64.c
@@ -788,6 +788,9 @@ arm64_init(int when)
if(!machdep->machspec->CONFIG_ARM64_KERNELPACMASK)
arm64_recalc_KERNELPACMASK();
+ vt->vmalloc_start = machdep->vmalloc_start();
+ machdep->init_kernel_pgd();
+
arm64_irq_stack_init();
arm64_overflow_stack_init();
arm64_stackframe_init();
@@ -1894,7 +1897,8 @@ arm64_init_kernel_pgd(void)
ulong value;
if (!kernel_symbol_exists("init_mm") ||
+ INVALID_MEMBER(mm_struct_pgd) ||
!readmem(symbol_value("init_mm") + OFFSET(mm_struct_pgd), KVADDR,
&value, sizeof(void *), "init_mm.pgd", RETURN_ON_ERROR)) {
if (kernel_symbol_exists("swapper_pg_dir"))
I'd lean towards this: it mirrors x86_64, adds no new state, and
fixes any percpu readmem() at POST_GDB, not just the irq_stack_ptr
branch. Let me know if this is acceptable, and I'll send a formal patch
Thanks
2 weeks, 6 days
Subject: [BUG/RFC] arm64: stack init functions fail in POST_GDB when percpu data is in vmalloc area
by chenhaixiang (A)
Subject: [BUG/RFC] arm64: stack init functions fail in POST_GDB when percpu data is in vmalloc area.
Hi,
While running crash on a large arm64 NUMA system, crash fails to
initialize the arm64 stack subsystems during arm64_init(POST_GDB) and
aborts startup with SEEK_ERROR on every percpu IRQ stack pointer.
Environment
-----------
- architecture : aarch64 (64K pages, CONFIG_ARM64_64K_PAGES=y)
- kernel : 6.x, large NUMA, many CPUs
- crash : current master (arm64.c)
Background on percpu allocation on this machine
-----------------------------------------------
On this arm64 machine the kernel's first percpu chunk does NOT live in
the linear map, it lives in the vmalloc area. The arm64 percpu setup
path in drivers/base/arch_numa.c:setup_per_cpu_areas() first tries
pcpu_embed_first_chunk() and, when that fails, falls back to
pcpu_page_first_chunk()
Observed symptom / reproduction log
-----------------------------------
With crash started on the vmcore, the percpu IRQ stack pointers are all in the
ffff800100xxxxxx range, which is the vmalloc area on this kernel.
Because vt->vmalloc_start is still 0 at POST_GDB, arm64_kvtop() takes
the early "VTOP()" shortcut and returns a bogus physical address far
beyond max_mapnr, so read_diskdump() reports SEEK_ERROR:
...
<readmem: ffff800100000050, KVADDR, "IRQ stack pointer", 8, (ROE), 1bfe2ef0>
<read_diskdump: addr: ffff800100000050 paddr: 800100000050 cnt: 8>
read_diskdump: SEEK_ERROR: paddr/pfn: 800100000050/800100000 max_mapnr: 60c000000
crash: seek error: kernel virtual address: ffff800100000050 type: "IRQ stack pointer"
<readmem: ffff800100021050, KVADDR, "IRQ stack pointer", 8, (ROE), 1bfe2ef8>
<read_diskdump: addr: ffff800100021050 paddr: 800100021050 cnt: 8>
read_diskdump: SEEK_ERROR: paddr/pfn: 800100021050/800100021 max_mapnr: 60c000000
crash: seek error: kernel virtual address: ffff800100021050 type: "IRQ stack pointer"
<readmem: ffff800100042050, KVADDR, "IRQ stack pointer", 8, (ROE), 1bfe2f00>
<read_diskdump: addr: ffff800100042050 paddr: 800100042050 cnt: 8>
read_diskdump: SEEK_ERROR: paddr/pfn: 800100042050/800100042 max_mapnr: 60c000000
crash: seek error: kernel virtual address: ffff800100042050 type: "IRQ stack pointer"
<readmem: ffff800100063050, KVADDR, "IRQ stack pointer", 8, (ROE), 1bfe2ef08>
<read_diskdump: addr: ffff800100063050 paddr: 800100063050 cnt: 8>
read_diskdump: SEEK_ERROR: paddr/pfn: 800100063050/800100063 max_mapnr: 60c000000
crash: seek error: kernel virtual address: ffff800100063050 type: "IRQ stack pointer"
... (one SEEK_ERROR per CPU)
Note the wrong physical addresses (800100xxxxxx) come from a direct
linear-map style translation (addr - PAGE_OFFSET) which is only valid
for the linear map, not for vmalloc-backed percpu memory.
Root cause
----------
In arm64_init() the three stack initializers are invoked in the
POST_GDB phase:
case POST_GDB:
...
if (!machdep->machspec->CONFIG_ARM64_KERNELPACMASK)
arm64_recalc_KERNELPACMASK();
arm64_irq_stack_init(); /* <-- readmem on percpu var */
arm64_overflow_stack_init(); /* <-- readmem on percpu var */
arm64_stackframe_init();
break;
These helpers end up calling readmem() on percpu addresses that, on
the machine described above, live in the vmalloc range. The
translation goes through arm64_kvtop() (arm64.c:1963), which for any
vmalloc address requires:
1. vt->vmalloc_start != 0 -> to reach the IS_VMALLOC_ADDR() branch
2. vt->kernel_pgd[0] -> to walk the kernel page tables
Both are populated by vm_init(), which runs in the POST_VM phase --
i.e. *after* POST_GDB. As a result, in arm64_kvtop() we hit:
if (!vt->vmalloc_start) {
*paddr = VTOP(kvaddr); /* wrong for vmalloc addresses */
return TRUE;
}
and the returned physical address is bogus, so the subsequent
readmem() fails with SEEK_ERROR (as shown in the log above: paddr
800100xxxxxx is the linear-map translation of a vmalloc address, and
it falls well beyond max_mapnr 0x60c000000).
Proposed fix
------------
Move the three stack initializers from POST_GDB to POST_VM, so that
vt->vmalloc_start and vt->kernel_pgd are guaranteed to be ready
before any readmem() that may touch vmalloc-backed percpu data.
diff --git a/arm64.c b/arm64.c
index d177092..654b972 100644
--- a/arm64.c
+++ b/arm64.c
@@ -787,7 +787,9 @@ arm64_init(int when)
*/
if(!machdep->machspec->CONFIG_ARM64_KERNELPACMASK)
arm64_recalc_KERNELPACMASK();
+ break;
+ case POST_VM:
arm64_irq_stack_init();
arm64_overflow_stack_init();
arm64_stackframe_init();
Testing
-------
Verified on:
- the failing large-NUMA arm64 system whose percpu first chunk was
created by pcpu_page_first_chunk() (i.e. percpu lives in vmalloc):
with the patch, the percpu IRQ stack pointers are translated
through the kernel page tables and crash starts cleanly. "bt" on
the panic task produces correct irq/overflow/standby stacks;
- a regular arm64 QEMU VM whose percpu first chunk was created by
pcpu_embed_first_chunk() (i.e. percpu lives in the linear map):
no regression, stack initialization still happens before any
"bt" command.
I'd like to get feedback on whether POST_VM is the preferred phase,
or whether the maintainers would rather guard the readmem() calls
inside the three helpers with a fallback (e.g. defer percpu
translation until later). I can send a formal patch with
Signed-off-by once the approach is agreed.
Thanks
3 weeks, 2 days
[announce] crash release 9.0.3
by Dave Young
Hi list:
Thank you all for your contributions to the crash-utility,
crash 9.0.3 is now available.
Download from:
https://crash-utility.github.io/
or
https://github.com/crash-utility/crash/releases
The GitHub master branch serves as a development branch that will
contain all patches that are queued for the next release:
$ git clone https://github.com/crash-utility/crash.git
The 9.0.3 release covers architecuter fixes and various kernel support
fixes eg. fixes for the mainline kernel 7.1 and 7.2
Changelog:
35457ac9015a update example nvr in README
1c8551c95cd1 rename .rh_rpm_package to crash-release
14f7004d0f77 Remove dead code of "BASELEVEL_REVISION"
cc2a942537da Prevent out-of-bounds access of note_buf
1491c42b7530 loongarch64: enable kaslr offset calculation
400874bb0351 loongarch64: fix SECTION_SIZE_BITS value
19d0582f4e3d task: show the NUMA information in task_cpu
bd0359714aa0 add cpu_to_nid function
fe61d650e135 riscv64: support leaf PTEs in page table walks
abcb45e7d443 symbols: use non-debug BFD to retrieve .rodata
9550178eca33 Refactor kernel version checking code
bb9a3fa67a79 Fix "kmem [-s]" command on Linux 7.1 and later
fa55e121672c Fix compound_head() and "bt -F" option on Linux 7.1 and later
08e9d02d2c46 symbols: optimize symval_hash_init with O(1) tail insertion
bfce25840cff x86_64: Make ORC_REG_SP and ORC_REG_PREV_SP independent from kernel version
f336dfa79ed9 (test-2) Fix "runq -g" option to display task_group name on Linux 6.15 and later
82039b0f21de Fix failure of "runq -g" option on Linux 7.2 and later kernels
97b7c0a60852 task: Introduce -Y option to ps command to display scheduling policy and priority
610e52023a42 task: Introduce -I option to ps command to exclude idle tasks
18d6e9f4852f riscv64: Add get_kvaddr_ranges callback for kernel address ranges
f5ce0c39b5b4 riscv64: Guard verbose output in vtop page table walk functions
f6084e0dbf36 riscv64: Set VMEMMAP flag to fix spurious mem_map[] warnings
2946d9bfa656 LoongArch64: print exception registers only with bt -f
a35fcc79393c LoongArch64: avoid replacing a valid RA with stack noise
a436361b2798 LoongArch64: unwind dumpfile active tasks from IRQ stacks
9e8d1784b910 LoongArch64: add initial ORC unwinder support
488777428c96 LoongArch64: resolve relocated exception vector addresses
bf279fbd474f LoongArch64: print exception return address as ERA
523981942d89 LoongArch64: Add dummy eframe_search to avoid bt -e segfault
40145cc1f001 LoongArch64: Fix stack frame loop bounds for exception frames
f3a02a49653b LoongArch64: Support backtracing across exception boundaries
a6817d475c75 LoongArch64: Fix pt_regs initialization for active tasks
32ac79ead778 LoongArch64: Fix CPU registers reading from dump notes
fdb94ed20ac8 Fix "swap" command on Linux 7.1 and later
c894f05d3cdc Fix "kmem -i" option to display swap usage on Linux 6.18 and later
d0ee428664f9 x86_64: Fix "bt" command to use correct ORC register values on Linux 7.1 and later
7b3f6e0f60be x86_64: Fix "bt" command for noreturn functions
28889183d3f5 add "files -n" command for an inode
5d494e101bdd xarray: add large folio support
744dfa3a190d add folio_order function
3a415c07a18a symbols: Add support for mod symtab with combined GPL and non-GPL symbols
7ff3146d46ea Fix get_xtime() for kernel 6.13 and higher
ad1e05f3d1ae RISCV64: add mapping symbol filter in riscv64_verify_symbol
ce0fc6b73de8 symbols: optimize symname_hash with larger table and FNV-1a hash
4f66cea2487d RISCV64: fix pmd address calculation in riscv64_vtop_4level_4k
fa4ba91e0383 arm64: fix gdb register feeding for exception frames and stack switching
5da4d159a5ef Fix unwinding with 32k stacks in ppc64le
c0097c84fed0 Fix wrong detection of "[LIVEPATCH]" on old kernels
4e4dc1c89310 RISCV64: introduce riscv64_VTOP() and riscv64_PTOV()
Full ChangeLog:
https://github.com/crash-utility/crash/compare/9.0.2...9.0.3
Thanks
Dave
3 weeks, 4 days
[PATCH 1/2] Remove dead code of "BASELEVEL_REVISION"
by Dave Young
Since commit e30a8743780a the BASELEVEL_REVISION is not used any more
Thus remove the dead code.
Signed-off-by: Dave Young <yangrr.2009(a)tsinghua.org.cn>
---
configure.c | 18 ------------------
1 file changed, 18 deletions(-)
diff --git a/configure.c b/configure.c
index e0626318b8d5..5eeddfe34135 100644
--- a/configure.c
+++ b/configure.c
@@ -653,24 +653,6 @@ get_release:
} else
fprintf(stderr,
"WARNING: .rh_rpm_package file does not exist!\n");
-
- if ((fp = fopen("defs.h", "r")) == NULL) {
- perror("defs.h");
- return;
- }
-
- while (fgets(buf, 512, fp)) {
- if (strncmp(buf, "#define BASELEVEL_REVISION",
- strlen("#define BASELEVEL_REVISION")) == 0) {
- p = strstr(buf, "\"") + 1;
- strip_linefeeds(p);
- p[strlen(p)-1] = '\0';
- strcpy(target_data.release, p);
- break;
- }
- }
-
- fclose(fp);
}
void
--
2.51.1
3 weeks, 4 days
Could you release 9.0.3?
by Jiri Slaby
Hi,
there are many post 9.0.2 changes to make kernels 7.1 and 7.2 working.
Could you release 9.0.3 with those?
thanks,
--
js
suse labs
3 weeks, 6 days
[PATCH] riscv64: handle leaf page table entries in vtop
by Rui Qi
RISC-V page table entries at any level may be leaf entries when V is
set and any of R/W/X is set. The riscv64 vtop walkers currently treat
non-zero intermediate entries as pointers to the next page table, so a
PMD leaf can be used as a page table address and make vtop fail with a
physical seek error.
Detect leaf entries at PGD/P4D/PUD/PMD levels and translate them using
the page size implied by that level. Continue walking only for non-leaf
table entries.
Signed-off-by: Rui Qi <qirui.001(a)bytedance.com>
---
riscv64.c | 54 ++++++++++++++++++++++++++++++++++++++++++++++++++++++
1 file changed, 54 insertions(+)
diff --git a/riscv64.c b/riscv64.c
index 4dcecc2f6b2c..ca26965421d7 100644
--- a/riscv64.c
+++ b/riscv64.c
@@ -602,6 +602,33 @@ riscv64_translate_pte(ulong pte, void *physaddr, ulonglong unused)
return page_present;
}
+static int
+riscv64_pte_is_leaf(ulong pte)
+{
+ return ((pte & _PAGE_PRESENT) &&
+ (pte & (_PAGE_READ | _PAGE_WRITE | _PAGE_EXEC)));
+}
+
+static int
+riscv64_handle_leaf_pte(ulong pte, ulong vaddr, int shift,
+ physaddr_t *paddr, int verbose)
+{
+ physaddr_t paddr_base, page_mask;
+
+ pte &= PTE_PFN_PROT_MASK;
+ paddr_base = PTOB(pte >> _PAGE_PFN_SHIFT);
+ page_mask = ~(((physaddr_t)1 << shift) - 1);
+ *paddr = (paddr_base & page_mask) + (vaddr & ~page_mask);
+
+ if (verbose) {
+ fprintf(fp, " PAGE: %016llx\n\n",
+ (ulonglong)(*paddr & page_mask));
+ riscv64_translate_pte(pte, 0, 0);
+ }
+
+ return TRUE;
+}
+
static void
riscv64_page_type_init(void)
{
@@ -655,6 +682,9 @@ riscv64_vtop_3level_4k(ulong *pgd, ulong vaddr, physaddr_t *paddr, int verbose)
fprintf(fp, " PGD: %lx => %lx\n", (ulong)pgd_ptr, pgd_val);
if (!pgd_val)
goto no_page;
+ if (riscv64_pte_is_leaf(pgd_val))
+ return riscv64_handle_leaf_pte(pgd_val, vaddr,
+ PGD_SHIFT_L3, paddr, verbose);
pgd_val &= PTE_PFN_PROT_MASK;
pmd_base = (pgd_val >> _PAGE_PFN_SHIFT) << PAGESHIFT();
@@ -666,6 +696,9 @@ riscv64_vtop_3level_4k(ulong *pgd, ulong vaddr, physaddr_t *paddr, int verbose)
fprintf(fp, " PMD: %016lx => %016lx\n", pmd_addr, pmd_val);
if (!pmd_val)
goto no_page;
+ if (riscv64_pte_is_leaf(pmd_val))
+ return riscv64_handle_leaf_pte(pmd_val, vaddr,
+ PMD_SHIFT, paddr, verbose);
pmd_val &= PTE_PFN_PROT_MASK;
pte_base = (pmd_val >> _PAGE_PFN_SHIFT) << PAGESHIFT();
@@ -1237,6 +1270,9 @@ riscv64_vtop_4level_4k(ulong *pgd, ulong vaddr, physaddr_t *paddr, int verbose)
fprintf(fp, " PGD: %lx => %lx\n", (ulong)pgd_ptr, pgd_val);
if (!pgd_val)
goto no_page;
+ if (riscv64_pte_is_leaf(pgd_val))
+ return riscv64_handle_leaf_pte(pgd_val, vaddr,
+ PGD_SHIFT_L4, paddr, verbose);
pgd_val &= PTE_PFN_PROT_MASK;
pud_base = (pgd_val >> _PAGE_PFN_SHIFT) << PAGESHIFT();
@@ -1248,6 +1284,9 @@ riscv64_vtop_4level_4k(ulong *pgd, ulong vaddr, physaddr_t *paddr, int verbose)
fprintf(fp, " PUD: %016lx => %016lx\n", pud_addr, pud_val);
if (!pud_val)
goto no_page;
+ if (riscv64_pte_is_leaf(pud_val))
+ return riscv64_handle_leaf_pte(pud_val, vaddr,
+ PUD_SHIFT, paddr, verbose);
pud_val &= PTE_PFN_PROT_MASK;
pmd_base = (pud_val >> _PAGE_PFN_SHIFT) << PAGESHIFT();
@@ -1259,6 +1298,9 @@ riscv64_vtop_4level_4k(ulong *pgd, ulong vaddr, physaddr_t *paddr, int verbose)
fprintf(fp, " PMD: %016lx => %016lx\n", pmd_addr, pmd_val);
if (!pmd_val)
goto no_page;
+ if (riscv64_pte_is_leaf(pmd_val))
+ return riscv64_handle_leaf_pte(pmd_val, vaddr,
+ PMD_SHIFT, paddr, verbose);
pmd_val &= PTE_PFN_PROT_MASK;
pte_base = (pmd_val >> _PAGE_PFN_SHIFT) << PAGESHIFT();
@@ -1316,6 +1358,9 @@ riscv64_vtop_5level_4k(ulong *pgd, ulong vaddr, physaddr_t *paddr, int verbose)
fprintf(fp, " PGD: %lx => %lx\n", (ulong)pgd_ptr, pgd_val);
if (!pgd_val)
goto no_page;
+ if (riscv64_pte_is_leaf(pgd_val))
+ return riscv64_handle_leaf_pte(pgd_val, vaddr,
+ PGD_SHIFT_L5, paddr, verbose);
pgd_val &= PTE_PFN_PROT_MASK;
p4d_base = (pgd_val >> _PAGE_PFN_SHIFT) << PAGESHIFT();
@@ -1327,6 +1372,9 @@ riscv64_vtop_5level_4k(ulong *pgd, ulong vaddr, physaddr_t *paddr, int verbose)
fprintf(fp, " P4D: %016lx => %016lx\n", p4d_addr, p4d_val);
if (!p4d_val)
goto no_page;
+ if (riscv64_pte_is_leaf(p4d_val))
+ return riscv64_handle_leaf_pte(p4d_val, vaddr,
+ P4D_SHIFT, paddr, verbose);
p4d_val &= PTE_PFN_PROT_MASK;
pud_base = (p4d_val >> _PAGE_PFN_SHIFT) << PAGESHIFT();
@@ -1338,6 +1386,9 @@ riscv64_vtop_5level_4k(ulong *pgd, ulong vaddr, physaddr_t *paddr, int verbose)
fprintf(fp, " PUD: %016lx => %016lx\n", pud_addr, pud_val);
if (!pud_val)
goto no_page;
+ if (riscv64_pte_is_leaf(pud_val))
+ return riscv64_handle_leaf_pte(pud_val, vaddr,
+ PUD_SHIFT, paddr, verbose);
pud_val &= PTE_PFN_PROT_MASK;
pmd_base = (pud_val >> _PAGE_PFN_SHIFT) << PAGESHIFT();
@@ -1349,6 +1400,9 @@ riscv64_vtop_5level_4k(ulong *pgd, ulong vaddr, physaddr_t *paddr, int verbose)
fprintf(fp, " PMD: %016lx => %016lx\n", pmd_addr, pmd_val);
if (!pmd_val)
goto no_page;
+ if (riscv64_pte_is_leaf(pmd_val))
+ return riscv64_handle_leaf_pte(pmd_val, vaddr,
+ PMD_SHIFT, paddr, verbose);
pmd_val &= PTE_PFN_PROT_MASK;
pte_base = (pmd_val >> _PAGE_PFN_SHIFT) << PAGESHIFT();
--
2.20.1
1 month