Skip to content

Keep PyLong loop carry at two-digit width in shift and division #156443

Description

@XiaohongGong

Feature or enhancement

Proposal:

Several functions in longobject.c (v_lshift, v_rshift, and x_divrem) narrowed a loop carry value to digit or sdigit, then widened it again on the next iteration.

On AArch64, that 64-to-32-to-64 conversion inserts an extra mov on the loop-carried critical path. Keeping the carry at two-digit width until the function returns drops that mov and shortens the carry chain. Results are unchanged for valid limbs.

Use v_lshift on AArch64 as an example. The code inside the loop is:

  ldr  w0, [x4, x2, lsl #2]    ; a[i]
  mov  w3, w3                  ; carry chain
  lsl  x0, x0, x24             ; a[i] << d
  orr  x0, x0, x3              ; carry chain
  and  w1, w0, #0x3fffffff
  ubfx x3, x0, #30, #32        ; carry chain
  str  w1, [x26, x2, lsl #2]
  add  x2, x2, #1
  cmp  x25, x2
  b.ne

removing the narrowing, the code will be optimized to:

  ldr  w0, [x4, x2, lsl #2]    ; a[i]
  lsl  x0, x0, x24             ; a[i] << d
  orr  x0, x0, x3              ; carry chain
  and  w1, w0, #0x3fffffff
  str  w1, [x26, x2, lsl #2]
  add  x2, x2, #1
  lsr  x3, x0, #30             ; carry chain
  cmp  x25, x2
  b.ne

The instruction count on the loop carried chain is reduced from 3 to 2.

Has this already been discussed elsewhere?

No response given

Links to previous discussion of this feature:

No response

Metadata

Metadata

Assignees

No one assigned

    Labels

    interpreter-core(Objects, Python, Grammar, and Parser dirs)pendingThe issue will be closed if no feedback is providedperformancePerformance or resource usagetype-featureA feature request or enhancement

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions