Skip to content

Repository files navigation

armlint

armlint examines AArch64 machine code to find suboptimal instruction sequences. For example, building the constant 0x66666666 as

movz w0, #0x6666
movk w0, #0x6666, lsl #16

is two instructions where one would do, because 0x66666666 is encodable as an AArch64 logical (bitmask) immediate:

mov w0, #0x66666666     ; orr w0, wzr, #0x66666666

armlint helps compiler writers and assembly authors generate tighter code, and documents corners of the A64 instruction set.

Design and limitations

armlint is a peephole analyzer. It decodes each 32-bit A64 instruction directly from the binary and matches it by mask and value, resolving aliases (for example MUL is MADD with a zero accumulator) so that both spellings of a pattern are caught. It then looks for a short window of adjacent instructions that a shorter or cheaper encoding can replace.

The overriding rule is soundness: armlint emits a finding only when the rewrite provably preserves the architectural result. For a tool that suggests code changes, a false positive is the worst failure, so it errs toward false negatives -- a missed opportunity is cheaper than a wrong one. Each check documents the exact conditions under which its rewrite is equivalent; the constraints below are the ones they share, and analyses.md's appendix collects the near-miss folds that are deliberately never matched (FP contraction, fcsel -> fmax, the SDIV remainder, and friends) with the argument against each.

  • Strict adjacency for matching; bounded lookahead for proof. The instructions of a matched pattern must be consecutive -- an unrelated instruction between a producer and its consumer suppresses the finding -- and armlint does not reorder code or match through intervening instructions. Emission, though, is not confined to the pattern window: as the following bullets describe, many findings are held back while a bounded forward scan (16 instructions) walks the fall-through path past the consumer to prove a register or the flags dead, so an instruction well after the pair can also suppress a finding. The scan only gates emission -- it never widens a match, so every reported rewrite still replaces only the adjacent instructions shown.
  • Liveness is proved structurally, or by a bounded forward scan. A producer-into-consumer fold fires when the consumer overwrites the producer's destination register, proving the intermediate value is dead. Folds whose saving is a deleted write with no such overwrite defer instead: a bounded forward scan of the fall-through path must see the register overwritten before any read or control transfer. This is how the address folds admit stores and loads into a fresh register -- add x8, sp, #32 ; str x0, [x8] folds to str x0, [sp, #0x20] only once x8 provably dies -- and how the producer folds (shift, funnel, extend, MUL/SMULL, NEG, MVN) admit consumers that write a register other than the producer's: lsl w8, w1, #3 ; add w9, w2, w8 folds to add w9, w2, w1, lsl #3 under the same proof. The single-bit and CSET branch folds additionally require the folded branch's taken edge to land inside that proven-clean span -- a general-purpose register, unlike NZCV, is routinely live into a branch target, so no block-locality assumption is made for it.
  • MOV-chain folds verify the constant dies. Folds that absorb a materialized constant -- MUL/MNEG/UDIV by a constant, MOV + ADD/AND/ORR/EOR/CCMP/FMOV, MOV #0, the register-offset and MOVI-zeroing folds -- report only once the consumer's own overwrite or the forward scan proves the constant register dead. The consumer rewrite itself stays valid regardless.
  • Flag liveness uses a bounded forward scan. The branch- and flag-folding checks drop a CMP/TST only after a bounded scan of the fall-through path confirms that no later instruction reads N/C/V before they are overwritten. Every NZCV reader is recognized -- the integer conditionals (B.cond, CSEL/CSINC/..., CCMP/CCMN, ADC/SBC) and the floating-point ones (FCSEL, FCCMP/FCCMPE) -- and any branch off the path whose destination the scan cannot see -- an unconditional B, or a conditional CBZ/CBNZ/TBZ/TBNZ (which do not themselves touch NZCV but whose taken target may still observe it) -- ends it conservatively. The scan does not follow the folded branch's own taken edge, so most of these folds assume N/C/V is dead at every branch target. That holds for most compiled code, where the flags are defined within a basic block, but not for hand-written assembly that deliberately keeps a flag live into a branch target -- including the stubs a JIT inlines into its output. The folds that meet such code prove the taken edge with a second scan starting at the target: the compare-and-branch fold (-m cmpbr), whose producer is a general two-register compare that clang's three-way comparator really does read again at the branch target; the CCMP chain and CMP-to-CMN folds; and the zero-compare CBZ/CBNZ and sign-bit TBZ/TBNZ folds, since HotSpot's lock fast paths return their result to the branch target in the flags. Unlike the fall-through scan, the one at the target follows a B to its destination and a CBZ/CBNZ/TBZ/TBNZ down both edges, within fixed bounds.

Findings are opportunities, not guaranteed speedups: some -- the pre- and post-indexed addressing folds -- are code-size and front-end wins that are backend-neutral. Each check's notes say what its rewrite actually saves.

Implemented analyses

Each row links to its full description -- mechanics, soundness, and what the rewrite saves -- in analyses.md. Candidate checks not yet implemented live in TODO.md.

Pattern Rewrite
movz/movn + movk (over-long constant) bitmask-immediate mov, or minimal movz/movn + movk chain
lsl/lsr/asr/ror + add/sub/and/orr/eor add Rd, Rn, Rm, <shift> #n
lsl/lsr + shifted orr/eor/add (funnel/rotate) extr Rd, Rhi, Rlo, #lsb (or ror when both halves match)
sxtw/uxtb/sxtb + add/sub add Rd, Rn, Wm, sxtw
cmp/cmn/tst zero-test + b.eq/b.ne (b.hi/b.ls after cmp) cbz/cbnz
cmp/cmn/tst zero-test + b.lt/b.ge/b.mi/b.pl tbnz/tbz Rn, #(msb)
tst #(1<<k) + b.eq/b.ne tbz/tbnz Rn, #k
tst #(1<<k) + cset/csetm ubfx/sbfx Rd, Rn, #k, #1
single-bit and/ubfx/lsr #31 + cbz/cbnz tbz/tbnz Rs, #k
cset + cbz/cbnz b.<cond> / b.<inverse cond>
cset + eor #1 cset <inverse cond>
cset + neg csetm
br x30 ret (engages the return-address predictor)
branch to the next instruction delete (both outcomes fall through; bl excluded)
lsl + lsr/asr ubfx/sbfx/ubfiz/sbfiz
lsr/asr + and #mask ubfx
and #mask + lsr ubfx
and #mask/uxtb/uxth/uxtw/mov + lsl (or lsr + lsl) ubfiz (or clearing and)
zeroing producer + uxtb/uxth/uxtw/and/mov Wd, Wd, adjacent or not drop the zero-extension
mov xd, xd remove (architectural no-op)
instruction recomputing a value its register or NZCV already holds, across a gap delete (an ADR/ADRP of a held page is reported apart: a relink fix)
register write overwritten before any read, across a gap delete the write
conditional branch the path to it already decides delete it (never taken) or b (always taken)
ALU operation whose inputs are all known values mov Rd, #result, or delete when Rd already holds it
and #mask/ubfx #0 of a value its known bits already keep inside the mask delete in place, mov Rd, Rn out of place
sign-extending producer + sxtb/sxth/sxtw drop the sign-extension
and/uxt*/sxt*/mov Wd, Wm + a narrower and/ands/tst/uxt*/sxt* the second reading the first's source, or one and of the intersected masks
and/orr/eor/sub/bic/orn/eon with Rs, Rs mov / zero / all-ones
ldr+ldr / str+str (consecutive) ldp/stp (integer or FP/SIMD, and ldpsw); notes where a slot is in flight, since Apple cores do not forward a store to a pair load or a pair store to a load
str wzr+str wzr (consecutive zero stores) str/stur xzr
stp wzr, wzr (W-form) str/stur xzr
movi #0 + vector cmeq/cmge/cmgt (or FP fcm*) cmeq/cmge/cmgt/cmle/cmlt Vd, X, #0 (drop the movi)
and + and/ubfiz + orr (clear/isolate/merge) bfxil/bfi
and #lowmask + orr Rm, lsl #k (or high mask + lsr; field reaches the top) bfi/bfxil
csel Rd, Rn, Rn, cond mov Rd, Rn
fcsel Vd, Vn, Vn, cond fmov Vd, Vn
add/sub Rd, Rn, #0 mov Rd, Rn, or remove
add/sub #a + add/sub #b (same register) one add/sub carrying the sum (or mov)
adds/subs/ands + cmp #0 + b.eq/b.ne drop the redundant cmp/tst
add/sub/and/bic + cmp #0 + b.eq/b.ne adds/subs/ands/bics (drop the cmp/tst)
sub + cmp / add + cmn of the same operands (either order) subs/adds (flag-exact; drop the compare)
subs/adds + cmp/cmn of its own operands drop the compare (flags already set)
cmp/cmn/tst/ccmp whose flags are overwritten unread delete (the compare writes no register)
cmp #0 + cset/csetm of lt/mi lsr/asr Rd, Rn, #(msb) (the sign bit; drop the compare)
mov #2^N + mul lsl, or add Rd, Ra, Ra, lsl #N
mov #C + mneg neg, or shifted neg/sub
mov #2^N + madd/msub add/sub Rd, Ra, Rn, lsl #N
mov #2^N + udiv lsr
mov #2^N + udiv + msub (remainder) and Rd, Rn, #(2^N-1)
mov #C + add/sub add/sub Rd, Rn, #C (sign-crossed add↔sub, cmp↔cmn for #-C)
movz/movn + movk of a constant below 2^24 + add/sub add/sub Rd, Rn, #hi, lsl #12 + add/sub Rd, Rd, #lo (add↔sub for a negative constant)
mov #C + and/orr/eor/ands or bic/orn/eon/bics and/orr/eor/ands Rd, Rn, #C (#~C for the inverting forms)
mov #INT_MIN or #INT_MAX + cmp + b.eq/b.ne cmp xzr, Xn or cmn Xn, #1 + b.vs/b.vc
mov #C + ccmp/ccmn ccmp/ccmn Rn, #C, #nzcv, cond (sign-crossed for #-C)
cmp/cmn + ccmp/ccmn on one register + b.cond/csel, one compare deciding it that cmp/cmn (or cbz/cbnz) with the reader's condition adjusted
mov #(2^32-k) + cmp Xn of a zero-extended value, read only through Z and C cmn Wn, #k
mov #1 + csel csinc Rd, Rn, wzr, cc (cset when the other operand is ZR)
mov #-1 + csel csinv Rd, Rn, wzr, cc (csetm when the other operand is ZR)
mov #C + lsl/lsr/asr/ror (register amount) immediate-form shift, amount C mod 32/64
mov #bits/#C + fmov/scvtf/ucvtf from GPR fmov Sd/Dd, #imm
fmov/scvtf/dup of wzr/xzr (or mov #0 + transfer) movi dN, #0 / movi vN.T, #0
sxtw/mov w, w + scvtf/ucvtf Xn scvtf/ucvtf of Wn
ldr w8 + scvtf/ucvtf from GPR ldr s0 + FP-side convert (no cross-file transfer)
ldr/ldrsw (literal) of an encodable constant mov #imm / fmov #imm8 / movi+mvni for Q (no memory access)
adr + ldr [x8]/br x16 ldr Rt, <literal> / b L
umov Wd, Vn.s[0] / Xd, Vn.d[0] fmov Wd, Sn / Xd, Dn (cheaper port; .h[0] under -m fp16)
and Xd, Xn, #0xffffffff / ubfx Xd, Xn, #0, #32 mov Wd, Wn (a rename on Neoverse rather than an ALU op)
cmp + csel (max/min shape) smax/smin/umax/umin (-m cssc)
cmp #0 + cneg abs (-m cssc)
rbit + clz ctz (-m cssc)
NEON popcount round trip cnt Xd, Xn (-m cssc)
cmp + b.<cond> cb<cond> Rn, Rm/#imm, label (-m cmpbr)
eor + eor (16B vectors) eor3 Vd, Vn, Vm, Va (-m sha3)
bic + eor (16B vectors) bcax Vd, Vn, Vm, Va (-m sha3)
autiasp/autibsp + ret retaa/retab (-m pauth; auto-armed on arm64e)
unsigned LR spill; raw br/blr; zero-discriminator braaz/blraaz audit-only review items (-a pac; auto-armed on arm64e)
mov chain + cmp/tst/ccmp/ldr whose immediate form the value misses audit-only review items, tallied by value, the constants within one step of an immediate first (-a imm)
ldxr/stxr fetch-op retry loop ldadd/ldset/ldeor/ldclr (+ mvn/neg/mov pre-op) (-m lse)
ldxr/stxr exchange retry loop swp (-m lse)
ldxr + cmp + b.ne + stxr CAS retry loop mov + cas + cmp (-m lse)
mov #C + orr Xd, x28, Xc (V8 cage base) add Xd, x28, #C (-m v8)
fmul + in-place fneg fnmul (bit-exact in every rounding mode)
mov #0 + str/add/and/csel/ccmp use use wzr/xzr
mov #C + ldr/str [xn, xc] ldr/str [xn, #C] (or ldur/stur)
mul + add/sub madd/msub (or mneg)
smull/umull + add/sub smaddl/umaddl/smsubl/umsubl
neg + add/sub sub/add
neg + csel csneg (inverted cond for the then slot)
mvn + csel csinv (inverted cond for the then slot)
add #1 + csel csinc (inverted cond for the then slot)
mvn + and/orr/eor/ands bic/orn/eon/bics
add + ldr/str [xt] ldr/str [xn, xm{, lsl #s}]
sxtw + ldr/str [xn, xt] ldr/str [xn, ws, sxtw {#s}]
ldrb/ldrh/ldr (or ldrsb/ldrsh Wt) + sxtb/sxth/sxtw ldrsb/ldrsh/ldrsw (Xt for the re-widened sign loads)
add #a + ldr/str [xt] (incl. mov xt, sp) ldr/str [xn, #a] / ldr [sp]
ldr/ldp [xn] + add/sub xn ldr [xn], #±imm / ldp [sp], #imm (post-index)
add/sub xn + ldr/stp [xn] ldr [xn, #±imm]! / stp [sp, #-imm]! (pre-index)

Compilation

armlint depends on Capstone and uses pkg-config to locate it. On macOS:

brew install capstone

On Debian/Ubuntu:

apt install libcapstone-dev pkg-config

Build:

git clone https://github.com/gaul/armlint.git armlint
cd armlint
make all

Two test suites are available. make test runs the unit tests against fabricated byte sequences, exercising the check registry directly. make integration-test runs the snapshot suite under fixtures/: each .s is assembled with clang -arch arm64 and armlint's output is diffed against a checked-in .expected file. The integration suite covers the Mach-O parser and the report formatting, which the unit tests bypass. It needs a clang that can assemble AArch64 and fails without one rather than reporting a pass it did not earn; off an arm64 host it names the target explicitly, so an x86-64 Linux box runs the suite too. After an intentional output change, regenerate the snapshots with make integration-test-regen and review the diff before committing.

Setting ARMLINT_LIVENESS_SWEEP=1 extends the unit tests' liveness cross-checks against Capstone to the entire 2^32 encoding space -- minutes of single-threaded CPU time, so CI runs it on pushes rather than PRs. Both classifiers are covered off one decode per word: the NZCV one (classify_liveness) and the register one (classify_reg_liveness), whose property is that a register Capstone reports as read must never be classified as dead or as no-effect. Three Capstone over-reports are documented and skipped; everything else is a defect on one side or the other, and the register half found four on its first run. ARMLINT_LIVENESS_SWEEP_THREADS divides the sweep across that many worker threads:

ARMLINT_LIVENESS_SWEEP=1 ARMLINT_LIVENESS_SWEEP_THREADS="$(sysctl -n hw.ncpu)" \
    make test

(nproc on Linux.)

Testing against Capstone 6

Capstone 6 (in alpha as of 2026) rewrote the AArch64 module from LLVM. armlint compiles against it unchanged via Capstone's compatibility header, and CI tracks the 6.0.0-Alpha11 tag. Before Alpha11 it tracked a commit instead: Alpha10 still leaves insn->alias_id stale on non-alias instructions (the sweep's oracle then reports a phantom x30 read after every ret) and still mis-reports the SYSL and integer-STLUR operands, and Alpha11 is the first tag with all three fixed. To reproduce locally:

git clone --depth 1 --branch 6.0.0-Alpha11 \
    https://github.com/capstone-engine/capstone.git capstone6
cmake -B capstone6/build -S capstone6 -DCMAKE_BUILD_TYPE=Release
cmake --build capstone6/build -j8
make clean   # never mix objects built against different Capstone ABIs
make CAPSTONE_CFLAGS="-I$PWD/capstone6/include -DCAPSTONE_AARCH64_COMPAT_HEADER" \
     CAPSTONE_LIBS="$PWD/capstone6/build/libcapstone.a" all

The overrides must be make arguments (not environment variables) to beat the Makefile's pkg-config defaults. The unit tests and the ARMLINT_LIVENESS_SWEEP=1 sweep pass under both major versions, and running the sweep under both is the point rather than a formality: the two model AArch64 register access differently, so each version corroborates cases the other cannot. Capstone 6's model is the sharper oracle -- it is what caught armlint treating a FEAT_MOPS main stage as a kill of its own source pointer -- while 5.x is what armlint ships against and so is what its corrections are written for; make integration-test is expected to show cosmetic snapshot diffs under v6 -- 22 of the 100 under Alpha11, which prints shift immediates in decimal and drops # on PC-relative operands (branch targets, adr, ldr literals), with every finding identical -- so the fixtures remain pinned to Capstone 5.x rendering until v6 stabilizes.

Usage

armlint is intended to be part of compiler test suites which should #include "armlint.h" and link libarmlint.a. Disassemble the just-emitted machine code with check_instructions; its return value is the number of opportunities found, which a test can assert is zero:

#include "armlint.h"   // also includes <capstone/capstone.h>

// code/code_len: the AArch64 bytes to check (e.g. a function the
// compiler just emitted); base_addr is the address they load at.
// Returns the opportunity count (0 == clean), or -1 on a decode error.
int lint(const uint8_t *code, size_t code_len, uint64_t base_addr)
{
    csh handle;
    if (cs_open(CS_ARCH_ARM64, CS_MODE_ARM, &handle) != CS_ERR_OK) {
        return -1;
    }
    cs_option(handle, CS_OPT_DETAIL, CS_OPT_ON);

    armlint_summary *summary = armlint_summary_create();
    int findings = check_instructions(
        handle, code, code_len, base_addr, /*verbose=*/true, summary,
        /*features=*/0,    // or ARMLINT_FEATURE_CSSC etc.
        /*symbols=*/NULL, /*nsymbols=*/0);   // see armlint_symbol
    armlint_summary_print(summary);   // optional by-type tally

    armlint_summary_destroy(summary);
    cs_close(&handle);
    return findings;
}

The summary is optional -- pass NULL to skip the by-type tally -- and verbose controls whether each opportunity is printed as it is found. armlint can also read arbitrary AArch64 binaries (ELF, thin Mach-O, or universal/fat Mach-O) directly:

./armlint /path/to/aarch64/binary
./armlint /bin/ls
./armlint -m cssc /bin/ls   # also suggest CSSC instructions
./armlint -m v8 jit.elf     # a V8 JIT dump from tools/v8dump2elf.py
./armlint -s all /bin/bash  # every ARM64 slice of a universal binary
./armlint -d /bin/ls        # also report functions that copy one another
./armlint -c /bin/ls        # also report the constants MOVZ/MOVK chains build

A universal binary can carry more than one ARM64 slice. macOS 27's /bin/bash, /usr/bin/ssh and /usr/lib/dyld each hold an arm64e and an arm64e.x1 slice, two compilations of the same program, so armlint scans one slice by default -- the lowest variant, arm64 before arm64e before arm64e.x1 -- and names the slices it skipped on stderr. -s all scans every ARM64 slice and adds up their findings; -s NAME (-s arm64e.x1) or -s INDEX (the slice's position in the fat header, as otool -f numbers it) picks one. arm64e.x1 code signs and authenticates return addresses with FEAT_PAuth_LR (pacibsppc, retabsppc), which Capstone 5 does not decode: a scan of that slice skips those words as data, and the PAC audit and -m pauth do not arm themselves there.

Linker-synthesized import glue is excluded from both the scan and the census: Mach-O __stubs, __stub_helper, and __objc_stubs, and ELF .plt, .iplt, and the .plt.* variants. Their shape is the dynamic-linking ABI's business, not the compiler's -- the classic lazy-binding __stub_helper alone would otherwise contribute one spurious "LDR literal foldable" finding per imported symbol (each fixed entry LDRs its lazy-bind-info offset into w16 from an inline literal), a couple of hundred lines of noise on a typical minos < 12 Mach-O binary.

-m <feature> enables checks whose rewrites use ISA-extension instructions the target must support: cssc (Armv8.9/9.4 Common Short Sequence Compression: smax/smin/umax/umin, abs, ctz), lrcpc2 (Armv8.4 unscaled store-release: stlur), pauth (Armv8.3 pointer authentication: retaa/retab), lse (Armv8.1 atomics: ldadd/ldset/ldeor/ldclr/swp), cmpbr (Armv9.6 compare-and-branch: cbgt/cbeq/...), sha3 (Armv8.2 three-operand vector logic: eor3, bcax), and fp16 (FEAT_FP16 half-precision transfers: fmov Wd, Hn). pauth arms automatically on arm64e slices, whose ABI mandates FEAT_PAuth (the same auto-arm as the PAC audit); the rest stay opt-in. cssc, sha3 and fp16 name extensions that are never mandatory at any architecture version, so those three assert a specific target rather than a version floor.

-m v8 is V8 the JavaScript engine, not an architecture version. It names an input rather than a target: a V8 JIT dump, the output of tools/v8dump2elf.py for a pointer-compressed build, and asserts the two facts about that stream that its checks need -- x28 is the 4GB-aligned pointer-compression cage base, and every ldr xzr, (literal) word opens a constant pool the scans step over. Both are knowledge about the scanned code rather than the hardware and unsound for arbitrary binaries, so the mode stays opt-in.

-a <audit> enables opt-in informational checks whose findings are review items rather than missed folds; pac audits the binary against the arm64e-style full pointer-authentication contract (return addresses spilled unsigned, unauthenticated br/blr, and zero-discriminator braaz/blraaz -- authenticated, but against a modifier of zero, so any same-key zero-discriminator pointer in the process substitutes). Audit findings are review items on a ladder: the raw-br check recognizes and auto-dismisses the clang jump-table idiom, so what remains is BLRs, linker veneers, and genuinely unclassified branches; the zero-discriminator rung marks where the C-ABI IA+0 signing floor could upgrade to __ptrauth-style diversified braa/blraa. The PAC audit arms automatically on arm64e slices (whose ABI already assumes full signing), so macOS system binaries surface their worklist with no flag; a plain arm64 slice never opted in, so it stays silent unless you pass -a pac explicitly. imm audits every materialized constant against its consumer's immediate encoding and tallies the misses by value, so the constants worth renumbering surface; see Immediate-misfit audit (-a imm) below.

By default armlint prints only a summary: the opportunities grouped by type and sorted by prevalence, so it is clear which to look at first, followed by a total and the number of instructions scanned. A large binary can have hundreds of thousands of opportunities, so the per-opportunity detail is suppressed unless requested:

$ ./armlint /bin/ls
Optimization opportunities by type:
      38  ADD + LDR foldable to immediate-offset LDR
       2  AUTIASP/AUTIBSP + RET foldable to RETAA/RETAB (PAuth)
       2  CBZ/CBNZ of a live single-bit test foldable to TBZ/TBNZ

42 optimization opportunities in 4153 instructions

Pass -v to also print each opportunity -- its one-line summary plus the offending instructions, as shown below -- ahead of the summary:

$ ./armlint -v /bin/ls
ADD + LDR foldable to immediate-offset LDR at offset: 0x60 <0x10000071c+0x44>: -> ldr w8, [x8, #0x2c] (2 instructions)
  add x8, x8, #0x2c
  ldr w8, [x8]
...

The <...> names the containing function so findings can be triaged per function: <_addhistnode+0x58> when the symbol table carries a name (Mach-O nlist, or the ELF .symtab with a .dynsym fallback), or the function's start address as above when a stripped binary's LC_FUNCTION_STARTS still records its boundaries. A binary carrying neither -- Go's linker, for one, emits its runtime symtab instead of either structure -- prints the historic unannotated form, as does a finding past the end of an ELF symbol whose recorded size stops short of it: a stripped libxul.so keeps only its exports, and the last one before 40 MB of unnamed code must not claim all of it.

The process exits non-zero when any opportunity is found, so armlint can gate a compiler test suite.

ISA census (-i)

armlint -i replaces the lint scan with a census: every instruction in the binary's executable sections, attributed to the FEAT_* group it requires, each group carrying the architecture version it became mandatory at -- LSE and CRC32 at Armv8.1, LRCPC/FCMA/JSCVT and the register-form PAC at 8.3, DotProd/LRCPC2/FlagM at 8.4, up through the MOPS memcpy instructions at 8.8. The resulting ladder answers "what was this binary compiled for": a -march=armv8.1-a build shows LSE atomics woven through every mutex, a baseline build shows LDXR/STXR loops with LSE only inside runtime-dispatched thunks (glibc's outline atomics).

$ ./armlint -i libc.so.6
ISA census: 275103 instructions, 1 undecodable words skipped
  Armv8.0 baseline: 275043
  mandatory from Armv8.1: LSE (21)
  mandatory from Armv8.3: none
  ...
  optional features: none
  branch protection (hint space): BTI (21), PAC (18)
  highest mandatory-from level: Armv8.1

Three groups never raise the ladder, each for a soundness reason of its own. Features that never become mandatory in the v8 line (the crypto extensions, FP16, SVE, MTE) are listed as optional: any of them can be bolted onto an old target with a single +feature flag, so their presence says nothing about -march. The hint-space branch-protection forms (PACIASP/AUTIASP/BTI/XPACLRI) execute as NOPs on cores without the extension -- -mbranch-protection=standard emits them precisely so the binary stays v8.0-compatible -- so they get their own line and no version claim; only the register-form PAC instructions (PACIA, RETAA, BRAA, ...) evidence a real Armv8.3 target. And an undecodable word is either data in text or an extension this Capstone build cannot decode, so the skipped count bounds what the census could have missed.

The census reports presence, not requirement: dispatched fast paths count even though the binary runs without them. Treat small exotic tallies in a binary with many skipped words with suspicion -- string pools embedded in text sometimes decode as valid SVE or atomics -- and use -v, which prints up to four sample addresses per feature, to check a surprising tally in a disassembler before believing it.

When the binary carries function boundaries (Mach-O LC_FUNCTION_STARTS/nlist, ELF .symtab/.dynsym -- the same sources that symbolize -v findings), the census adds per-function pac-ret coverage:

  pac-ret coverage: 866 of 1113 functions sign the return address

A function counts as signed when its span contains PACIASP/PACIBSP or the register-form pacia/pacib x30, sp. The ratio is a fingerprint, not a target: leaf functions never spill the return address and legitimately never sign, so Apple's fully signed arm64e binaries sit around 78-85% (zsh 866/1113, ssh 912/1073, sshd 360/430) with the remainder leaves -- the -a pac audit separately confirms zero unsigned spills there, which is the claim that matters. The signature of never opting in looks entirely different: Homebrew's plain-arm64 libcapstone reads 0 of 1441. The line is omitted when no boundary information exists (Go binaries carry neither structure).

Immediate-misfit audit (-a imm)

armlint -a imm adds an audit to the lint scan that is the complement of the mov folds. Those report a movz/movn/movk chain whose consumer could have taken the value as an immediate; the audit reports the chain whose consumer has an immediate form the value misses: movz x16, #0x100, lsl #16 ; cmp x0, x16, where 0x1000000 is one step past the largest cmp immediate. Four consumer classes are checked, each against its own encoding: add/sub/adds/subs and cmp/cmn (a 12-bit immediate, optionally shifted by 12, in either sign -- cmp x0, #-5 is cmn x0, #5), the logical ops and/orr/eor/ands/tst (a bitmask immediate; the complement for bic/orn/eon/bics), ccmp/ccmn (a 5-bit immediate, ccmn covering -1..-32), and a register-offset ldr/str indexed by the constant (the scaled 12-bit or the unscaled signed 9-bit offset). The consumer must follow the chain directly, the operand rules are the folds' own (the constant in an immediate-capable slot, the other operand neither the zero register nor the constant's), zero and all-ones are skipped, and the chain and the consumer may differ in width where the value survives the change (a W consumer reads the low half of an X chain, a W chain zero-extends into an X consumer).

Nothing at the site can be rewritten -- the chain is already the cheapest spelling of that value -- so the finding is informational, like the PAC audit's, and the review item is the constant itself. The exception is a plain add or sub of a constant below 2^24, which needs no register at all: when the constant's register dies, the two-immediate fold reports the site too, and the audit still lists the value, since renumbering it saves one more instruction. A value the code's author chose (an object-size limit, a sentinel, the bit assignment of a flags field) that lands one step outside the encoding costs a materialization and a register operand at every use, and choosing it one step differently turns every site into the immediate form. The audit therefore needs no liveness proof -- a constant hoisted into a live register amortizes its materialization, but each use still pays for the register operand the immediate form would not need -- and each finding names what would have encoded. The summary tallies the audit by value, most frequent first, so the constants worth renumbering sort to the top:

$ ./armlint -a imm libxul.so
Optimization opportunities by type:
   80391  MOV + ADD/SUB/CMP: constant has no add/sub immediate form (imm audit)
   ...
   10279  MOV + AND/ORR/EOR/TST: constant is not a bitmask immediate (imm audit)
   ...
    6364  MOV + CCMP/CCMN: constant is outside imm5 (imm audit)
   ...
    1228  MOV + register-offset LDR/STR: constant has no immediate-offset form (imm audit)
   ...

Immediate misfits by value (-a imm), within one step of an immediate form:
    2707  bitmask imm    w  #0x50         nearest bitmask immediates 0x10 and 0x40 (1 bit away)
    2501  add/sub imm12  w  #0x270f       nearest add/sub immediates 0x2000 and 0x3000
    1489  bitmask imm    x  #0xfffb000000000000 nearest bitmask immediate 0xffff000000000000 (1 bit away)
    1000  add/sub imm12  x  #0x15f90      nearest add/sub immediates 0x15000 and 0x16000
    ...
     159  ldr/str offset b  #0x10d9       immediate offsets cover -256..255 unscaled, 0..4095 x size scaled
    ...
  (7309 more distinct values)

Immediate misfits by value (-a imm), beyond one step (informational):
   11978  add/sub imm12  x  #0x8000000000000000 largest add/sub immediate is 0xfff000
    5738  add/sub imm12  x  #0x7ffffffffffffffd largest add/sub immediate is 0xfff000
    ...
  (7906 more distinct values)

135936 optimization opportunities in 28583984 instructions

Each row is a count of sites, the consumer class, a width letter (the consumer's register width, w or x; for a load or store its access size, b/h/w/x), the constant (a load/store misfit is tallied as its byte offset), and a hint naming what would have encoded: the two nearest 4 KiB multiples, the bitmask immediates fewest bit flips away, the imm5 or offset range. A bitmask hint sometimes adds a widened mask -- 0xdf, the byte with its ASCII case bit clear, is not a bitmask immediate but 0xffffffdf is, and the two agree on any operand whose bits above the mask are known clear -- which is a compiler's fix (known bits) rather than a renumbering. Each table prints its top 40 rows and counts the rest.

Two tables, because a real corpus is dominated by constants no renumbering reaches. The first holds the constants within one step of an encodable neighbour: an add/sub magnitude up to 0x1000000, so a 4 KiB multiple lies within 4 KiB of it; a bitmask at most two bit flips away; a ccmp magnitude up to 63; a load/store byte offset within twice the scaled range of its access size, or the 256 bytes below the unscaled range. That table is the renumbering worklist, and its rows have owners: in the Firefox libxul.so above, 0x50 under tst is JS::shadow::Zone::GCState with Sweep = 4 and Compact = 6, a two-bit mask one flip from contiguous; 0x270f is nsAtom's kAtomGCThreshold = 10000, tested as ++count >= 10000 and so compared against 9999, where 8193 would encode; the byte loads at 0x10d9 are Document bit-field bytes just past the 4 KiB byte-load range. The second table holds everything beyond reach, whose whole magnitude is the point: INT64_MIN under cmp (the niche rustc gives the dataless variants of a Vec- or String-carrying enum, plus Gecko's TimeDuration sentinel), the JS::Value tags, the words of the interface IDs every QueryInterface compares.

To find the sites of one value, -v prints each finding as MOV + ADD/SUB/CMP: ... at offset: 0x... <function+0x...>: -> #0x270f; nearest ..., so a grep for the constant lists them with their containing functions, which point at the definition to change. The findings count toward the total and the non-zero exit status like any other, so a test suite that gates on a clean run should not pass -a imm; the worklist is for reading. analyses.md walks the rules and two corpora in full: this libxul.so run and V8's JetStream 3 JIT output (-m v8 -a imm on a tools/v8dump2elf.py dump), where the table opened with a 16 MiB reservation bound one step past 0xfff000 at 92,109 sites, since fixed in V8 by comparing against the largest encodable bound.

Duplicate code (-d)

armlint -d adds a report on which functions are copies of one another, and on which findings are the same advice counted again in each copy. A function is a span between the boundaries -v uses to name functions: Mach-O nlist and LC_FUNCTION_STARTS, or the ELF .symtab (.dynsym as a fallback), where a sized symbol ends at its size. An executable section with no boundary inside it counts as one function, which is how a JIT dump's code blobs arrive. Trailing NOP and zero padding is not part of a function.

Each function is keyed twice:

  • The identical key is the code as it executes. A branch, adr, adrp or literal load that reaches outside the function counts by the absolute address it resolves to. One that stays inside keeps its encoding, which is relative to the function. Two functions with the same identical key behave the same wherever they are loaded, so a linker's identical-code folding could keep just one.

  • The up to constants key also masks the fields that tell instances of one piece of code apart:

    • movz/movn/movk immediates;
    • the words the function's own literal loads read (and V8 constant pools, under -m v8);
    • every target outside the function;
    • the low 12 bits an add or a load/store adds to an adrp page.

    This merges a template's or a generic's instances that differ only in what they call and address, and a JIT's copies of one stub that differ only in the pointers they embed. It also merges two stubs that differ only in a constant they guard. That is the intended reading of "duplicate code", but keep it in mind when reading the numbers.

$ ./armlint -d librustc_driver.dylib
Optimization opportunities by type:
   ...

Duplicate code (-d): 174219 functions, 103924556 bytes
  identical: 36353 copies of 5900 functions, 4591292 bytes (4.4%)
  up to constants: 86331 copies of 11651 functions, 24890780 bytes (24.0%)
  how often each distinct function appears, up to constants: once 76237, 2-10 times 10101, 11-100 times 1484, more 66
  findings: 18673 of 96161 (19.4%) repeat an earlier copy's at the same offset
      7510 of  27595  ADD/SUB immediate chain foldable to one
      3086 of  16891  register already holds the recomputed value
      2028 of   6156  branch to the next instruction is a no-op
   ...

96161 optimization opportunities in 25974168 instructions

In scan order, the first function with a given key is the original and the later ones are its copies. Percentages are of the bytes in functions. The rustc above keeps 4.4% of its code in identical copies that the linker did not fold; one example is 350 copies of a single 168-byte LLVM DenseMap lookup. A further fifth of its code is generic and template instances that differ only in what they call and address.

A finding repeats when an earlier copy of its function (up to constants) has a finding of the same type at the same offset: the same advice to whatever emitted the code, counted again. The by-type table, the total and the exit status still count every finding. The findings lines say how much of that count is repetition, overall and by type. -v marks each repeated finding by appending [repeat] to its header line. tools/v8_analyze_findings.py and tools/jsc_analyze_findings.py read that marker and add a repeats row under their per-tier tables. -v also lists the most duplicated functions. Here it is on SpiderMonkey's Octane code (a tools/smdump2elf.py dump):

$ ./armlint -d -v sm-octane.elf
...
  most duplicated up to constants, by bytes past the original:
      copies variants    bytes findings distinct  original
          86       86     3568     1118       13  <BL_f1e37185c0>
          86       86     3424     1290       15  <BL_f1e371c3b0>
         118      118      240      472        4  <ION_f1e3379700>
...

The columns:

  • variants: how many distinct identical keys the copies have. 1 means every copy is the same function byte for byte.
  • findings: the findings in all the copies together.
  • distinct: how many of those findings are not repeats.

The first row above is 86 Baseline blocks of 3,568 bytes each. They are the same code up to constants, and no two are byte-identical. Their 1,118 findings are 13 pieces of advice.

corpus functions identical up to constants findings repeated
librustc_driver 174,219 4.4% 24.0% 19.4%
HotSpot, javac (every tier) 24,739 4.6% 10.9% 24.9%
SpiderMonkey, Octane 5,579 0.7% 9.1% 14.6%
V8, Octane 2,534 0.9% 1.7% 0.6%

Most of HotSpot's repeats (51,417 of 66,222) are in C1's per-method stubs, every one of whose 86,790 findings is a "suboptimal MOVZ/MOVK sequence". C2's own code is barely repeated: 382 of its 21,371 findings.

Limitations:

  • Keys are 64-bit hashes plus the function's length, so a false merge is possible but improbable.
  • In an unlinked object file every relocated field reads as zero, so functions that differ only in a relocated target key alike.
  • A stripped ELF with only .dynsym has boundaries for its exports alone, so -d sees only those functions.
  • With -s all, every slice feeds one table. A function that both compilations contain counts as a copy, and its findings as repeats.
  • The adrp page tracking is linear, not flow-sensitive.
  • Pointers stored as data inside a blob, such as a JIT's absolute dispatch table, are not masked.

Constant chains (-c)

armlint -c adds a report on the constants the code builds with movz/movn + movk chains. A chain is a run of two or more of them into one register at one width: the runs the "suboptimal MOVZ/MOVK sequence" check judges. The -a imm audit reports a constant that just misses its consumer's immediate form. This report is its complement: it counts what every constant too wide for one instruction costs. Constants are tallied by width and value, so w #0x1deb8 and x #0x1deb8 are two values, and ranked by the instructions spent on them.

Within a function, a chain that builds a value an earlier chain of the same function already built is a rebuild. The functions are the ones -d uses. A rebuild spends the instructions again instead of keeping the value in a register. The usual causes are a call that clobbered the register and a loop invariant that was not hoisted. Not every rebuild is a miss: the two arms of an if/else both need their build, and a rebuild after a call can be cheaper than a spill. The rebuild counts bound what keeping the values could save; they do not measure it.

$ ./armlint -c librustc_driver.dylib
Optimization opportunities by type:
   ...

Constant chains (-c): 25981139 words, 174219 functions
  every chain: 73533 building 17385 distinct values, 208685 instructions (0.80%)
  rebuilds within a function: 22192 in 4922 functions, 59481 instructions (0.23%; 28.5% of chain instructions)
  how often a function builds a value it rebuilds: twice 5825, 3 times 1381, 4 times 676, 5 or more times 1220
  most instructions spent building one value:
    instructions   chains  value
           19792     4948  x #0xf1357aea2e62a9c5
            6180     1545  x #0xbf58476d1ce4e5b9
            3774     1887  w #0x7a3e8
            3302     1651  w #0x1deb8
   ...
    (17365 more distinct values)
  most instructions spent rebuilding one value within a function:
    instructions rebuilds functions  value
            6616     1654       826  x #0xf1357aea2e62a9c5
            1952      976       246  w #0x1deb8
   ...
    (3323 more distinct values)

96161 optimization opportunities in 25974168 instructions

Every word of the executable sections is scanned, except V8 constant pools under -m v8, and the percentages are of those words. Each ranking lists its top 20 values. Here x #0xf1357aea2e62a9c5 is the multiplier of rustc-hash's FxHasher, built with four instructions at every inlined hash, and w #0x1deb8 is a field offset into rustc's global context. -v adds where to look. It gives the addresses of a value's first three chains; for a rebuilt value, it gives the function that builds it most often and its first builds there:

    instructions rebuilds functions  value
            6616     1654       826  x #0xf1357aea2e62a9c5  built 120 times in <__RNvNtCshPlmC27tfnj_16rustc_query_impl9execution25collect_active_query_jobs>: 0x2603918, 0x2603f44, 0x260452c (+117 more)
corpus chain instructions distinct values rebuilt
librustc_driver 0.80% 17,385 28.5%
clang-24 1.10% 22,502 33.6%
Firefox libxul.so 0.61% 11,753 not counted
HotSpot, javac (every tier) 17.8% 44,350 63.9%
SpiderMonkey, Octane 10.2% 7,465 58.2%
V8, Octane (-m v8) 3.9% 485 73.6%

The chain instructions are a share of the words scanned, and the rebuilt ones a share of the chain instructions.

  • clang's costliest value is x #0xbf58476d1ce4e5b9, splitmix64's multiplier. LLVM's DenseMap hashes every pointer key by multiplying it by that constant (densemap::detail::mix). The value takes 45,804 instructions in 11,451 chains, 19% of clang's chain instructions.
  • libxul.so is stripped. Its .dynsym bounds 5 functions in the text, so its rebuilds go uncounted. Its totals are led by SpiderMonkey's UndefinedValue() (x #0xfff9800000000000, 7,661 chains) and by nsresult codes: NS_ERROR_ILLEGAL_VALUE (w #0x80070057, 3,616) and NS_ERROR_FAILURE (w #0x80004005, 3,337).
  • HotSpot's costliest value is x #0x0: 112,251 three-instruction chains, 26% of its chain instructions, all but 8 of them in the stubs C1 and C2 append to each method. The stub isb ; mov x12, #0 ; movk ; movk ; mov x8, #0 ; movk ; movk ; br x8 calls the interpreter, and its method and entry address are placeholders, patched when the call is resolved. A method's stubs are one function, so the same placeholders also lead the rebuilds.
  • SpiderMonkey's costliest value is the JS::Value Int32 tag, x #0xfff8800000000000 (12,865 chains). UndefinedValue() (8,901), Int32Value(2) and Int32Value(1) join it in the top six, between runtime pointers. One Baseline blob rebuilds the Int32 tag 2,773 times.
  • V8's costliest value, x #0x17cb0007c000, is the address of a typed array's backing store. Maglev and TurboFan embed it in the code of Octane's Mandreel benchmark, the Bullet physics engine compiled to JavaScript: 3,992 chains in 146 functions, most of them followed directly by an indexed access such as str w1, [x3, x0, lsl #2]. One Maglev function rebuilds it 222 times.

Limitations:

  • Only move-wide chains count. A constant loaded from a literal pool, built by an orr + movk hybrid, or encoded in one instruction is not counted.
  • A W chain zero-extends into its X register, so w #0x1deb8 and x #0x1deb8 leave the same register contents. They still count as two values, and a build of one does not make the other a rebuild.
  • A rebuild is a build of the same value earlier in the function, in address order. Neither dominance nor liveness is checked.
  • A JIT patches many of its constants at runtime. The census cannot tell a patch site's placeholder from a constant.
  • The functions are -d's. A stripped ELF with only .dynsym has functions for its exports alone. A binary with no boundaries at all is one function per section, which makes every repeated value in a section a rebuild.
  • With -s all, every slice feeds one table.

Mining tools

tools/ holds the research utilities that feed armlint's check backlog -- and one that checks the checks -- built separately with make tools:

  • tools/pairscan counts adjacent-instruction pairs by normalized shape (registers collapsed to classes, immediates to #0/#i) across the executable sections of ELF and Mach-O binaries, surfacing frequent patterns worth a new check. -e SUBSTR prints example sites for shapes matching a substring.
  • tools/defuse profiles block-local def-to-use distances (how far a value's sole consumer sits from its producer) and multi-instruction redundancies no pair statistic can see: dead definitions, redundant reloads of the same address, re-materialized constants, and zero compares of a value whose producer could have set the flags.
  • tools/shapescan.py counts a fixed list of specific candidates -- the rows TODO.md tracks -- with their real operand, range and encodability conditions applied, which is what separates a population from a pair count. adrp + add is the standing example: 753,648 adjacent dependent pairs across the corpus, of which 43,434 have a target inside ADR's reach. Run tools/shapescan.py --selftest first: it assembles every reference instance with clang and checks each mask in both directions -- that it matches no instruction belonging to another mask, and that it matches every spelling of its own listed in ALSO. The second half is the one that matters most, because a mask too narrow to see the stur, the ldp q, or the 64-bit form of its shape reports a small number rather than a wrong one, and nothing looks broken. -e SHAPE prints example sites. Needs numpy.
  • tools/addpairscan.py sizes the ADD-immediate + LDP/STP family, which one shape mask cannot split honestly: the half whose combined offset fits the pair's signed, pre-scaled imm7 folds two instructions into one, while the half that overflows it can only be split into two singles at no size saving. 8,775 against 17,565 across the corpus. Carries its own --selftest; needs numpy.
  • tools/rwfuzz checks a shipped check's advice by running it. It generates random AArch64 programs, applies each rewrite armlint suggests for one check to a copy, and runs original and rewrite natively from the same random registers, flags and memory: any difference is a soundness bug. A control arm applies the rewrite where armlint refused to, so a harness that stopped seeing differences would show it. Modes cover the redundant zero-extension, shift + mask, CMP + CCMP chain, value-numbering, dead-write and CMP-to-CMN checks (tools/rwfuzz -n 100000 zext); a new rewrite gets a mode of its own. Needs an AArch64 host, since the programs run natively, and links libarmlint.a.
  • tools/tblgen_audit.py checks armlint's flag and register model against LLVM's own instruction definitions. llvm-tblgen --dump-json on AArch64.td lists every A64 instruction with its encoding and its implicit NZCV and register defs and uses; the tool builds one word per record, has tools/tblgen_probe print Capstone's access lists next to armlint's classifiers for it, and then runs armlint itself over mov ; W ; cbz, mov ; W ; mov and cmp ; b.ne ; W ; b.ne probes, where a finding means a write or a read of W's went unmodelled. That second step is what proves a blind spot, since armlint already corrects much of what Capstone omits. It found FJCVTZS, SUBPS, the SVE compares and the MOPS source pointer; --masks derives the encoding masks for a new family. Needs a built llvm-tblgen and clang.

The workflow that produced several of the current checks: compile a representative corpus, run pairscan to rank pair shapes, classify the top shapes as by-design or foldable, then use defuse to decide whether a candidate needs adjacency only or a liveness window, and shapescan.py to size the survivor with its real conditions applied before writing any code. pairscan and defuse lean on Capstone's register-access model, which mis-reports the compare aliases (CMP/CMN/TST mark their first operand as a write); defuse corrects for this, and armlint's own checks decode the raw encodings precisely to avoid that class of problem. shapescan.py decodes raw encodings for the same reason, and self-tests them because that is where its own bugs live: every wrong figure it has produced came from a mask one bit too loose or a destination modelled as write-only when the instruction merges into it.

References

License

Copyright (C) 2026 Andrew Gaul

Licensed under the Apache License, Version 2.0

About

Examine AArch64 machine code to find suboptimal instruction sequences

Resources

Stars

9 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages