armlint examines AArch64 machine code to find suboptimal instruction
sequences. For example, building the constant 0x66666666 as
movz w0, #0x6666
movk w0, #0x6666, lsl #16is two instructions where one would do, because 0x66666666 is encodable
as an AArch64 logical (bitmask) immediate:
mov w0, #0x66666666 ; orr w0, wzr, #0x66666666armlint helps compiler writers and assembly authors generate tighter code, and documents corners of the A64 instruction set.
armlint is a peephole analyzer. It decodes each 32-bit A64 instruction
directly from the binary and matches it by mask and value, resolving
aliases (for example MUL is MADD with a zero accumulator) so that
both spellings of a pattern are caught. It then looks for a short window
of adjacent instructions that a shorter or cheaper encoding can replace.
The overriding rule is soundness: armlint emits a finding only when the
rewrite provably preserves the architectural result. For a tool that
suggests code changes, a false positive is the worst failure, so it errs
toward false negatives -- a missed opportunity is cheaper than a wrong
one. Each check documents the exact conditions under which its rewrite is
equivalent; the constraints below are the ones they share, and
analyses.md's appendix
collects the near-miss folds that are deliberately never matched (FP
contraction, fcsel -> fmax, the SDIV remainder, and friends) with
the argument against each.
- Strict adjacency for matching; bounded lookahead for proof. The instructions of a matched pattern must be consecutive -- an unrelated instruction between a producer and its consumer suppresses the finding -- and armlint does not reorder code or match through intervening instructions. Emission, though, is not confined to the pattern window: as the following bullets describe, many findings are held back while a bounded forward scan (16 instructions) walks the fall-through path past the consumer to prove a register or the flags dead, so an instruction well after the pair can also suppress a finding. The scan only gates emission -- it never widens a match, so every reported rewrite still replaces only the adjacent instructions shown.
- Liveness is proved structurally, or by a bounded forward scan. A
producer-into-consumer fold fires when the consumer overwrites the
producer's destination register, proving the intermediate value is
dead. Folds whose saving is a deleted write with no such overwrite
defer instead: a bounded forward scan of the fall-through path must
see the register overwritten before any read or control transfer.
This is how the address folds admit stores and loads into a fresh
register --
add x8, sp, #32 ; str x0, [x8]folds tostr x0, [sp, #0x20]only oncex8provably dies -- and how the producer folds (shift, funnel, extend,MUL/SMULL,NEG,MVN) admit consumers that write a register other than the producer's:lsl w8, w1, #3 ; add w9, w2, w8folds toadd w9, w2, w1, lsl #3under the same proof. The single-bit and CSET branch folds additionally require the folded branch's taken edge to land inside that proven-clean span -- a general-purpose register, unlike NZCV, is routinely live into a branch target, so no block-locality assumption is made for it. - MOV-chain folds verify the constant dies. Folds that absorb a
materialized constant --
MUL/MNEG/UDIVby a constant,MOV+ADD/AND/ORR/EOR/CCMP/FMOV,MOV #0, the register-offset and MOVI-zeroing folds -- report only once the consumer's own overwrite or the forward scan proves the constant register dead. The consumer rewrite itself stays valid regardless. - Flag liveness uses a bounded forward scan. The branch- and
flag-folding checks drop a
CMP/TSTonly after a bounded scan of the fall-through path confirms that no later instruction reads N/C/V before they are overwritten. Every NZCV reader is recognized -- the integer conditionals (B.cond,CSEL/CSINC/...,CCMP/CCMN,ADC/SBC) and the floating-point ones (FCSEL,FCCMP/FCCMPE) -- and any branch off the path whose destination the scan cannot see -- an unconditionalB, or a conditionalCBZ/CBNZ/TBZ/TBNZ(which do not themselves touch NZCV but whose taken target may still observe it) -- ends it conservatively. The scan does not follow the folded branch's own taken edge, so most of these folds assume N/C/V is dead at every branch target. That holds for most compiled code, where the flags are defined within a basic block, but not for hand-written assembly that deliberately keeps a flag live into a branch target -- including the stubs a JIT inlines into its output. The folds that meet such code prove the taken edge with a second scan starting at the target: the compare-and-branch fold (-m cmpbr), whose producer is a general two-register compare that clang's three-way comparator really does read again at the branch target; theCCMPchain andCMP-to-CMNfolds; and the zero-compareCBZ/CBNZand sign-bitTBZ/TBNZfolds, since HotSpot's lock fast paths return their result to the branch target in the flags. Unlike the fall-through scan, the one at the target follows aBto its destination and aCBZ/CBNZ/TBZ/TBNZdown both edges, within fixed bounds.
Findings are opportunities, not guaranteed speedups: some -- the pre- and post-indexed addressing folds -- are code-size and front-end wins that are backend-neutral. Each check's notes say what its rewrite actually saves.
Each row links to its full description -- mechanics, soundness, and what the rewrite saves -- in analyses.md. Candidate checks not yet implemented live in TODO.md.
| Pattern | Rewrite |
|---|---|
movz/movn + movk (over-long constant) |
bitmask-immediate mov, or minimal movz/movn + movk chain |
lsl/lsr/asr/ror + add/sub/and/orr/eor |
add Rd, Rn, Rm, <shift> #n |
lsl/lsr + shifted orr/eor/add (funnel/rotate) |
extr Rd, Rhi, Rlo, #lsb (or ror when both halves match) |
sxtw/uxtb/sxtb + add/sub |
add Rd, Rn, Wm, sxtw |
cmp/cmn/tst zero-test + b.eq/b.ne (b.hi/b.ls after cmp) |
cbz/cbnz |
cmp/cmn/tst zero-test + b.lt/b.ge/b.mi/b.pl |
tbnz/tbz Rn, #(msb) |
tst #(1<<k) + b.eq/b.ne |
tbz/tbnz Rn, #k |
tst #(1<<k) + cset/csetm |
ubfx/sbfx Rd, Rn, #k, #1 |
single-bit and/ubfx/lsr #31 + cbz/cbnz |
tbz/tbnz Rs, #k |
cset + cbz/cbnz |
b.<cond> / b.<inverse cond> |
cset + eor #1 |
cset <inverse cond> |
cset + neg |
csetm |
br x30 |
ret (engages the return-address predictor) |
| branch to the next instruction | delete (both outcomes fall through; bl excluded) |
lsl + lsr/asr |
ubfx/sbfx/ubfiz/sbfiz |
lsr/asr + and #mask |
ubfx |
and #mask + lsr |
ubfx |
and #mask/uxtb/uxth/uxtw/mov + lsl (or lsr + lsl) |
ubfiz (or clearing and) |
zeroing producer + uxtb/uxth/uxtw/and/mov Wd, Wd, adjacent or not |
drop the zero-extension |
mov xd, xd |
remove (architectural no-op) |
| instruction recomputing a value its register or NZCV already holds, across a gap | delete (an ADR/ADRP of a held page is reported apart: a relink fix) |
| register write overwritten before any read, across a gap | delete the write |
| conditional branch the path to it already decides | delete it (never taken) or b (always taken) |
| ALU operation whose inputs are all known values | mov Rd, #result, or delete when Rd already holds it |
and #mask/ubfx #0 of a value its known bits already keep inside the mask |
delete in place, mov Rd, Rn out of place |
sign-extending producer + sxtb/sxth/sxtw |
drop the sign-extension |
and/uxt*/sxt*/mov Wd, Wm + a narrower and/ands/tst/uxt*/sxt* |
the second reading the first's source, or one and of the intersected masks |
and/orr/eor/sub/bic/orn/eon with Rs, Rs |
mov / zero / all-ones |
ldr+ldr / str+str (consecutive) |
ldp/stp (integer or FP/SIMD, and ldpsw); notes where a slot is in flight, since Apple cores do not forward a store to a pair load or a pair store to a load |
str wzr+str wzr (consecutive zero stores) |
str/stur xzr |
stp wzr, wzr (W-form) |
str/stur xzr |
movi #0 + vector cmeq/cmge/cmgt (or FP fcm*) |
cmeq/cmge/cmgt/cmle/cmlt Vd, X, #0 (drop the movi) |
and + and/ubfiz + orr (clear/isolate/merge) |
bfxil/bfi |
and #lowmask + orr Rm, lsl #k (or high mask + lsr; field reaches the top) |
bfi/bfxil |
csel Rd, Rn, Rn, cond |
mov Rd, Rn |
fcsel Vd, Vn, Vn, cond |
fmov Vd, Vn |
add/sub Rd, Rn, #0 |
mov Rd, Rn, or remove |
add/sub #a + add/sub #b (same register) |
one add/sub carrying the sum (or mov) |
adds/subs/ands + cmp #0 + b.eq/b.ne |
drop the redundant cmp/tst |
add/sub/and/bic + cmp #0 + b.eq/b.ne |
adds/subs/ands/bics (drop the cmp/tst) |
sub + cmp / add + cmn of the same operands (either order) |
subs/adds (flag-exact; drop the compare) |
subs/adds + cmp/cmn of its own operands |
drop the compare (flags already set) |
cmp/cmn/tst/ccmp whose flags are overwritten unread |
delete (the compare writes no register) |
cmp #0 + cset/csetm of lt/mi |
lsr/asr Rd, Rn, #(msb) (the sign bit; drop the compare) |
mov #2^N + mul |
lsl, or add Rd, Ra, Ra, lsl #N |
mov #C + mneg |
neg, or shifted neg/sub |
mov #2^N + madd/msub |
add/sub Rd, Ra, Rn, lsl #N |
mov #2^N + udiv |
lsr |
mov #2^N + udiv + msub (remainder) |
and Rd, Rn, #(2^N-1) |
mov #C + add/sub |
add/sub Rd, Rn, #C (sign-crossed add↔sub, cmp↔cmn for #-C) |
movz/movn + movk of a constant below 2^24 + add/sub |
add/sub Rd, Rn, #hi, lsl #12 + add/sub Rd, Rd, #lo (add↔sub for a negative constant) |
mov #C + and/orr/eor/ands or bic/orn/eon/bics |
and/orr/eor/ands Rd, Rn, #C (#~C for the inverting forms) |
mov #INT_MIN or #INT_MAX + cmp + b.eq/b.ne |
cmp xzr, Xn or cmn Xn, #1 + b.vs/b.vc |
mov #C + ccmp/ccmn |
ccmp/ccmn Rn, #C, #nzcv, cond (sign-crossed for #-C) |
cmp/cmn + ccmp/ccmn on one register + b.cond/csel, one compare deciding it |
that cmp/cmn (or cbz/cbnz) with the reader's condition adjusted |
mov #(2^32-k) + cmp Xn of a zero-extended value, read only through Z and C |
cmn Wn, #k |
mov #1 + csel |
csinc Rd, Rn, wzr, cc (cset when the other operand is ZR) |
mov #-1 + csel |
csinv Rd, Rn, wzr, cc (csetm when the other operand is ZR) |
mov #C + lsl/lsr/asr/ror (register amount) |
immediate-form shift, amount C mod 32/64 |
mov #bits/#C + fmov/scvtf/ucvtf from GPR |
fmov Sd/Dd, #imm |
fmov/scvtf/dup of wzr/xzr (or mov #0 + transfer) |
movi dN, #0 / movi vN.T, #0 |
sxtw/mov w, w + scvtf/ucvtf Xn |
scvtf/ucvtf of Wn |
ldr w8 + scvtf/ucvtf from GPR |
ldr s0 + FP-side convert (no cross-file transfer) |
ldr/ldrsw (literal) of an encodable constant |
mov #imm / fmov #imm8 / movi+mvni for Q (no memory access) |
adr + ldr [x8]/br x16 |
ldr Rt, <literal> / b L |
umov Wd, Vn.s[0] / Xd, Vn.d[0] |
fmov Wd, Sn / Xd, Dn (cheaper port; .h[0] under -m fp16) |
and Xd, Xn, #0xffffffff / ubfx Xd, Xn, #0, #32 |
mov Wd, Wn (a rename on Neoverse rather than an ALU op) |
cmp + csel (max/min shape) |
smax/smin/umax/umin (-m cssc) |
cmp #0 + cneg |
abs (-m cssc) |
rbit + clz |
ctz (-m cssc) |
| NEON popcount round trip | cnt Xd, Xn (-m cssc) |
cmp + b.<cond> |
cb<cond> Rn, Rm/#imm, label (-m cmpbr) |
eor + eor (16B vectors) |
eor3 Vd, Vn, Vm, Va (-m sha3) |
bic + eor (16B vectors) |
bcax Vd, Vn, Vm, Va (-m sha3) |
autiasp/autibsp + ret |
retaa/retab (-m pauth; auto-armed on arm64e) |
unsigned LR spill; raw br/blr; zero-discriminator braaz/blraaz |
audit-only review items (-a pac; auto-armed on arm64e) |
mov chain + cmp/tst/ccmp/ldr whose immediate form the value misses |
audit-only review items, tallied by value, the constants within one step of an immediate first (-a imm) |
ldxr/stxr fetch-op retry loop |
ldadd/ldset/ldeor/ldclr (+ mvn/neg/mov pre-op) (-m lse) |
ldxr/stxr exchange retry loop |
swp (-m lse) |
ldxr + cmp + b.ne + stxr CAS retry loop |
mov + cas + cmp (-m lse) |
mov #C + orr Xd, x28, Xc (V8 cage base) |
add Xd, x28, #C (-m v8) |
fmul + in-place fneg |
fnmul (bit-exact in every rounding mode) |
mov #0 + str/add/and/csel/ccmp use |
use wzr/xzr |
mov #C + ldr/str [xn, xc] |
ldr/str [xn, #C] (or ldur/stur) |
mul + add/sub |
madd/msub (or mneg) |
smull/umull + add/sub |
smaddl/umaddl/smsubl/umsubl |
neg + add/sub |
sub/add |
neg + csel |
csneg (inverted cond for the then slot) |
mvn + csel |
csinv (inverted cond for the then slot) |
add #1 + csel |
csinc (inverted cond for the then slot) |
mvn + and/orr/eor/ands |
bic/orn/eon/bics |
add + ldr/str [xt] |
ldr/str [xn, xm{, lsl #s}] |
sxtw + ldr/str [xn, xt] |
ldr/str [xn, ws, sxtw {#s}] |
ldrb/ldrh/ldr (or ldrsb/ldrsh Wt) + sxtb/sxth/sxtw |
ldrsb/ldrsh/ldrsw (Xt for the re-widened sign loads) |
add #a + ldr/str [xt] (incl. mov xt, sp) |
ldr/str [xn, #a] / ldr [sp] |
ldr/ldp [xn] + add/sub xn |
ldr [xn], #±imm / ldp [sp], #imm (post-index) |
add/sub xn + ldr/stp [xn] |
ldr [xn, #±imm]! / stp [sp, #-imm]! (pre-index) |
armlint depends on Capstone and uses
pkg-config to locate it. On macOS:
brew install capstoneOn Debian/Ubuntu:
apt install libcapstone-dev pkg-configBuild:
git clone https://github.com/gaul/armlint.git armlint
cd armlint
make allTwo test suites are available. make test runs the unit tests against
fabricated byte sequences, exercising the check registry directly.
make integration-test runs the snapshot suite under fixtures/:
each .s is assembled with clang -arch arm64 and armlint's
output is diffed against a checked-in .expected file. The
integration suite covers the Mach-O parser and the report formatting,
which the unit tests bypass. It needs a clang that can assemble
AArch64 and fails without one rather than reporting a pass it did not
earn; off an arm64 host it names the target explicitly, so an x86-64
Linux box runs the suite too. After an intentional output change,
regenerate the
snapshots with make integration-test-regen and review the diff
before committing.
Setting ARMLINT_LIVENESS_SWEEP=1 extends the unit tests' liveness
cross-checks against Capstone to the entire 2^32 encoding space --
minutes of single-threaded CPU time, so CI runs it on pushes rather
than PRs. Both classifiers are covered off one decode per word: the
NZCV one (classify_liveness) and the register one
(classify_reg_liveness), whose property is that a register Capstone
reports as read must never be classified as dead or as no-effect.
Three Capstone over-reports are documented and skipped; everything
else is a defect on one side or the other, and the register half found
four on its first run. ARMLINT_LIVENESS_SWEEP_THREADS divides the
sweep across that many worker threads:
ARMLINT_LIVENESS_SWEEP=1 ARMLINT_LIVENESS_SWEEP_THREADS="$(sysctl -n hw.ncpu)" \
make test(nproc on Linux.)
Capstone 6 (in alpha as of 2026) rewrote the AArch64 module from
LLVM. armlint compiles against it unchanged via Capstone's
compatibility header, and CI tracks the 6.0.0-Alpha11 tag. Before
Alpha11 it tracked a commit instead: Alpha10 still leaves
insn->alias_id stale on non-alias instructions (the sweep's oracle
then reports a phantom x30 read after every ret) and still
mis-reports the SYSL and integer-STLUR operands, and Alpha11 is the
first tag with all three fixed. To reproduce locally:
git clone --depth 1 --branch 6.0.0-Alpha11 \
https://github.com/capstone-engine/capstone.git capstone6
cmake -B capstone6/build -S capstone6 -DCMAKE_BUILD_TYPE=Release
cmake --build capstone6/build -j8
make clean # never mix objects built against different Capstone ABIs
make CAPSTONE_CFLAGS="-I$PWD/capstone6/include -DCAPSTONE_AARCH64_COMPAT_HEADER" \
CAPSTONE_LIBS="$PWD/capstone6/build/libcapstone.a" allThe overrides must be make arguments (not environment variables) to
beat the Makefile's pkg-config defaults. The unit tests and the
ARMLINT_LIVENESS_SWEEP=1 sweep pass under both major versions, and
running the sweep under both is the point rather than a
formality: the two model AArch64 register access differently, so each
version corroborates cases the other cannot. Capstone 6's model is the
sharper oracle -- it is what caught armlint treating a FEAT_MOPS main
stage as a kill of its own source pointer -- while 5.x is what armlint
ships against and so is what its corrections are written for;
make integration-test is expected to show cosmetic snapshot diffs
under v6 -- 22 of the 100 under Alpha11, which prints shift immediates
in decimal and drops # on PC-relative operands (branch targets,
adr, ldr literals), with every finding identical -- so the
fixtures remain pinned to Capstone 5.x rendering until v6 stabilizes.
armlint is intended to be part of compiler test suites which should
#include "armlint.h" and link libarmlint.a. Disassemble the
just-emitted machine code with check_instructions; its return value is
the number of opportunities found, which a test can assert is zero:
#include "armlint.h" // also includes <capstone/capstone.h>
// code/code_len: the AArch64 bytes to check (e.g. a function the
// compiler just emitted); base_addr is the address they load at.
// Returns the opportunity count (0 == clean), or -1 on a decode error.
int lint(const uint8_t *code, size_t code_len, uint64_t base_addr)
{
csh handle;
if (cs_open(CS_ARCH_ARM64, CS_MODE_ARM, &handle) != CS_ERR_OK) {
return -1;
}
cs_option(handle, CS_OPT_DETAIL, CS_OPT_ON);
armlint_summary *summary = armlint_summary_create();
int findings = check_instructions(
handle, code, code_len, base_addr, /*verbose=*/true, summary,
/*features=*/0, // or ARMLINT_FEATURE_CSSC etc.
/*symbols=*/NULL, /*nsymbols=*/0); // see armlint_symbol
armlint_summary_print(summary); // optional by-type tally
armlint_summary_destroy(summary);
cs_close(&handle);
return findings;
}The summary is optional -- pass NULL to skip the by-type tally --
and verbose controls whether each opportunity is printed as it is
found. armlint can also read arbitrary AArch64 binaries (ELF, thin
Mach-O, or universal/fat Mach-O) directly:
./armlint /path/to/aarch64/binary
./armlint /bin/ls
./armlint -m cssc /bin/ls # also suggest CSSC instructions
./armlint -m v8 jit.elf # a V8 JIT dump from tools/v8dump2elf.py
./armlint -s all /bin/bash # every ARM64 slice of a universal binary
./armlint -d /bin/ls # also report functions that copy one another
./armlint -c /bin/ls # also report the constants MOVZ/MOVK chains buildA universal binary can carry more than one ARM64 slice. macOS 27's
/bin/bash, /usr/bin/ssh and /usr/lib/dyld each hold an arm64e
and an arm64e.x1 slice, two compilations of the same program, so
armlint scans one slice by default -- the lowest variant, arm64
before arm64e before arm64e.x1 -- and names the slices it skipped
on stderr. -s all scans every ARM64 slice and adds up their
findings; -s NAME (-s arm64e.x1) or -s INDEX (the slice's
position in the fat header, as otool -f numbers it) picks one.
arm64e.x1 code signs and authenticates return addresses with
FEAT_PAuth_LR (pacibsppc, retabsppc), which Capstone 5 does not
decode: a scan of that slice skips those words as data, and the PAC
audit and -m pauth do not arm themselves there.
Linker-synthesized import glue is excluded from both the scan and
the census: Mach-O __stubs, __stub_helper, and __objc_stubs,
and ELF .plt, .iplt, and the .plt.* variants. Their shape is
the dynamic-linking ABI's business, not the compiler's -- the
classic lazy-binding __stub_helper alone would otherwise
contribute one spurious "LDR literal foldable" finding per imported
symbol (each fixed entry LDRs its lazy-bind-info offset into w16
from an inline literal), a couple of hundred lines of noise on a
typical minos < 12 Mach-O binary.
-m <feature> enables checks whose rewrites use ISA-extension
instructions the target must support: cssc (Armv8.9/9.4 Common
Short Sequence Compression: smax/smin/umax/umin, abs,
ctz), lrcpc2 (Armv8.4 unscaled store-release: stlur), pauth
(Armv8.3 pointer authentication: retaa/retab), lse
(Armv8.1 atomics: ldadd/ldset/ldeor/ldclr/swp), cmpbr
(Armv9.6 compare-and-branch: cbgt/cbeq/...), sha3
(Armv8.2 three-operand vector logic: eor3, bcax), and fp16
(FEAT_FP16 half-precision transfers: fmov Wd, Hn). pauth
arms automatically on arm64e slices, whose ABI mandates FEAT_PAuth
(the same auto-arm as the PAC audit); the rest stay opt-in. cssc,
sha3 and fp16 name extensions that are never mandatory at any
architecture version, so those three assert a specific target rather
than a version floor.
-m v8 is V8 the JavaScript engine, not an architecture version. It
names an input rather than a target: a V8 JIT dump, the output of
tools/v8dump2elf.py for a pointer-compressed build, and asserts the
two facts about that stream that its checks need -- x28 is the
4GB-aligned pointer-compression cage base, and every
ldr xzr, (literal) word opens a constant pool the scans step over.
Both are knowledge about the scanned code rather than the hardware
and unsound for arbitrary binaries, so the mode stays opt-in.
-a <audit> enables opt-in informational checks whose findings are
review items rather than missed folds; pac audits the binary against
the arm64e-style full pointer-authentication contract (return
addresses spilled unsigned, unauthenticated br/blr, and
zero-discriminator braaz/blraaz -- authenticated, but against a
modifier of zero, so any same-key zero-discriminator pointer in the
process substitutes). Audit findings are review items on a ladder:
the raw-br check recognizes and auto-dismisses the clang jump-table
idiom, so what remains is BLRs, linker veneers, and genuinely
unclassified branches; the zero-discriminator rung marks where the
C-ABI IA+0 signing floor could upgrade to __ptrauth-style
diversified braa/blraa. The PAC audit arms automatically on
arm64e slices (whose ABI already assumes full signing), so macOS
system binaries surface their worklist with no flag; a plain arm64
slice never opted in, so it stays silent unless you pass -a pac
explicitly. imm audits every materialized constant against its
consumer's immediate encoding and tallies the misses by value, so the
constants worth renumbering surface; see Immediate-misfit audit
(-a imm) below.
By default armlint prints only a summary: the opportunities grouped by type and sorted by prevalence, so it is clear which to look at first, followed by a total and the number of instructions scanned. A large binary can have hundreds of thousands of opportunities, so the per-opportunity detail is suppressed unless requested:
$ ./armlint /bin/ls
Optimization opportunities by type:
38 ADD + LDR foldable to immediate-offset LDR
2 AUTIASP/AUTIBSP + RET foldable to RETAA/RETAB (PAuth)
2 CBZ/CBNZ of a live single-bit test foldable to TBZ/TBNZ
42 optimization opportunities in 4153 instructionsPass -v to also print each opportunity -- its one-line summary plus
the offending instructions, as shown below -- ahead of the summary:
$ ./armlint -v /bin/ls
ADD + LDR foldable to immediate-offset LDR at offset: 0x60 <0x10000071c+0x44>: -> ldr w8, [x8, #0x2c] (2 instructions)
add x8, x8, #0x2c
ldr w8, [x8]
...The <...> names the containing function so findings can be triaged
per function: <_addhistnode+0x58> when the symbol table carries a
name (Mach-O nlist, or the ELF .symtab with a .dynsym
fallback), or
the function's start address as above when a stripped binary's
LC_FUNCTION_STARTS still records its boundaries. A binary carrying
neither -- Go's linker, for one, emits its runtime symtab instead of
either structure -- prints the historic unannotated form, as does a
finding past the end of an ELF symbol whose recorded size stops short
of it: a stripped libxul.so keeps only its exports, and the last one
before 40 MB of unnamed code must not claim all of it.
The process exits non-zero when any opportunity is found, so armlint can gate a compiler test suite.
armlint -i replaces the lint scan with a census: every instruction in
the binary's executable sections, attributed to the FEAT_* group it
requires, each group carrying the architecture version it became
mandatory at -- LSE and CRC32 at Armv8.1, LRCPC/FCMA/JSCVT and the
register-form PAC at 8.3, DotProd/LRCPC2/FlagM at 8.4, up through the
MOPS memcpy instructions at 8.8. The resulting ladder answers "what was
this binary compiled for": a -march=armv8.1-a build shows LSE atomics
woven through every mutex, a baseline build shows LDXR/STXR loops with
LSE only inside runtime-dispatched thunks (glibc's outline atomics).
$ ./armlint -i libc.so.6
ISA census: 275103 instructions, 1 undecodable words skipped
Armv8.0 baseline: 275043
mandatory from Armv8.1: LSE (21)
mandatory from Armv8.3: none
...
optional features: none
branch protection (hint space): BTI (21), PAC (18)
highest mandatory-from level: Armv8.1Three groups never raise the ladder, each for a soundness reason of its
own. Features that never become mandatory in the v8 line (the crypto
extensions, FP16, SVE, MTE) are listed as optional: any of them can
be bolted onto an old target with a single +feature flag, so their
presence says nothing about -march. The hint-space branch-protection
forms (PACIASP/AUTIASP/BTI/XPACLRI) execute as NOPs on cores without
the extension -- -mbranch-protection=standard emits them precisely so
the binary stays v8.0-compatible -- so they get their own line and no
version claim; only the register-form PAC instructions (PACIA, RETAA,
BRAA, ...) evidence a real Armv8.3 target. And an undecodable word is
either data in text or an extension this Capstone build cannot decode,
so the skipped count bounds what the census could have missed.
The census reports presence, not requirement: dispatched fast paths
count even though the binary runs without them. Treat small exotic
tallies in a binary with many skipped words with suspicion -- string
pools embedded in text sometimes decode as valid SVE or atomics -- and
use -v, which prints up to four sample addresses per feature, to
check a surprising tally in a disassembler before believing it.
When the binary carries function boundaries (Mach-O
LC_FUNCTION_STARTS/nlist, ELF .symtab/.dynsym -- the same
sources that symbolize -v findings), the census adds per-function
pac-ret coverage:
pac-ret coverage: 866 of 1113 functions sign the return address
A function counts as signed when its span contains PACIASP/PACIBSP or
the register-form pacia/pacib x30, sp. The ratio is a fingerprint,
not a target: leaf functions never spill the return address and
legitimately never sign, so Apple's fully signed arm64e binaries sit
around 78-85% (zsh 866/1113, ssh 912/1073, sshd 360/430) with the
remainder leaves -- the -a pac audit separately confirms zero
unsigned spills there, which is the claim that matters. The
signature of never opting in looks entirely different: Homebrew's
plain-arm64 libcapstone reads 0 of 1441. The line is omitted when no
boundary information exists (Go binaries carry neither structure).
armlint -a imm adds an audit to the lint scan that is the complement
of the mov folds. Those report a movz/movn/movk chain whose
consumer could have taken the value as an immediate; the audit reports
the chain whose consumer has an immediate form the value misses:
movz x16, #0x100, lsl #16 ; cmp x0, x16, where 0x1000000 is one
step past the largest cmp immediate. Four consumer classes are
checked, each against its own encoding: add/sub/adds/subs and
cmp/cmn (a 12-bit immediate, optionally shifted by 12, in either
sign -- cmp x0, #-5 is cmn x0, #5), the logical ops
and/orr/eor/ands/tst (a bitmask immediate; the complement for
bic/orn/eon/bics), ccmp/ccmn (a 5-bit immediate, ccmn
covering -1..-32), and a register-offset ldr/str indexed by the
constant (the scaled 12-bit or the unscaled signed 9-bit offset). The
consumer must follow the chain directly, the operand rules are the
folds' own (the constant in an immediate-capable slot, the other
operand neither the zero register nor the constant's), zero and
all-ones are skipped, and the chain and the consumer may differ in
width where the value survives the change (a W consumer reads the low
half of an X chain, a W chain zero-extends into an X consumer).
Nothing at the site can be rewritten -- the chain is already the
cheapest spelling of that value -- so the finding is informational,
like the PAC audit's, and the review item is the constant itself.
The exception is a plain add or sub of a constant below 2^24,
which needs no register at all: when the constant's register dies,
the two-immediate fold
reports the site too, and the audit still lists the value, since
renumbering it saves one more instruction. A value the code's author
chose (an object-size limit, a sentinel, the bit assignment of a
flags field) that lands one step outside the encoding costs a
materialization and a register operand at every use,
and choosing it one step differently turns every site into the
immediate form. The audit therefore needs no liveness proof -- a
constant hoisted into a live register amortizes its materialization,
but each use still pays for the register operand the immediate form
would not need -- and each finding names what would have encoded. The
summary tallies the audit by value, most frequent first, so the
constants worth renumbering sort to the top:
$ ./armlint -a imm libxul.so
Optimization opportunities by type:
80391 MOV + ADD/SUB/CMP: constant has no add/sub immediate form (imm audit)
...
10279 MOV + AND/ORR/EOR/TST: constant is not a bitmask immediate (imm audit)
...
6364 MOV + CCMP/CCMN: constant is outside imm5 (imm audit)
...
1228 MOV + register-offset LDR/STR: constant has no immediate-offset form (imm audit)
...
Immediate misfits by value (-a imm), within one step of an immediate form:
2707 bitmask imm w #0x50 nearest bitmask immediates 0x10 and 0x40 (1 bit away)
2501 add/sub imm12 w #0x270f nearest add/sub immediates 0x2000 and 0x3000
1489 bitmask imm x #0xfffb000000000000 nearest bitmask immediate 0xffff000000000000 (1 bit away)
1000 add/sub imm12 x #0x15f90 nearest add/sub immediates 0x15000 and 0x16000
...
159 ldr/str offset b #0x10d9 immediate offsets cover -256..255 unscaled, 0..4095 x size scaled
...
(7309 more distinct values)
Immediate misfits by value (-a imm), beyond one step (informational):
11978 add/sub imm12 x #0x8000000000000000 largest add/sub immediate is 0xfff000
5738 add/sub imm12 x #0x7ffffffffffffffd largest add/sub immediate is 0xfff000
...
(7906 more distinct values)
135936 optimization opportunities in 28583984 instructionsEach row is a count of sites, the consumer class, a width letter (the
consumer's register width, w or x; for a load or store its access
size, b/h/w/x), the constant (a load/store misfit is tallied
as its byte offset), and a hint naming what would have encoded: the two
nearest 4 KiB multiples, the bitmask immediates fewest bit flips away,
the imm5 or offset range. A bitmask hint sometimes adds a widened
mask -- 0xdf, the byte with its ASCII case bit clear, is not a
bitmask immediate but 0xffffffdf is, and the two agree on any operand
whose bits above the mask are known clear -- which is a compiler's fix
(known bits) rather than a renumbering. Each table prints its top 40
rows and counts the rest.
Two tables, because a real corpus is dominated by constants no
renumbering reaches. The first holds the constants within one step of
an encodable neighbour: an add/sub magnitude up to 0x1000000, so a
4 KiB multiple lies within 4 KiB of it; a bitmask at most two bit
flips away; a ccmp magnitude up to 63; a load/store byte offset
within twice the scaled range of its access size, or the 256 bytes
below the unscaled range. That table is the renumbering worklist, and
its rows have owners: in the Firefox libxul.so above, 0x50 under
tst is JS::shadow::Zone::GCState with Sweep = 4 and Compact = 6,
a two-bit mask one flip from contiguous; 0x270f is nsAtom's
kAtomGCThreshold = 10000, tested as ++count >= 10000 and so
compared against 9999, where 8193 would encode; the byte loads at
0x10d9 are Document bit-field bytes just past the 4 KiB byte-load
range. The second table holds everything beyond reach, whose whole
magnitude is the point: INT64_MIN under cmp (the niche rustc
gives the dataless variants of a Vec- or String-carrying enum,
plus Gecko's TimeDuration sentinel), the JS::Value tags, the words
of the interface IDs every QueryInterface compares.
To find the sites of one value, -v prints each finding as
MOV + ADD/SUB/CMP: ... at offset: 0x... <function+0x...>: -> #0x270f; nearest ..., so a grep for the constant lists them with their
containing functions, which point at the definition to change. The
findings count toward the total and the non-zero exit status like any
other, so a test suite that gates on a clean run should not pass
-a imm; the worklist is for reading.
analyses.md walks the
rules and two corpora in full: this libxul.so run and V8's JetStream 3
JIT output (-m v8 -a imm on a tools/v8dump2elf.py dump), where
the table opened with a 16 MiB reservation bound one step past
0xfff000 at 92,109 sites, since fixed in V8 by comparing against the
largest encodable bound.
armlint -d adds a report on which functions are copies of one
another, and on which findings are the same advice counted again in
each copy. A function is a span between the boundaries -v uses to
name functions: Mach-O nlist and LC_FUNCTION_STARTS, or the ELF
.symtab (.dynsym as a fallback), where a sized symbol ends at its
size. An executable section with no boundary inside it counts as one
function, which is how a JIT dump's code blobs arrive. Trailing NOP
and zero padding is not part of a function.
Each function is keyed twice:
-
The identical key is the code as it executes. A branch,
adr,adrpor literal load that reaches outside the function counts by the absolute address it resolves to. One that stays inside keeps its encoding, which is relative to the function. Two functions with the same identical key behave the same wherever they are loaded, so a linker's identical-code folding could keep just one. -
The up to constants key also masks the fields that tell instances of one piece of code apart:
movz/movn/movkimmediates;- the words the function's own literal loads read (and V8 constant
pools, under
-m v8); - every target outside the function;
- the low 12 bits an
addor a load/store adds to anadrppage.
This merges a template's or a generic's instances that differ only in what they call and address, and a JIT's copies of one stub that differ only in the pointers they embed. It also merges two stubs that differ only in a constant they guard. That is the intended reading of "duplicate code", but keep it in mind when reading the numbers.
$ ./armlint -d librustc_driver.dylib
Optimization opportunities by type:
...
Duplicate code (-d): 174219 functions, 103924556 bytes
identical: 36353 copies of 5900 functions, 4591292 bytes (4.4%)
up to constants: 86331 copies of 11651 functions, 24890780 bytes (24.0%)
how often each distinct function appears, up to constants: once 76237, 2-10 times 10101, 11-100 times 1484, more 66
findings: 18673 of 96161 (19.4%) repeat an earlier copy's at the same offset
7510 of 27595 ADD/SUB immediate chain foldable to one
3086 of 16891 register already holds the recomputed value
2028 of 6156 branch to the next instruction is a no-op
...
96161 optimization opportunities in 25974168 instructionsIn scan order, the first function with a given key is the original
and the later ones are its copies. Percentages are of the bytes in
functions. The rustc above keeps 4.4% of its code in identical copies
that the linker did not fold; one example is 350 copies of a single
168-byte LLVM DenseMap lookup. A further fifth of its code is
generic and template instances that differ only in what they call and
address.
A finding repeats when an earlier copy of its function (up to
constants) has a finding of the same type at the same offset: the same
advice to whatever emitted the code, counted again. The by-type table,
the total and the exit status still count every finding. The
findings lines say how much of that count is repetition, overall and
by type. -v marks each repeated finding by appending [repeat] to
its header line. tools/v8_analyze_findings.py and
tools/jsc_analyze_findings.py read that marker and add a repeats row
under their per-tier tables. -v also lists the most duplicated
functions. Here it is on SpiderMonkey's Octane code (a
tools/smdump2elf.py dump):
$ ./armlint -d -v sm-octane.elf
...
most duplicated up to constants, by bytes past the original:
copies variants bytes findings distinct original
86 86 3568 1118 13 <BL_f1e37185c0>
86 86 3424 1290 15 <BL_f1e371c3b0>
118 118 240 472 4 <ION_f1e3379700>
...The columns:
variants: how many distinct identical keys the copies have. 1 means every copy is the same function byte for byte.findings: the findings in all the copies together.distinct: how many of those findings are not repeats.
The first row above is 86 Baseline blocks of 3,568 bytes each. They are the same code up to constants, and no two are byte-identical. Their 1,118 findings are 13 pieces of advice.
| corpus | functions | identical | up to constants | findings repeated |
|---|---|---|---|---|
| librustc_driver | 174,219 | 4.4% | 24.0% | 19.4% |
| HotSpot, javac (every tier) | 24,739 | 4.6% | 10.9% | 24.9% |
| SpiderMonkey, Octane | 5,579 | 0.7% | 9.1% | 14.6% |
| V8, Octane | 2,534 | 0.9% | 1.7% | 0.6% |
Most of HotSpot's repeats (51,417 of 66,222) are in C1's per-method stubs, every one of whose 86,790 findings is a "suboptimal MOVZ/MOVK sequence". C2's own code is barely repeated: 382 of its 21,371 findings.
Limitations:
- Keys are 64-bit hashes plus the function's length, so a false merge is possible but improbable.
- In an unlinked object file every relocated field reads as zero, so functions that differ only in a relocated target key alike.
- A stripped ELF with only
.dynsymhas boundaries for its exports alone, so-dsees only those functions. - With
-s all, every slice feeds one table. A function that both compilations contain counts as a copy, and its findings as repeats. - The
adrppage tracking is linear, not flow-sensitive. - Pointers stored as data inside a blob, such as a JIT's absolute dispatch table, are not masked.
armlint -c adds a report on the constants the code builds with
movz/movn + movk chains. A chain is a run of two or more of them
into one register at one width: the runs the "suboptimal MOVZ/MOVK
sequence" check judges. The -a imm audit
reports a constant that just misses its consumer's immediate form.
This report is its complement: it counts what every constant too wide
for one instruction costs. Constants are tallied by width and value,
so w #0x1deb8 and x #0x1deb8 are two values, and ranked by the
instructions spent on them.
Within a function, a chain that builds a value an earlier chain of the
same function already built is a rebuild. The functions are the
ones -d uses. A rebuild spends the instructions again instead of
keeping the value in a register. The usual causes are a call that
clobbered the register and a loop invariant that was not hoisted. Not
every rebuild is a miss: the two arms of an if/else both need their
build, and a rebuild after a call can be cheaper than a spill. The
rebuild counts bound what keeping the values could save; they do not
measure it.
$ ./armlint -c librustc_driver.dylib
Optimization opportunities by type:
...
Constant chains (-c): 25981139 words, 174219 functions
every chain: 73533 building 17385 distinct values, 208685 instructions (0.80%)
rebuilds within a function: 22192 in 4922 functions, 59481 instructions (0.23%; 28.5% of chain instructions)
how often a function builds a value it rebuilds: twice 5825, 3 times 1381, 4 times 676, 5 or more times 1220
most instructions spent building one value:
instructions chains value
19792 4948 x #0xf1357aea2e62a9c5
6180 1545 x #0xbf58476d1ce4e5b9
3774 1887 w #0x7a3e8
3302 1651 w #0x1deb8
...
(17365 more distinct values)
most instructions spent rebuilding one value within a function:
instructions rebuilds functions value
6616 1654 826 x #0xf1357aea2e62a9c5
1952 976 246 w #0x1deb8
...
(3323 more distinct values)
96161 optimization opportunities in 25974168 instructionsEvery word of the executable sections is scanned, except V8 constant
pools under -m v8, and the percentages are of those words. Each
ranking lists its top 20 values. Here x #0xf1357aea2e62a9c5 is the
multiplier of rustc-hash's FxHasher, built with four instructions at
every inlined hash, and w #0x1deb8 is a field offset into rustc's
global context. -v adds where to look. It gives the addresses of a
value's first three chains; for a rebuilt value, it gives the function
that builds it most often and its first builds there:
instructions rebuilds functions value
6616 1654 826 x #0xf1357aea2e62a9c5 built 120 times in <__RNvNtCshPlmC27tfnj_16rustc_query_impl9execution25collect_active_query_jobs>: 0x2603918, 0x2603f44, 0x260452c (+117 more)| corpus | chain instructions | distinct values | rebuilt |
|---|---|---|---|
| librustc_driver | 0.80% | 17,385 | 28.5% |
| clang-24 | 1.10% | 22,502 | 33.6% |
| Firefox libxul.so | 0.61% | 11,753 | not counted |
| HotSpot, javac (every tier) | 17.8% | 44,350 | 63.9% |
| SpiderMonkey, Octane | 10.2% | 7,465 | 58.2% |
V8, Octane (-m v8) |
3.9% | 485 | 73.6% |
The chain instructions are a share of the words scanned, and the rebuilt ones a share of the chain instructions.
- clang's costliest value is
x #0xbf58476d1ce4e5b9, splitmix64's multiplier. LLVM'sDenseMaphashes every pointer key by multiplying it by that constant (densemap::detail::mix). The value takes 45,804 instructions in 11,451 chains, 19% of clang's chain instructions. - libxul.so is stripped. Its
.dynsymbounds 5 functions in the text, so its rebuilds go uncounted. Its totals are led by SpiderMonkey'sUndefinedValue()(x #0xfff9800000000000, 7,661 chains) and by nsresult codes:NS_ERROR_ILLEGAL_VALUE(w #0x80070057, 3,616) andNS_ERROR_FAILURE(w #0x80004005, 3,337). - HotSpot's costliest value is
x #0x0: 112,251 three-instruction chains, 26% of its chain instructions, all but 8 of them in the stubs C1 and C2 append to each method. The stubisb ; mov x12, #0 ; movk ; movk ; mov x8, #0 ; movk ; movk ; br x8calls the interpreter, and its method and entry address are placeholders, patched when the call is resolved. A method's stubs are one function, so the same placeholders also lead the rebuilds. - SpiderMonkey's costliest value is the
JS::ValueInt32 tag,x #0xfff8800000000000(12,865 chains).UndefinedValue()(8,901),Int32Value(2)andInt32Value(1)join it in the top six, between runtime pointers. One Baseline blob rebuilds the Int32 tag 2,773 times. - V8's costliest value,
x #0x17cb0007c000, is the address of a typed array's backing store. Maglev and TurboFan embed it in the code of Octane's Mandreel benchmark, the Bullet physics engine compiled to JavaScript: 3,992 chains in 146 functions, most of them followed directly by an indexed access such asstr w1, [x3, x0, lsl #2]. One Maglev function rebuilds it 222 times.
Limitations:
- Only move-wide chains count. A constant loaded from a literal pool,
built by an
orr+movkhybrid, or encoded in one instruction is not counted. - A W chain zero-extends into its X register, so
w #0x1deb8andx #0x1deb8leave the same register contents. They still count as two values, and a build of one does not make the other a rebuild. - A rebuild is a build of the same value earlier in the function, in address order. Neither dominance nor liveness is checked.
- A JIT patches many of its constants at runtime. The census cannot tell a patch site's placeholder from a constant.
- The functions are
-d's. A stripped ELF with only.dynsymhas functions for its exports alone. A binary with no boundaries at all is one function per section, which makes every repeated value in a section a rebuild. - With
-s all, every slice feeds one table.
tools/ holds the research utilities that feed armlint's check
backlog -- and one that checks the checks -- built separately with
make tools:
tools/pairscancounts adjacent-instruction pairs by normalized shape (registers collapsed to classes, immediates to#0/#i) across the executable sections of ELF and Mach-O binaries, surfacing frequent patterns worth a new check.-e SUBSTRprints example sites for shapes matching a substring.tools/defuseprofiles block-local def-to-use distances (how far a value's sole consumer sits from its producer) and multi-instruction redundancies no pair statistic can see: dead definitions, redundant reloads of the same address, re-materialized constants, and zero compares of a value whose producer could have set the flags.tools/shapescan.pycounts a fixed list of specific candidates -- the rows TODO.md tracks -- with their real operand, range and encodability conditions applied, which is what separates a population from a pair count.adrp+addis the standing example: 753,648 adjacent dependent pairs across the corpus, of which 43,434 have a target inside ADR's reach. Runtools/shapescan.py --selftestfirst: it assembles every reference instance with clang and checks each mask in both directions -- that it matches no instruction belonging to another mask, and that it matches every spelling of its own listed inALSO. The second half is the one that matters most, because a mask too narrow to see thestur, theldp q, or the 64-bit form of its shape reports a small number rather than a wrong one, and nothing looks broken.-e SHAPEprints example sites. Needs numpy.tools/addpairscan.pysizes the ADD-immediate +LDP/STPfamily, which one shape mask cannot split honestly: the half whose combined offset fits the pair's signed, pre-scaledimm7folds two instructions into one, while the half that overflows it can only be split into two singles at no size saving. 8,775 against 17,565 across the corpus. Carries its own--selftest; needs numpy.tools/rwfuzzchecks a shipped check's advice by running it. It generates random AArch64 programs, applies each rewrite armlint suggests for one check to a copy, and runs original and rewrite natively from the same random registers, flags and memory: any difference is a soundness bug. A control arm applies the rewrite where armlint refused to, so a harness that stopped seeing differences would show it. Modes cover the redundant zero-extension, shift + mask, CMP + CCMP chain, value-numbering, dead-write and CMP-to-CMN checks (tools/rwfuzz -n 100000 zext); a new rewrite gets a mode of its own. Needs an AArch64 host, since the programs run natively, and linkslibarmlint.a.tools/tblgen_audit.pychecks armlint's flag and register model against LLVM's own instruction definitions.llvm-tblgen --dump-jsononAArch64.tdlists every A64 instruction with its encoding and its implicit NZCV and register defs and uses; the tool builds one word per record, hastools/tblgen_probeprint Capstone's access lists next to armlint's classifiers for it, and then runs armlint itself overmov ; W ; cbz,mov ; W ; movandcmp ; b.ne ; W ; b.neprobes, where a finding means a write or a read of W's went unmodelled. That second step is what proves a blind spot, since armlint already corrects much of what Capstone omits. It found FJCVTZS, SUBPS, the SVE compares and the MOPS source pointer;--masksderives the encoding masks for a new family. Needs a builtllvm-tblgenand clang.
The workflow that produced several of the current checks: compile a
representative corpus, run pairscan to rank pair shapes, classify
the top shapes as by-design or foldable, then use defuse to decide
whether a candidate needs adjacency only or a liveness window, and
shapescan.py to size the survivor with its real conditions applied
before writing any code. pairscan and defuse lean on Capstone's
register-access model, which mis-reports the compare aliases
(CMP/CMN/TST mark their first operand as a write); defuse
corrects for this, and armlint's own checks decode the raw encodings
precisely to avoid that class of problem. shapescan.py decodes raw
encodings for the same reason, and self-tests them because that is
where its own bugs live: every wrong figure it has produced came from
a mask one bit too loose or a destination modelled as write-only when
the instruction merges into it.
- Arm A-profile A64 Instruction Set Architecture - per-instruction reference, including alias conditions
- Arm Cortex-A optimization guides - per-microarchitecture tuning notes
- Apple Silicon CPU Optimization Guide - Apple M-series tuning notes
- Encoding of immediate values on AArch64 - the bitmask-immediate scheme explained
- Capstone disassembly framework - library to parse instructions
- x86lint - x86-64 equivalent of armlint
Copyright (C) 2026 Andrew Gaul
Licensed under the Apache License, Version 2.0