An AVX-512 IPv6 text-to-bytes parser, written from scratch in C
(needs AVX-512 VBMI/VBMI2/BW/VL). It produces 16 network-order bytes and handles
full 8-group addresses and :: zero compression on the SIMD path; embedded
IPv4 and malformed input fall back to a scalar reference matching inet_pton.
On a Ryzen 9 9950X3D, parsing 100k random full-form addresses with boost disabled, the GCC build takes 12.61 cycles/address versus 21.51 for Shreesh Adiga's parser: 1.71x throughput, or 41% fewer cycles.
| parser | compiler | ns/addr | Mv/s | cyc/addr | instr/addr |
|---|---|---|---|---|---|
| parse_ipv6 | gcc | 2.922 | 342.3 | 12.61 | 63 |
| parse_ipv6 | clang | 3.063 | 326.5 | 13.21 | 58 |
| Adiga/Lemire | g++ | 4.982 | 200.7 | 21.51 | 110 |
| inet_pton | gcc | 130.812 | 7.6 | 564.13 | 820.6 |
The parser came out of Daniel Lemire's "Parsing IP addresses crazily fast"
post; the bundled reference is Shreesh Adiga's parse_ipv6_avx512.
./compare checks both implementations against inet_pton with explicit
accept/reject cases, four million generated valid addresses including ::
and IPv4 tails, and every address in any supplied input file. Successful
results must match all 16 output bytes.
The merged validity check also passed 1,646,384 additional cases against the scalar reference under GCC and Clang: all-byte substitutions across full and compressed forms with group widths 1..4 and every legal gap position, one million arbitrary-byte and syntax-biased inputs, and guarded page-boundary loads. These checks cover embedded NUL, high-bit bytes, and masked-load bounds.
The headline results were measured on 2026-10-06 with GCC/G++ 16.2.1,
Clang 23.1.1, and Linux 7.1.11-arch1-1. Builds use -O3 -march=native without
PGO. Each parser inlines into its own loop, uses the same 100k-address input
set within a run, and materializes its output through the same keep16()
barrier. Hardware counters collect user-space cycles and retired instructions.
The table reports medians of three independent processes, each using its
fastest timed pass after warming the parser. CPU 9 runs the benchmark at
FIFO 99; CPU 11 runs the controller. CPUs 9-11 and 25-27 are isolated, the
SMT siblings remain idle, and no other experiment runs during measurement.
The governor is performance, boost is disabled, and the measured clock is
4.31 GHz. Both the GCC and Clang comparisons use the reference compiled by
g++, its faster compiler in the earlier comparison.
Both SIMD paths end in the same gather tail. For the clean 8-group form,
vpcompressb pulls the colon offsets to give each group its (start,end),
then a single vpermb gathers every group's nibbles into a right-aligned
slot and vpmaddubsw + vpmovwb fold the 16 bytes out. The :: path runs
the same (start,end) step on the contiguous fields, drops the zero-width gap
with vpcompressb, vpexpandbs the survivors back into 8 slots with the
implied zero groups inserted, and joins the gather tail. Hex digits are
translated by one vpermi2b lookup. Validation ORs the original bytes with
the translated nibbles before extracting their sign bits, then masks off
inactive lanes and colons. This catches both invalid table entries and
high-bit input bytes that would otherwise alias ASCII in the 7-bit lookup.
Group-offset calculation and the gather contain a dependent chain of cross-lane permutes; nibble translation can run in parallel with the offset calculation. This parser issues less work along that chain than the compress-then-expand reference.
./build.sh # builds ./compare and ./bench; CC/CXX to override compilers
CC=clang CXX=g++ ./build.sh # Clang parser with the same G++ reference./compare is the full check. It validates both this parser and the reference
against inet_pton (self-tests, a 4M-iteration fuzz, and a per-address
cross-check on any file you pass), then benchmarks all four. The default build
uses gcc for this C parser and g++ for the C++ reference. Each parser inlines
into its own bench loop, driven over identical inputs by the same timer.
The reference is bundled at ext/avx512ip.h, so no extra checkout is needed.
Run it on a pinned core at real-time priority, to keep the numbers clean:
taskset -c 9 chrt -f 99 ./compare # random addresses
taskset -c 9 chrt -f 99 ./compare addresses.txt # one address per line./bench is the standalone path: just main.c and the parser header, no C++
and no third-party reference. Same tests, fuzz, and benchmark, only without the
comparison.
taskset -c 9 chrt -f 99 ./benchThe current full-form results are in the headline table. GCC takes about 4.5% fewer cycles than Clang. The reference retires 110 instructions/address versus 63 for the GCC parser; its higher instruction parallelism does not offset the additional work.
CPU 9 is on the 32 MiB L3 CCD and has the highest CPPC performance rank. CPU 1 is the highest-ranked core on the 96 MiB V-cache CCD. Earlier CPU-selection measurements used 13.36-13.37 cycles/address on CPU 9 and 15.11-15.13 on CPU 1. The 6.48 MiB input set fits either L3; the V-cache CCD becomes relevant when the working set exceeds 32 MiB.
Both paths now extract the sign bits of v | nib once. Inactive source bytes
are zero and ':' has no sign bit, so masking this combined result with
active & ~mcolon preserves the previous checks. GCC replaces two
vpmovb2m instructions with one vector OR and one vpmovb2m, and removes
associated mask/GPR work. Retired instructions fall from 65 to 63 on the
clean path and from 100 to 96 on the compressed path. Clang's corresponding
counts fall from 61 to 58 and from 98 to 94.
The paired experiment used 100k generated addresses, fixed seed
0x123456789abcdef0, and alternating variant order. Clean input uses
make_ipv6_string flavor 0; compressed input uses flavor 1, which forces four
zero groups after the first hextet. The driver uses the normal heap allocation
and timing loop. These are medians of two processes per variant, on the same
CPU and fixed-clock setup as the headline comparison:
| compiler | path | before cycles/addr | after cycles/addr | reduction |
|---|---|---|---|---|
| gcc | clean | 13.48 | 12.61 | 6.5% |
| gcc | compressed | 27.65 | 26.87 | 2.8% |
| clang | clean | 13.53 | 13.17 | 2.7% |
| clang | compressed | 27.80 | 26.43 | 4.9% |
The gain also held across varied gap positions and a randomized 80/20 clean/compressed workload. The latter used 21.74 to 20.80 cycles/address with GCC and 22.16 to 21.84 with Clang.
Two related alternatives were rejected. Mapping colons to zero and filling
inactive lanes with '0' removes the remaining validity mask, but increases
GCC's clean-path cost from 13.48 to 14.82 cycles/address; Clang's additional
improvement is small. Reconstructing compressed starts from expanded ends
removes a compress/expand pair but adds dependent shift/add/broadcast work
and scalar mask generation. It increases compressed-path cost from 27.65 to
32.71 cycles with GCC and from 27.80 to 30.38 with Clang. Combining it with
the validity changes also regresses compressed parsing. Accepting colons in
the table while keeping the zero-masked load and & active trades small gains
between workloads without a clear advantage over the retained change.
Original README results, retained unchanged. Random 100k addresses,
taskset -c 1 chrt -f 99. The cycle and time columns give an effective clock
of about 5.44 GHz; the boost state was not recorded:
| parser | compiler | ns/addr | Mv/s | cyc/addr | instr | i/c |
|---|---|---|---|---|---|---|
| parse_ipv6 | gcc | 3.12 | 320.9 | 16.98 | 66 | 3.89 |
| parse_ipv6 | clang | 3.12 | 320.9 | 16.98 | 61 | 3.59 |
| Adiga/Lemire | g++ | 4.80 | 208.5 | 26.14 | 109 | 4.17 |
| Adiga/Lemire | clang++ | 4.87 | 205.5 | 26.52 | 119 | 4.49 |
| inet_pton | gcc | 91.03 | 11.0 | 495.3 | 829 | 1.67 |
Real TUM hitlist, 1M shuffled, ~20% with ::: 5.4 ns / 29 c here vs 6.8 ns /
37 c for the reference.
The cycle count is identical on gcc and clang, unlike in the original benchmark - see below for why. The reference shows a higher i/c, but that is not a win for it: it issues more, more-parallel instructions to do the same job. The cycle and instruction columns are where this parser pulls ahead.
The separate boost-disabled pre-swap baseline at 4.45-4.46 GHz measured 16.81-16.82 cycles/address for this parser and 26.04-26.06 for Adiga/Lemire. Against those ranges, preferred CPU 9 on the boost-disabled 9950X3D uses about 25% fewer cycles for this parser and 17% fewer for the reference.
The reference benchmark sums the 16 result bytes so the optimizer cannot
delete the parse. That sum is nearly half of what gets timed (about 15 cycles
on top of a ~17-cycle parse), and it compiles differently per compiler: gcc
builds a scalar add chain, clang folds it to one vpsadbw. So the clang-vs-gcc
gap in the original numbers is the sum loop, not the parser.
bench_ctx.h replaces the sum with keep16(): an empty asm with an "m"
constraint on the 16 result bytes. It emits zero instructions, just tells the
compiler the bytes are observed so the store cannot be elided. Same idea as
Google Benchmark's DoNotOptimize. The timed loop includes parsing, input
traversal, and counting successes, without the extra byte-summing work.
parse_ipv6.h the parser (header-only, inlines into the caller's loop)
perf.h perf_event_open wrapper (cycles + instructions)
bench_ctx.h shared bench context and the keep16 barrier
main.c driver: rng, tests, fuzz, file loader, timing
bench_lemire.cpp the reference parser's bench loop, its own C++ TU
ext/avx512ip.h the reference parser, third-party
The TUM hitlist row above is the TUM IPv6 Hitlist Service's open list of
responsive addresses. It is one canonical address per line, which is what
./compare <file> expects, so decompress it and pass it straight in:
curl -O https://alcatraz.net.in.tum.de/ipv6-hitlist-service/open/responsive-addresses.txt.xz
unxz responsive-addresses.txt.xz
taskset -c 9 chrt -f 99 ./compare responsive-addresses.txtNo registration is needed for that list. main.c skips blank lines, #
comments, and anything over 45 characters, and ./compare cross-checks every
loaded address against inet_pton before timing.
MIT, Copyright (c) 2026 Peter Fors.
ext/avx512ip.h is Shreesh Adiga's parser from Daniel Lemire's blog,
reproduced with permission; it keeps its own attribution and is not covered by
the line above.