Skip to content

About

AVX-512 IPv6 text-to-bytes parser in C for Zen 5 — ~13.5 cycles/address on random inputs, with :: compression and inet_pton validation.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

5 Commits

Folders and files

Repository files navigation

parse_ipv6

An AVX-512 IPv6 text-to-bytes parser, written from scratch in C (needs AVX-512 VBMI/VBMI2/BW/VL). It produces 16 network-order bytes and handles full 8-group addresses and :: zero compression on the SIMD path; embedded IPv4 and malformed input fall back to a scalar reference matching inet_pton.

On a Ryzen 9 9950X3D, parsing 100k random full-form addresses with boost disabled, the GCC build takes 12.61 cycles/address versus 21.51 for Shreesh Adiga's parser: 1.71x throughput, or 41% fewer cycles.

parser compiler ns/addr Mv/s cyc/addr instr/addr
parse_ipv6 gcc 2.922 342.3 12.61 63
parse_ipv6 clang 3.063 326.5 13.21 58
Adiga/Lemire g++ 4.982 200.7 21.51 110
inet_pton gcc 130.812 7.6 564.13 820.6

Correctness and measurement

The parser came out of Daniel Lemire's "Parsing IP addresses crazily fast" post; the bundled reference is Shreesh Adiga's parse_ipv6_avx512. ./compare checks both implementations against inet_pton with explicit accept/reject cases, four million generated valid addresses including :: and IPv4 tails, and every address in any supplied input file. Successful results must match all 16 output bytes.

The merged validity check also passed 1,646,384 additional cases against the scalar reference under GCC and Clang: all-byte substitutions across full and compressed forms with group widths 1..4 and every legal gap position, one million arbitrary-byte and syntax-biased inputs, and guarded page-boundary loads. These checks cover embedded NUL, high-bit bytes, and masked-load bounds.

The headline results were measured on 2026-10-06 with GCC/G++ 16.2.1, Clang 23.1.1, and Linux 7.1.11-arch1-1. Builds use -O3 -march=native without PGO. Each parser inlines into its own loop, uses the same 100k-address input set within a run, and materializes its output through the same keep16() barrier. Hardware counters collect user-space cycles and retired instructions.

The table reports medians of three independent processes, each using its fastest timed pass after warming the parser. CPU 9 runs the benchmark at FIFO 99; CPU 11 runs the controller. CPUs 9-11 and 25-27 are isolated, the SMT siblings remain idle, and no other experiment runs during measurement. The governor is performance, boost is disabled, and the measured clock is 4.31 GHz. Both the GCC and Clang comparisons use the reference compiled by g++, its faster compiler in the earlier comparison.

How it works

Both SIMD paths end in the same gather tail. For the clean 8-group form, vpcompressb pulls the colon offsets to give each group its (start,end), then a single vpermb gathers every group's nibbles into a right-aligned slot and vpmaddubsw + vpmovwb fold the 16 bytes out. The :: path runs the same (start,end) step on the contiguous fields, drops the zero-width gap with vpcompressb, vpexpandbs the survivors back into 8 slots with the implied zero groups inserted, and joins the gather tail. Hex digits are translated by one vpermi2b lookup. Validation ORs the original bytes with the translated nibbles before extracting their sign bits, then masks off inactive lanes and colons. This catches both invalid table entries and high-bit input bytes that would otherwise alias ASCII in the 7-bit lookup.

Group-offset calculation and the gather contain a dependent chain of cross-lane permutes; nibble translation can run in parallel with the offset calculation. This parser issues less work along that chain than the compress-then-expand reference.

Build and run

./build.sh          # builds ./compare and ./bench; CC/CXX to override compilers
CC=clang CXX=g++ ./build.sh   # Clang parser with the same G++ reference

./compare is the full check. It validates both this parser and the reference against inet_pton (self-tests, a 4M-iteration fuzz, and a per-address cross-check on any file you pass), then benchmarks all four. The default build uses gcc for this C parser and g++ for the C++ reference. Each parser inlines into its own bench loop, driven over identical inputs by the same timer. The reference is bundled at ext/avx512ip.h, so no extra checkout is needed. Run it on a pinned core at real-time priority, to keep the numbers clean:

taskset -c 9 chrt -f 99 ./compare                 # random addresses
taskset -c 9 chrt -f 99 ./compare addresses.txt   # one address per line

./bench is the standalone path: just main.c and the parser header, no C++ and no third-party reference. Same tests, fuzz, and benchmark, only without the comparison.

taskset -c 9 chrt -f 99 ./bench

Numbers

Ryzen 9 9950X3D

The current full-form results are in the headline table. GCC takes about 4.5% fewer cycles than Clang. The reference retires 110 instructions/address versus 63 for the GCC parser; its higher instruction parallelism does not offset the additional work.

CPU 9 is on the 32 MiB L3 CCD and has the highest CPPC performance rank. CPU 1 is the highest-ranked core on the 96 MiB V-cache CCD. Earlier CPU-selection measurements used 13.36-13.37 cycles/address on CPU 9 and 15.11-15.13 on CPU 1. The 6.48 MiB input set fits either L3; the V-cache CCD becomes relevant when the working set exceeds 32 MiB.

Merged validity check

Both paths now extract the sign bits of v | nib once. Inactive source bytes are zero and ':' has no sign bit, so masking this combined result with active & ~mcolon preserves the previous checks. GCC replaces two vpmovb2m instructions with one vector OR and one vpmovb2m, and removes associated mask/GPR work. Retired instructions fall from 65 to 63 on the clean path and from 100 to 96 on the compressed path. Clang's corresponding counts fall from 61 to 58 and from 98 to 94.

The paired experiment used 100k generated addresses, fixed seed 0x123456789abcdef0, and alternating variant order. Clean input uses make_ipv6_string flavor 0; compressed input uses flavor 1, which forces four zero groups after the first hextet. The driver uses the normal heap allocation and timing loop. These are medians of two processes per variant, on the same CPU and fixed-clock setup as the headline comparison:

compiler path before cycles/addr after cycles/addr reduction
gcc clean 13.48 12.61 6.5%
gcc compressed 27.65 26.87 2.8%
clang clean 13.53 13.17 2.7%
clang compressed 27.80 26.43 4.9%

The gain also held across varied gap positions and a randomized 80/20 clean/compressed workload. The latter used 21.74 to 20.80 cycles/address with GCC and 22.16 to 21.84 with Clang.

Two related alternatives were rejected. Mapping colons to zero and filling inactive lanes with '0' removes the remaining validity mask, but increases GCC's clean-path cost from 13.48 to 14.82 cycles/address; Clang's additional improvement is small. Reconstructing compressed starts from expanded ends removes a compress/expand pair but adds dependent shift/add/broadcast work and scalar mask generation. It increases compressed-path cost from 27.65 to 32.71 cycles with GCC and from 27.80 to 30.38 with Clang. Combining it with the validity changes also regresses compressed parsing. Accepting colons in the table while keeping the zero-masked load and & active trades small gains between workloads without a clear advantage over the retained change.

Ryzen 9 7950X

Original README results, retained unchanged. Random 100k addresses, taskset -c 1 chrt -f 99. The cycle and time columns give an effective clock of about 5.44 GHz; the boost state was not recorded:

parser compiler ns/addr Mv/s cyc/addr instr i/c
parse_ipv6 gcc 3.12 320.9 16.98 66 3.89
parse_ipv6 clang 3.12 320.9 16.98 61 3.59
Adiga/Lemire g++ 4.80 208.5 26.14 109 4.17
Adiga/Lemire clang++ 4.87 205.5 26.52 119 4.49
inet_pton gcc 91.03 11.0 495.3 829 1.67

Real TUM hitlist, 1M shuffled, ~20% with ::: 5.4 ns / 29 c here vs 6.8 ns / 37 c for the reference.

The cycle count is identical on gcc and clang, unlike in the original benchmark - see below for why. The reference shows a higher i/c, but that is not a win for it: it issues more, more-parallel instructions to do the same job. The cycle and instruction columns are where this parser pulls ahead.

The separate boost-disabled pre-swap baseline at 4.45-4.46 GHz measured 16.81-16.82 cycles/address for this parser and 26.04-26.06 for Adiga/Lemire. Against those ranges, preferred CPU 9 on the boost-disabled 9950X3D uses about 25% fewer cycles for this parser and 17% fewer for the reference.

The benchmark, and why it changed

The reference benchmark sums the 16 result bytes so the optimizer cannot delete the parse. That sum is nearly half of what gets timed (about 15 cycles on top of a ~17-cycle parse), and it compiles differently per compiler: gcc builds a scalar add chain, clang folds it to one vpsadbw. So the clang-vs-gcc gap in the original numbers is the sum loop, not the parser.

bench_ctx.h replaces the sum with keep16(): an empty asm with an "m" constraint on the 16 result bytes. It emits zero instructions, just tells the compiler the bytes are observed so the store cannot be elided. Same idea as Google Benchmark's DoNotOptimize. The timed loop includes parsing, input traversal, and counting successes, without the extra byte-summing work.

Layout

parse_ipv6.h        the parser (header-only, inlines into the caller's loop)
perf.h              perf_event_open wrapper (cycles + instructions)
bench_ctx.h         shared bench context and the keep16 barrier
main.c              driver: rng, tests, fuzz, file loader, timing
bench_lemire.cpp    the reference parser's bench loop, its own C++ TU
ext/avx512ip.h      the reference parser, third-party

Test data

The TUM hitlist row above is the TUM IPv6 Hitlist Service's open list of responsive addresses. It is one canonical address per line, which is what ./compare <file> expects, so decompress it and pass it straight in:

curl -O https://alcatraz.net.in.tum.de/ipv6-hitlist-service/open/responsive-addresses.txt.xz
unxz responsive-addresses.txt.xz
taskset -c 9 chrt -f 99 ./compare responsive-addresses.txt

No registration is needed for that list. main.c skips blank lines, # comments, and anything over 45 characters, and ./compare cross-checks every loaded address against inet_pton before timing.

License

MIT, Copyright (c) 2026 Peter Fors.

ext/avx512ip.h is Shreesh Adiga's parser from Daniel Lemire's blog, reproduced with permission; it keeps its own attribution and is not covered by the line above.

About

AVX-512 IPv6 text-to-bytes parser in C for Zen 5 — ~13.5 cycles/address on random inputs, with :: compression and inet_pton validation.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages