Skip to content

base64 AVX2 - Use half avx2 to smooth the performance cliffs on small payloads when encoding - #1041

Open
gaspardpetit wants to merge 1 commit into
simdutf:masterfrom
gaspardpetit:master
Open

gaspardpetit wants to merge 1 commit into
simdutf:masterfrom
gaspardpetit:master

Conversation

@gaspardpetit

@gaspardpetit gaspardpetit commented Sep 19, 2026

Copy link
Copy Markdown
Contributor

Short title (summary):

Smooth AVX2 Base64 encoding ramp-up with 12/24-byte tails

Description

Add 12-byte and 24-byte AVX2 encoding paths that reuse the existing AVX2 packing and alphabet-translation pipeline. This avoids falling back entirely to scalar encoding for inputs and tails that are too short for the regular AVX2 loop, substantially reducing the performance cliffs at 28-byte intervals.

This is very similar to what is already being done in the tail handling of src/icelake/icelake_base64.inl.cpp using AVX-512.

NEON also does this partially (src/arm64/arm_base64.cpp) but not fully - we could add a custom 12 byte tail handling there as well. I'll submit it later if I can find the time.

Type of change

  • Bug fix
  • Optimization
  • New feature
  • Refactor / cleanup
  • Documentation / tests
  • Other (please describe):

How to verify / test

  • Add additional tests to verify bugs or new features.
  • If you claim performance gains, you should provide benchmark numbers using high quality benchmarking code.

Checklist before submitting

  • I added/updated tests covering my change (if applicable)
  • Code builds locally and passes my check
  • Documentation / README updated if needed
  • Commits are atomic and messages are clear
  • I linked the related issue (if applicable)

Final notes

This PR address the performance cliff observed on small payloads as AVX2 improvements are degraded by tails until another AVX2 window can be used. The strategy is to run AVX2 against a partial vector. There is a point where the overhead of a partial vector balances the overhead of the tail - this is where it is interesting to switch.

The effect compounds with #1037 as shown below.

Bytes Master Master + half-AVX2 #1037+ half-AVX2
1 3.73 ns 4.38 ns (+17.3%) 3.86 ns (+3.5%)
2 4.05 ns 4.35 ns (+7.5%) 3.96 ns (-2.2%)
3 4.29 ns 4.59 ns (+7.1%) 3.85 ns (-10.3%)
4 4.96 ns 5.44 ns (+9.5%) 4.31 ns (-13.1%)
5 5.36 ns 5.50 ns (+2.7%) 4.76 ns (-11.2%)
6 5.37 ns 5.86 ns (+9.0%) 4.57 ns (-15.0%)
7 5.97 ns 6.41 ns (+7.3%) 4.97 ns (-16.7%)
8 6.47 ns 6.64 ns (+2.6%) 5.36 ns (-17.2%)
9 6.52 ns 6.81 ns (+4.5%) 5.24 ns (-19.7%)
10 7.16 ns 7.44 ns (+3.8%) 5.67 ns (-20.9%)
11 7.53 ns 7.76 ns (+3.1%) 5.93 ns (-21.2%)
12 7.84 ns 5.69 ns (-27.4%) 5.52 ns (-29.5%)
13 8.57 ns 5.93 ns (-30.9%) 5.78 ns (-32.6%)
14 8.73 ns 6.27 ns (-28.3%) 6.16 ns (-29.5%)
15 8.74 ns 6.69 ns (-23.4%) 6.02 ns (-31.1%)
16 9.43 ns 7.07 ns (-25.0%) 6.29 ns (-33.3%)
17 9.92 ns 7.70 ns (-22.4%) 6.37 ns (-35.8%)
18 10.16 ns 7.80 ns (-23.2%) 6.33 ns (-37.7%)
19 10.65 ns 8.28 ns (-22.3%) 6.59 ns (-38.1%)
20 10.89 ns 8.66 ns (-20.5%) 7.20 ns (-33.9%)
21 11.22 ns 9.08 ns (-19.1%) 6.96 ns (-37.9%)
22 11.39 ns 9.72 ns (-14.6%) 7.57 ns (-33.5%)
23 12.13 ns 10.15 ns (-16.3%) 7.70 ns (-36.5%)
24 12.26 ns 5.48 ns (-55.3%) 5.40 ns (-55.9%)
25 12.73 ns 6.08 ns (-52.3%) 5.83 ns (-54.2%)
26 12.81 ns 6.45 ns (-49.6%) 6.10 ns (-52.4%)
27 13.31 ns 6.85 ns (-48.5%) 6.12 ns (-54.0%)
28 6.58 ns 6.46 ns (-1.9%) 5.57 ns (-15.5%)

Notice the gap in performance on byte 28 in master - this is when we introduce the first AVX2 block processing - and then we have another ramp up:

Bytes Master Master + half-AVX2 #1037+ half-AVX2
29 6.71 ns 7.04 ns (+5.0%) 6.01 ns (-10.4%)
30 7.07 ns 7.02 ns (-0.7%) 5.76 ns (-18.5%)
31 7.49 ns 7.72 ns (+3.0%) 6.15 ns (-17.8%)
32 7.82 ns 7.83 ns (+0.2%) 6.30 ns (-19.5%)
33 8.06 ns 8.05 ns (-0.1%) 6.50 ns (-19.3%)
34 8.54 ns 8.59 ns (+0.6%) 6.68 ns (-21.7%)
35 9.10 ns 8.95 ns (-1.7%) 7.06 ns (-22.5%)
36 9.00 ns 6.88 ns (-23.5%) 6.78 ns (-24.7%)
37 9.90 ns 7.35 ns (-25.7%) 7.06 ns (-28.7%)
38 10.11 ns 7.81 ns (-22.8%) 7.56 ns (-25.2%)
39 10.32 ns 8.01 ns (-22.4%) 7.21 ns (-30.1%)
40 11.02 ns 8.64 ns (-21.6%) 7.84 ns (-28.8%)
41 11.31 ns 8.95 ns (-20.9%) 8.06 ns (-28.8%)
42 11.55 ns 9.06 ns (-21.6%) 8.28 ns (-28.4%)
43 12.25 ns 9.62 ns (-21.5%) 8.41 ns (-31.4%)
44 12.43 ns 10.00 ns (-19.5%) 8.43 ns (-32.2%)
45 12.35 ns 10.03 ns (-18.8%) 8.43 ns (-31.7%)
46 13.17 ns 10.54 ns (-20.0%) 8.61 ns (-34.6%)
47 13.29 ns 11.24 ns (-15.4%) 9.07 ns (-31.8%)
48 13.65 ns 6.87 ns (-49.6%) 6.88 ns (-49.6%)
49 13.99 ns 7.10 ns (-49.2%) 7.18 ns (-48.6%)
50 14.51 ns 7.55 ns (-48.0%) 7.27 ns (-49.9%)
51 14.58 ns 7.73 ns (-47.0%) 7.38 ns (-49.4%)
52 7.00 ns 7.17 ns (+2.5%) 6.28 ns (-10.3%)

Here we have another performance cliff, followed by another ramp up:

Bytes Master Master + half-AVX2 #1037+ half-AVX2
53 7.60 ns 7.45 ns (-2.0%) 6.59 ns (-13.4%)
54 7.59 ns 7.60 ns (+0.2%) 6.56 ns (-13.6%)
55 8.14 ns 8.32 ns (+2.2%) 6.89 ns (-15.4%)
56 8.52 ns 8.63 ns (+1.3%) 7.14 ns (-16.3%)
57 8.89 ns 8.66 ns (-2.5%) 7.19 ns (-19.1%)
58 9.45 ns 9.21 ns (-2.6%) 7.55 ns (-20.1%)
59 9.46 ns 9.75 ns (+3.1%) 7.74 ns (-18.2%)
60 9.87 ns 7.54 ns (-23.6%) 7.53 ns (-23.7%)
61 10.35 ns 8.36 ns (-19.3%) 8.45 ns (-18.3%)
62 10.58 ns 8.65 ns (-18.2%) 8.29 ns (-21.7%)
63 11.12 ns 8.69 ns (-21.8%) 8.33 ns (-25.1%)
64 11.71 ns 9.46 ns (-19.3%) 8.59 ns (-26.6%)
65 11.79 ns 9.85 ns (-16.4%) 8.95 ns (-24.0%)
66 11.79 ns 10.18 ns (-13.6%) 8.83 ns (-25.0%)
67 12.64 ns 10.53 ns (-16.7%) 9.31 ns (-26.4%)
68 13.01 ns 11.19 ns (-14.0%) 9.50 ns (-26.9%)
69 13.27 ns 11.22 ns (-15.4%) 9.51 ns (-28.4%)
70 14.09 ns 11.55 ns (-18.0%) 9.73 ns (-31.0%)
71 14.17 ns 12.15 ns (-14.3%) 9.95 ns (-29.8%)
72 14.65 ns 7.54 ns (-48.5%) 7.57 ns (-48.3%)
73 14.76 ns 8.33 ns (-43.6%) 7.87 ns (-46.7%)
74 15.52 ns 8.52 ns (-45.1%) 8.30 ns (-46.6%)
75 15.55 ns 8.75 ns (-43.8%) 8.12 ns (-47.8%)
76 7.64 ns 8.07 ns (+5.6%) 7.02 ns (-8.1%)

And another here.

Bytes Master Master + half-AVX2 #1037+ half-AVX2
77 8.60 ns 8.56 ns (-0.4%) 7.35 ns (-14.6%)
78 8.54 ns 8.65 ns (+1.3%) 7.19 ns (-15.9%)
79 9.41 ns 9.10 ns (-3.4%) 7.54 ns (-19.9%)
80 9.60 ns 9.70 ns (+1.0%) 7.68 ns (-20.0%)
81 9.75 ns 9.80 ns (+0.6%) 7.80 ns (-20.0%)
82 10.35 ns 10.34 ns (-0.1%) 8.29 ns (-19.9%)
83 10.14 ns 10.66 ns (+5.1%) 8.52 ns (-16.0%)
84 10.77 ns 8.17 ns (-24.1%) 8.51 ns (-21.0%)
85 11.24 ns 9.19 ns (-18.2%) 8.84 ns (-21.3%)
86 11.56 ns 8.95 ns (-22.6%) 8.89 ns (-23.1%)
87 11.67 ns 9.29 ns (-20.4%) 9.12 ns (-21.9%)
88 12.56 ns 9.96 ns (-20.7%) 9.22 ns (-26.6%)
89 12.45 ns 10.49 ns (-15.8%) 9.50 ns (-23.7%)
90 12.96 ns 10.41 ns (-19.6%) 9.50 ns (-26.7%)
91 13.82 ns 11.21 ns (-18.8%) 9.89 ns (-28.4%)
92 13.95 ns 11.81 ns (-15.3%) 9.97 ns (-28.5%)
93 14.06 ns 11.70 ns (-16.8%) 10.04 ns (-28.6%)
94 14.33 ns 12.53 ns (-12.6%) 10.11 ns (-29.4%)
95 14.30 ns 12.47 ns (-12.8%) 10.45 ns (-26.9%)
96 14.99 ns 8.13 ns (-45.8%) 8.17 ns (-45.5%)

So the approach significantly reduces the cliffs. It can be further improved with SSE (i.e. transitioning from scalar -> partial SSE -> full SSE -> partial AVX2 -> full AVX2 -> full AVX2 + scalar -> ...

going to partial AVX 512 would also be interesting

But adding SSE substeps would have increased the complexity here and the gains are not as important. If interested, I am happy to provide it as another PR.

@gaspardpetit gaspardpetit changed the title Use half avx2 to smooth the performance cliffs on small payloads base64 AVX2 - Use half avx2 to smooth the performance cliffs on small payloads Sep 19, 2026
@gaspardpetit
gaspardpetit marked this pull request as draft September 19, 2026 05:36
@gaspardpetit gaspardpetit changed the title base64 AVX2 - Use half avx2 to smooth the performance cliffs on small payloads base64 AVX2 - Use half avx2 to smooth the performance cliffs on small payloads when encoding Sep 19, 2026
@gaspardpetit
gaspardpetit marked this pull request as ready for review September 19, 2026 05:54

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant