diff --git a/README.md b/README.md
index 90cd68f..9e3bcce 100644
--- a/README.md
+++ b/README.md
@@ -70,6 +70,7 @@ These optimizations are based on Firedancer's specialized Base58 algorithms and
- Precomputed multiplication tables replace expensive division operations
- Converts binary data to base 58^5 limbs, then to raw base58 digits
- Matrix multiplication approach processes 5 base58 digits simultaneously
+- The 32/64-byte encode and decode matrix kernels are SIMD-accelerated with `Vector256`/`Vector128` (widest available width, scalar fallback) — in addition to the vectorized leading-zero count
- Separate encode/decode tables for 32-byte and 64-byte fixed sizes
- Achieves ~2.5x speedup through table-based optimizations vs iterative division
@@ -81,43 +82,43 @@ These optimizations are based on Firedancer's specialized Base58 algorithms and
```
-BenchmarkDotNet v0.15.8, Windows 11 (10.0.26200.8246/25H2/2025Update/HudsonValley2)
+BenchmarkDotNet v0.15.8, Windows 11 (10.0.26200.8655/25H2/2025Update/HudsonValley2)
13th Gen Intel Core i7-13700KF 3.40GHz, 1 CPU, 24 logical and 16 physical cores
-.NET SDK 10.0.203
- [Host] : .NET 10.0.7 (10.0.7, 10.0.726.21808), X64 RyuJIT x86-64-v3
- DefaultJob : .NET 10.0.7 (10.0.7, 10.0.726.21808), X64 RyuJIT x86-64-v3
+.NET SDK 11.0.100-preview.3.26207.106
+ [Host] : .NET 10.0.10 (10.0.10, 10.0.1026.32716), X64 RyuJIT x86-64-v3
+ DefaultJob : .NET 10.0.10 (10.0.10, 10.0.1026.32716), X64 RyuJIT x86-64-v3
Job=DefaultJob
```
-| Method | VectorType | Mean | Ratio | Gen0 | Allocated | Alloc Ratio |
-|--------------------------- |--------------- |------------:|------:|-------:|----------:|------------:|
-| **'Our Base58 Encode'** | **BitcoinAddress** | **537.17 ns** | **1.00** | **0.0057** | **96 B** | **1.00** |
-| 'SimpleBase Base58 Encode' | BitcoinAddress | 776.69 ns | 1.45 | 0.0057 | 96 B | 1.00 |
-| 'Our Base58 Decode' | BitcoinAddress | 160.88 ns | 0.30 | 0.0033 | 56 B | 0.58 |
-| 'SimpleBase Base58 Decode' | BitcoinAddress | 353.19 ns | 0.66 | 0.0033 | 56 B | 0.58 |
-| | | | | | | |
-| **'Our Base58 Encode'** | **SolanaAddress** | **94.07 ns** | **1.00** | **0.0070** | **112 B** | **1.00** |
-| 'SimpleBase Base58 Encode' | SolanaAddress | 1,433.92 ns | 15.24 | 0.0057 | 112 B | 1.00 |
-| 'Our Base58 Decode' | SolanaAddress | 104.19 ns | 1.11 | 0.0035 | 56 B | 0.50 |
-| 'SimpleBase Base58 Decode' | SolanaAddress | 703.66 ns | 7.48 | 0.0029 | 56 B | 0.50 |
-| | | | | | | |
-| **'Our Base58 Encode'** | **SolanaTx** | **239.21 ns** | **1.00** | **0.0124** | **200 B** | **1.00** |
-| 'SimpleBase Base58 Encode' | SolanaTx | 7,166.10 ns | 29.96 | 0.0076 | 200 B | 1.00 |
-| 'Our Base58 Decode' | SolanaTx | 180.37 ns | 0.75 | 0.0055 | 88 B | 0.44 |
-| 'SimpleBase Base58 Decode' | SolanaTx | 2,957.77 ns | 12.36 | 0.0038 | 88 B | 0.44 |
-| | | | | | | |
-| **'Our Base58 Encode'** | **IPFSHash** | **1,084.69 ns** | **1.00** | **0.0076** | **120 B** | **1.00** |
-| 'SimpleBase Base58 Encode' | IPFSHash | 1,617.11 ns | 1.49 | 0.0076 | 120 B | 1.00 |
-| 'Our Base58 Decode' | IPFSHash | 318.15 ns | 0.29 | 0.0038 | 64 B | 0.53 |
-| 'SimpleBase Base58 Decode' | IPFSHash | 854.47 ns | 0.79 | 0.0038 | 64 B | 0.53 |
-| | | | | | | |
-| **'Our Base58 Encode'** | **MoneroAddress** | **4,917.65 ns** | **1.00** | **0.0076** | **216 B** | **1.00** |
-| 'SimpleBase Base58 Encode' | MoneroAddress | 8,621.98 ns | 1.75 | - | 216 B | 1.00 |
-| 'Our Base58 Decode' | MoneroAddress | 1,198.92 ns | 0.24 | 0.0057 | 96 B | 0.44 |
-| 'SimpleBase Base58 Decode' | MoneroAddress | 3,844.43 ns | 0.78 | - | 96 B | 0.44 |
-
-
-## License
-
-This project is available under the MIT License.
+| Method | Categories | VectorType | Mean | Ratio | Gen0 | Allocated | Alloc Ratio |
+|--------------------------- |----------- |--------------- |------------:|------:|-------:|----------:|------------:|
+| **'Our Base58 Decode'** | **Decode** | **BitcoinAddress** | **178.71 ns** | **1.00** | **0.0033** | **56 B** | **1.00** |
+| 'SimpleBase Base58 Decode' | Decode | BitcoinAddress | 364.51 ns | 2.04 | 0.0033 | 56 B | 1.00 |
+| | | | | | | | |
+| **'Our Base58 Decode'** | **Decode** | **SolanaAddress** | **83.27 ns** | **1.00** | **0.0035** | **56 B** | **1.00** |
+| 'SimpleBase Base58 Decode' | Decode | SolanaAddress | 598.64 ns | 7.19 | 0.0029 | 56 B | 1.00 |
+| | | | | | | | |
+| **'Our Base58 Decode'** | **Decode** | **SolanaTx** | **167.00 ns** | **1.00** | **0.0055** | **88 B** | **1.00** |
+| 'SimpleBase Base58 Decode' | Decode | SolanaTx | 4,190.80 ns | 25.10 | 0.0038 | 88 B | 1.00 |
+| | | | | | | | |
+| **'Our Base58 Decode'** | **Decode** | **IPFSHash** | **339.26 ns** | **1.00** | **0.0038** | **64 B** | **1.00** |
+| 'SimpleBase Base58 Decode' | Decode | IPFSHash | 641.17 ns | 1.89 | 0.0038 | 64 B | 1.00 |
+| | | | | | | | |
+| **'Our Base58 Decode'** | **Decode** | **MoneroAddress** | **1,391.11 ns** | **1.00** | **0.0057** | **96 B** | **1.00** |
+| 'SimpleBase Base58 Decode' | Decode | MoneroAddress | 3,882.84 ns | 2.79 | - | 96 B | 1.00 |
+| | | | | | | | |
+| **'Our Base58 Encode'** | **Encode** | **BitcoinAddress** | **532.69 ns** | **1.00** | **0.0057** | **96 B** | **1.00** |
+| 'SimpleBase Base58 Encode' | Encode | BitcoinAddress | 774.83 ns | 1.45 | 0.0057 | 96 B | 1.00 |
+| | | | | | | | |
+| **'Our Base58 Encode'** | **Encode** | **SolanaAddress** | **95.15 ns** | **1.00** | **0.0070** | **112 B** | **1.00** |
+| 'SimpleBase Base58 Encode' | Encode | SolanaAddress | 1,521.61 ns | 15.99 | 0.0057 | 112 B | 1.00 |
+| | | | | | | | |
+| **'Our Base58 Encode'** | **Encode** | **SolanaTx** | **196.50 ns** | **1.00** | **0.0126** | **200 B** | **1.00** |
+| 'SimpleBase Base58 Encode' | Encode | SolanaTx | 7,338.14 ns | 37.34 | 0.0076 | 200 B | 1.00 |
+| | | | | | | | |
+| **'Our Base58 Encode'** | **Encode** | **IPFSHash** | **1,085.19 ns** | **1.00** | **0.0076** | **120 B** | **1.00** |
+| 'SimpleBase Base58 Encode' | Encode | IPFSHash | 1,690.70 ns | 1.56 | 0.0076 | 120 B | 1.00 |
+| | | | | | | | |
+| **'Our Base58 Encode'** | **Encode** | **MoneroAddress** | **4,959.86 ns** | **1.00** | **0.0076** | **216 B** | **1.00** |
+| 'SimpleBase Base58 Encode' | Encode | MoneroAddress | 8,734.32 ns | 1.76 | - | 216 B | 1.00 |
diff --git a/src/Base58Encoding.Benchmarks/Base58ComparisonBenchmark.cs b/src/Base58Encoding.Benchmarks/Base58ComparisonBenchmark.cs
index f5635ca..286fc7e 100644
--- a/src/Base58Encoding.Benchmarks/Base58ComparisonBenchmark.cs
+++ b/src/Base58Encoding.Benchmarks/Base58ComparisonBenchmark.cs
@@ -1,11 +1,14 @@
using Base58Encoding.Benchmarks.Common;
using BenchmarkDotNet.Attributes;
+using BenchmarkDotNet.Configs;
using BenchmarkDotNet.Diagnosers;
namespace Base58Encoding.Benchmarks;
[MemoryDiagnoser]
+[CategoriesColumn]
+[GroupBenchmarksBy(BenchmarkLogicalGroupRule.ByCategory)]
[HideColumns("Job", "Error", "StdDev", "Median", "RatioSD")]
public class Base58ComparisonBenchmark
{
@@ -28,25 +31,25 @@ public void Setup()
_base58Encoded = Base58.Bitcoin.Encode(_testData);
}
- [Benchmark(Description = "Our Base58 Encode", Baseline = true)]
+ [BenchmarkCategory("Encode"), Benchmark(Description = "Our Base58 Encode", Baseline = true)]
public string Encode_OurBase58()
{
return Base58.Bitcoin.Encode(_testData);
}
- [Benchmark(Description = "SimpleBase Base58 Encode")]
+ [BenchmarkCategory("Encode"), Benchmark(Description = "SimpleBase Base58 Encode")]
public string Encode_SimpleBase58()
{
return SimpleBase.Base58.Bitcoin.Encode(_testData);
}
- [Benchmark(Description = "Our Base58 Decode")]
+ [BenchmarkCategory("Decode"), Benchmark(Description = "Our Base58 Decode", Baseline = true)]
public byte[] Decode_OurBase58()
{
return Base58.Bitcoin.Decode(_base58Encoded);
}
- [Benchmark(Description = "SimpleBase Base58 Decode")]
+ [BenchmarkCategory("Decode"), Benchmark(Description = "SimpleBase Base58 Decode")]
public byte[] Decode_SimpleBase58()
{
return SimpleBase.Base58.Bitcoin.Decode(_base58Encoded);
diff --git a/src/Base58Encoding.Benchmarks/Base58Encoding.Benchmarks.csproj b/src/Base58Encoding.Benchmarks/Base58Encoding.Benchmarks.csproj
index 3d63021..de4f780 100644
--- a/src/Base58Encoding.Benchmarks/Base58Encoding.Benchmarks.csproj
+++ b/src/Base58Encoding.Benchmarks/Base58Encoding.Benchmarks.csproj
@@ -2,11 +2,15 @@
Exe
+
+ pdbonly
+ true
+
diff --git a/src/Base58Encoding.Benchmarks/EndToEndDecodeBenchmark.cs b/src/Base58Encoding.Benchmarks/EndToEndDecodeBenchmark.cs
new file mode 100644
index 0000000..86be123
--- /dev/null
+++ b/src/Base58Encoding.Benchmarks/EndToEndDecodeBenchmark.cs
@@ -0,0 +1,210 @@
+using System.Buffers.Binary;
+using System.Numerics;
+using System.Runtime.CompilerServices;
+using System.Runtime.InteropServices;
+using System.Runtime.Intrinsics;
+
+using BenchmarkDotNet.Attributes;
+using BenchmarkDotNet.Jobs;
+
+namespace Base58Encoding.Benchmarks;
+
+// End-to-end 64-byte decode A/B. Both arms run the identical full decode path (validate,
+// intermediate build, matmul, carry-reduce, leading-zero check, byte output) and differ ONLY
+// in the matmul step (SIMD TensorDot over transposed ulong columns vs jagged scalar). Measured
+// in a single run, so the ratio is the real end-to-end effect of the matmul change.
+[SimpleJob(RuntimeMoniker.Net10_0)]
+[HideColumns("Error", "StdDev", "RatioSD")]
+public class EndToEndDecodeBenchmark
+{
+ private string _sig64 = default!;
+ private ulong[] _transposed64 = default!;
+ private uint[][] _jagged64 = default!;
+
+ [GlobalSetup]
+ public void Setup()
+ {
+ var bytes = new byte[64];
+ Random.Shared.NextBytes(bytes);
+ _sig64 = Base58.Bitcoin.Encode(bytes);
+
+ // Production tables on this branch are transposed column-major ulong.
+ _transposed64 = Base58BitcoinTables.DecodeTable64;
+
+ // Rebuild the original jagged row-major layout: jagged[i][j] = transposed[j*IntermediateSz64 + i].
+ int rows = Base58BitcoinTables.IntermediateSz64;
+ int cols = Base58BitcoinTables.BinarySz64;
+ _jagged64 = new uint[rows][];
+ for (int i = 0; i < rows; i++)
+ {
+ _jagged64[i] = new uint[cols];
+ for (int j = 0; j < cols; j++)
+ {
+ _jagged64[i][j] = (uint)_transposed64[j * rows + i];
+ }
+ }
+
+ // Correctness: both matmul variants must reproduce the production decode exactly.
+ Span a = stackalloc byte[64];
+ Span b = stackalloc byte[64];
+ DecodeFull64(_sig64, a, useSimd: true);
+ DecodeFull64(_sig64, b, useSimd: false);
+ byte[] expected = Base58.Bitcoin.Decode(_sig64);
+ if (!a.SequenceEqual(expected) || !b.SequenceEqual(expected))
+ {
+ throw new InvalidOperationException("DecodeFull64 does not match production Decode");
+ }
+ }
+
+ [Benchmark(Baseline = true)]
+ public int Full_Jagged()
+ {
+ Span dest = stackalloc byte[64];
+ return DecodeFull64(_sig64, dest, useSimd: false);
+ }
+
+ [Benchmark]
+ public int Full_Simd()
+ {
+ Span dest = stackalloc byte[64];
+ return DecodeFull64(_sig64, dest, useSimd: true);
+ }
+
+ [SkipLocalsInit]
+ private int DecodeFull64(ReadOnlySpan encoded, Span destination, bool useSimd)
+ {
+ int charCount = encoded.Length;
+
+ Span rawBase58 = stackalloc byte[Base58BitcoinTables.Raw58Sz64];
+ ReadOnlySpan decodeTable = BitcoinAlphabet.DecodeTable;
+
+ int prepend0 = Base58BitcoinTables.Raw58Sz64 - charCount;
+ for (int j = 0; j < Base58BitcoinTables.Raw58Sz64; j++)
+ {
+ if (j < prepend0)
+ {
+ rawBase58[j] = 0;
+ }
+ else
+ {
+ int c = encoded[j - prepend0];
+ if ((uint)c >= 128 || decodeTable[c] == 255)
+ {
+ ThrowHelper.ThrowInvalidCharacter((char)c);
+ }
+
+ rawBase58[j] = decodeTable[c];
+ }
+ }
+
+ Span intermediate = stackalloc ulong[Base58BitcoinTables.IntermediateSz64];
+ for (int i = 0; i < Base58BitcoinTables.IntermediateSz64; i++)
+ {
+ intermediate[i] = (ulong)rawBase58[5 * i + 0] * 11316496UL +
+ (ulong)rawBase58[5 * i + 1] * 195112UL +
+ (ulong)rawBase58[5 * i + 2] * 3364UL +
+ (ulong)rawBase58[5 * i + 3] * 58UL +
+ (ulong)rawBase58[5 * i + 4] * 1UL;
+ }
+
+ Span binary = stackalloc ulong[Base58BitcoinTables.BinarySz64];
+ if (useSimd)
+ {
+ var table = _transposed64;
+ int rows = Base58BitcoinTables.IntermediateSz64;
+ for (int j = 0; j < Base58BitcoinTables.BinarySz64; j++)
+ {
+ binary[j] = TensorDot(intermediate, table.AsSpan(j * rows, rows));
+ }
+ }
+ else
+ {
+ var table = _jagged64;
+ for (int j = 0; j < Base58BitcoinTables.BinarySz64; j++)
+ {
+ ulong acc = 0UL;
+ for (int i = 0; i < Base58BitcoinTables.IntermediateSz64; i++)
+ {
+ acc += intermediate[i] * table[i][j];
+ }
+
+ binary[j] = acc;
+ }
+ }
+
+ for (int i = Base58BitcoinTables.BinarySz64 - 1; i > 0; i--)
+ {
+ binary[i - 1] += binary[i] >> 32;
+ binary[i] &= 0xFFFFFFFFUL;
+ }
+
+ if (binary[0] > 0xFFFFFFFFUL)
+ {
+ return -1;
+ }
+
+ int outputLeadingZeros = 0;
+ for (int i = 0; i < Base58BitcoinTables.BinarySz64; i++)
+ {
+ uint v = (uint)binary[i];
+ if (v != 0)
+ {
+ outputLeadingZeros += BitOperations.LeadingZeroCount(v) / 8;
+ break;
+ }
+ outputLeadingZeros += 4;
+ }
+
+ int inputLeadingOnes = Base58.CountLeadingCharacters(encoded, '1');
+ if (outputLeadingZeros != inputLeadingOnes)
+ {
+ return -1;
+ }
+
+ for (int i = 0; i < Base58BitcoinTables.BinarySz64; i++)
+ {
+ BinaryPrimitives.WriteUInt32BigEndian(destination.Slice(i * sizeof(uint), sizeof(uint)), (uint)binary[i]);
+ }
+
+ return 64;
+ }
+
+ private static ulong TensorDot(ReadOnlySpan x, ReadOnlySpan y)
+ {
+ ref ulong xr = ref MemoryMarshal.GetReference(x);
+ ref ulong yr = ref MemoryMarshal.GetReference(y);
+ int len = x.Length;
+ int i = 0;
+ ulong sum = 0UL;
+
+ if (Vector256.IsHardwareAccelerated && len >= Vector256.Count)
+ {
+ Vector256 acc = Vector256.Zero;
+ int upper = len - Vector256.Count;
+ for (; i <= upper; i += Vector256.Count)
+ {
+ acc += Vector256.LoadUnsafe(ref xr, (nuint)i) * Vector256.LoadUnsafe(ref yr, (nuint)i);
+ }
+
+ sum += Vector256.Sum(acc);
+ }
+ else if (Vector128.IsHardwareAccelerated && len >= Vector128.Count)
+ {
+ Vector128 acc = Vector128.Zero;
+ int upper = len - Vector128.Count;
+ for (; i <= upper; i += Vector128.Count)
+ {
+ acc += Vector128.LoadUnsafe(ref xr, (nuint)i) * Vector128.LoadUnsafe(ref yr, (nuint)i);
+ }
+
+ sum += Vector128.Sum(acc);
+ }
+
+ for (; i < len; i++)
+ {
+ sum += Unsafe.Add(ref xr, i) * Unsafe.Add(ref yr, i);
+ }
+
+ return sum;
+ }
+}
diff --git a/src/Base58Encoding.Benchmarks/EndToEndEncodeBenchmark.cs b/src/Base58Encoding.Benchmarks/EndToEndEncodeBenchmark.cs
new file mode 100644
index 0000000..2f56752
--- /dev/null
+++ b/src/Base58Encoding.Benchmarks/EndToEndEncodeBenchmark.cs
@@ -0,0 +1,296 @@
+using System.Buffers.Binary;
+using System.Numerics.Tensors;
+using System.Runtime.CompilerServices;
+using System.Runtime.InteropServices;
+using System.Runtime.Intrinsics;
+using System.Text;
+
+using BenchmarkDotNet.Attributes;
+using BenchmarkDotNet.Jobs;
+
+namespace Base58Encoding.Benchmarks;
+
+// End-to-end 32/64-byte encode. Each arm runs the full encode path and differs only in the matmul
+// multiply-add step (intermediate[1..] += binary[i] * EncodeTableRow[i]): hand-written SIMD (baseline) vs
+// TensorPrimitives.MultiplyAdd, with the production scalar loop available as a commented reference.
+// TMode is a struct marker, so the typeof checks fold at JIT time and each arm is fully specialized.
+// Disasm: dotnet run -c Release -- --filter "*EndToEndEncode*" --disasm --disasmDepth 3
+[MemoryDiagnoser]
+[SimpleJob(RuntimeMoniker.Net10_0)]
+[HideColumns("RatioSD")]
+public class EndToEndEncodeBenchmark
+{
+ // Empty value-type markers: one per matmul strategy. Constrained to `struct` so the JIT
+ // specializes each generic instantiation and elides the typeof comparisons.
+ private interface IMatmulMode { }
+ private readonly struct ScalarMode : IMatmulMode { }
+ private readonly struct SimdMode : IMatmulMode { }
+ private readonly struct TensorMode : IMatmulMode { }
+
+ [Params(32, 64)]
+ public int Size { get; set; }
+
+ private byte[] _data = default!;
+ private uint[][] _jagged = default!;
+ private ulong[] _rowMajor = default!;
+
+ [GlobalSetup]
+ public void Setup()
+ {
+ var rng = new Random(42);
+ _data = new byte[Size];
+ rng.NextBytes(_data);
+
+ // Scalar arm reads the production jagged tables directly; the vectorized arms read each
+ // encode row laid out contiguously as ulong. Row length is IntermediateSz - 1 (the matmul
+ // only touches columns 0..IntermediateSz-2, accumulating into intermediate[1..]).
+ _jagged = Size == 32 ? Base58BitcoinTables.EncodeTable32 : Base58BitcoinTables.EncodeTable64;
+ int binarySz = Size == 32 ? Base58BitcoinTables.BinarySz32 : Base58BitcoinTables.BinarySz64;
+ int rowLen = (Size == 32 ? Base58BitcoinTables.IntermediateSz32 : Base58BitcoinTables.IntermediateSz64) - 1;
+
+ _rowMajor = new ulong[binarySz * rowLen];
+ for (int i = 0; i < binarySz; i++)
+ {
+ for (int j = 0; j < rowLen; j++)
+ {
+ _rowMajor[i * rowLen + j] = _jagged[i][j];
+ }
+ }
+
+ // Correctness: every arm must reproduce the production encode exactly.
+ byte[] expected = Encoding.ASCII.GetBytes(Base58.Bitcoin.Encode(_data));
+ AssertMatches(expected);
+ AssertMatches(expected);
+ AssertMatches(expected);
+ }
+
+ private void AssertMatches(byte[] expected)
+ where TMode : struct, IMatmulMode
+ {
+ Span scratch = stackalloc byte[128];
+ int len = EncodeFull(_data, scratch);
+ if (!scratch[..len].SequenceEqual(expected))
+ {
+ throw new InvalidOperationException($"EncodeFull<{typeof(TMode).Name}> does not match production Encode for size {Size}");
+ }
+ }
+
+ // Reference arm: the current production scalar matmul over the jagged uint[][]. Left commented
+ // so runs compare Simd vs Tensor only. Uncomment to include it; Simd stays the baseline (BDN
+ // allows exactly one Baseline).
+ //[Benchmark]
+ //public int Scalar()
+ //{
+ // Span dest = stackalloc byte[128];
+ // return EncodeFull(_data, dest);
+ //}
+
+ [Benchmark(Baseline = true)]
+ public int Simd()
+ {
+ Span dest = stackalloc byte[128];
+ return EncodeFull(_data, dest);
+ }
+
+ [Benchmark]
+ public int Tensor()
+ {
+ Span dest = stackalloc byte[128];
+ return EncodeFull(_data, dest);
+ }
+
+ private int EncodeFull(ReadOnlySpan data, Span destination)
+ where TMode : struct, IMatmulMode
+ {
+ return data.Length == 32
+ ? EncodeFull32(data, destination)
+ : EncodeFull64(data, destination);
+ }
+
+ [SkipLocalsInit]
+ private int EncodeFull32(ReadOnlySpan data, Span destination)
+ where TMode : struct, IMatmulMode
+ {
+ int inLeadingZeros = Base58.CountLeadingZeros(data);
+
+ Span binary = stackalloc uint[Base58BitcoinTables.BinarySz32];
+ for (int i = 0; i < Base58BitcoinTables.BinarySz32; i++)
+ {
+ binary[i] = BinaryPrimitives.ReadUInt32BigEndian(data.Slice(i * sizeof(uint), sizeof(uint)));
+ }
+
+ Span intermediate = stackalloc ulong[Base58BitcoinTables.IntermediateSz32];
+ intermediate.Clear();
+
+ int rowLen = Base58BitcoinTables.IntermediateSz32 - 1;
+ AccumulateRows(intermediate, binary, 0, Base58BitcoinTables.BinarySz32, rowLen);
+
+ Span rawBase58 = stackalloc byte[Base58BitcoinTables.Raw58Sz32];
+ return ReduceExtractEmit(intermediate, Base58BitcoinTables.IntermediateSz32, rawBase58, Base58BitcoinTables.Raw58Sz32, inLeadingZeros, destination);
+ }
+
+ [SkipLocalsInit]
+ private int EncodeFull64(ReadOnlySpan data, Span destination)
+ where TMode : struct, IMatmulMode
+ {
+ int inLeadingZeros = Base58.CountLeadingZeros(data);
+
+ Span binary = stackalloc uint[Base58BitcoinTables.BinarySz64];
+ for (int i = 0; i < Base58BitcoinTables.BinarySz64; i++)
+ {
+ binary[i] = BinaryPrimitives.ReadUInt32BigEndian(data.Slice(i * sizeof(uint), sizeof(uint)));
+ }
+
+ Span intermediate = stackalloc ulong[Base58BitcoinTables.IntermediateSz64];
+ intermediate.Clear();
+
+ int rowLen = Base58BitcoinTables.IntermediateSz64 - 1;
+
+ // Split matmul with an interleaved mini-reduction to keep the limbs from overflowing,
+ // matching production ComputeBitcoin64FastRaw exactly.
+ AccumulateRows(intermediate, binary, 0, 8, rowLen);
+ intermediate[15] += intermediate[16] / Base58BitcoinTables.R1Div;
+ intermediate[16] %= Base58BitcoinTables.R1Div;
+ AccumulateRows(intermediate, binary, 8, Base58BitcoinTables.BinarySz64, rowLen);
+
+ Span rawBase58 = stackalloc byte[Base58BitcoinTables.Raw58Sz64];
+ return ReduceExtractEmit(intermediate, Base58BitcoinTables.IntermediateSz64, rawBase58, Base58BitcoinTables.Raw58Sz64, inLeadingZeros, destination);
+ }
+
+ // The only step that differs between arms: intermediate[1..] += binary[i] * row[i] for each
+ // row. TMode is a value type, so every `typeof(TMode) == typeof(...)` below is a compile-time
+ // constant and only the matching block survives in each specialized instantiation.
+ private void AccumulateRows(
+ Span intermediate,
+ ReadOnlySpan binary,
+ int startRow,
+ int endRow,
+ int rowLen)
+ where TMode : struct, IMatmulMode
+ {
+ if (typeof(TMode) == typeof(ScalarMode))
+ {
+ uint[][] jagged = _jagged;
+ for (int i = startRow; i < endRow; i++)
+ {
+ uint[] row = jagged[i];
+ for (int j = 0; j < rowLen; j++)
+ {
+ intermediate[j + 1] += (ulong)binary[i] * row[j];
+ }
+ }
+
+ return;
+ }
+
+ ulong[] rowMajor = _rowMajor;
+ Span target = intermediate.Slice(1, rowLen);
+
+ if (typeof(TMode) == typeof(TensorMode))
+ {
+ for (int i = startRow; i < endRow; i++)
+ {
+ ReadOnlySpan row = rowMajor.AsSpan(i * rowLen, rowLen);
+ // destination[k] = (row[k] * binary[i]) + target[k]; addend and destination are the
+ // same span (allowed: they begin at the same location).
+ TensorPrimitives.MultiplyAdd(row, (ulong)binary[i], target, target);
+ }
+
+ return;
+ }
+
+ // SimdMode
+ for (int i = startRow; i < endRow; i++)
+ {
+ ReadOnlySpan row = rowMajor.AsSpan(i * rowLen, rowLen);
+ MultiplyAddSimd(row, binary[i], target);
+ }
+ }
+
+ // acc[k] += x[k] * scale, widest available width first, then a scalar tail / fallback.
+ private static void MultiplyAddSimd(ReadOnlySpan x, ulong scale, Span acc)
+ {
+ ref ulong xr = ref MemoryMarshal.GetReference(x);
+ ref ulong ar = ref MemoryMarshal.GetReference(acc);
+ int len = x.Length;
+ int i = 0;
+
+ if (Vector256.IsHardwareAccelerated && len >= Vector256.Count)
+ {
+ Vector256 s = Vector256.Create(scale);
+ int upper = len - Vector256.Count;
+ for (; i <= upper; i += Vector256.Count)
+ {
+ Vector256 a = Vector256.LoadUnsafe(ref ar, (nuint)i);
+ Vector256 v = Vector256.LoadUnsafe(ref xr, (nuint)i);
+ (a + (v * s)).StoreUnsafe(ref ar, (nuint)i);
+ }
+ }
+ else if (Vector128.IsHardwareAccelerated && len >= Vector128.Count)
+ {
+ Vector128 s = Vector128.Create(scale);
+ int upper = len - Vector128.Count;
+ for (; i <= upper; i += Vector128.Count)
+ {
+ Vector128 a = Vector128.LoadUnsafe(ref ar, (nuint)i);
+ Vector128 v = Vector128.LoadUnsafe(ref xr, (nuint)i);
+ (a + (v * s)).StoreUnsafe(ref ar, (nuint)i);
+ }
+ }
+
+ for (; i < len; i++)
+ {
+ Unsafe.Add(ref ar, i) += Unsafe.Add(ref xr, i) * scale;
+ }
+ }
+
+ private static int ReduceExtractEmit(
+ Span intermediate,
+ int intermediateSz,
+ Span rawBase58,
+ int raw58Sz,
+ int inLeadingZeros,
+ Span destination)
+ {
+ // Reduce each limb to less than 58^5.
+ for (int i = intermediateSz - 1; i > 0; i--)
+ {
+ intermediate[i - 1] += intermediate[i] / Base58BitcoinTables.R1Div;
+ intermediate[i] %= Base58BitcoinTables.R1Div;
+ }
+
+ // Convert each limb to five base58 digits.
+ for (int i = 0; i < intermediateSz; i++)
+ {
+ uint v = (uint)intermediate[i];
+ rawBase58[5 * i + 4] = (byte)((v / 1U) % 58U);
+ rawBase58[5 * i + 3] = (byte)((v / 58U) % 58U);
+ rawBase58[5 * i + 2] = (byte)((v / 3364U) % 58U);
+ rawBase58[5 * i + 1] = (byte)((v / 195112U) % 58U);
+ rawBase58[5 * i + 0] = (byte)(v / 11316496U);
+ }
+
+ int rawLeadingZeros = 0;
+ while (rawLeadingZeros < raw58Sz && rawBase58[rawLeadingZeros] == 0)
+ {
+ rawLeadingZeros++;
+ }
+
+ int digitCount = raw58Sz - rawLeadingZeros;
+ int outputLength = inLeadingZeros + digitCount;
+
+ if (inLeadingZeros > 0)
+ {
+ destination[..inLeadingZeros].Fill((byte)'1');
+ }
+
+ int index = inLeadingZeros;
+ ReadOnlySpan alphabet = BitcoinAlphabet.Characters;
+ for (int k = rawLeadingZeros; k < raw58Sz; k++)
+ {
+ destination[index++] = alphabet[rawBase58[k]];
+ }
+
+ return outputLength;
+ }
+}
diff --git a/src/Base58Encoding.Benchmarks/JaggedVsMultidimensionalArrayBenchmark.cs b/src/Base58Encoding.Benchmarks/JaggedVsMultidimensionalArrayBenchmark.cs
index f1725b4..d8ed15a 100644
--- a/src/Base58Encoding.Benchmarks/JaggedVsMultidimensionalArrayBenchmark.cs
+++ b/src/Base58Encoding.Benchmarks/JaggedVsMultidimensionalArrayBenchmark.cs
@@ -17,11 +17,11 @@ public class JaggedVsMultidimensionalArrayBenchmark
// Jagged arrays (current implementation)
private static readonly uint[][] JaggedEncodeTable32 = Base58BitcoinTables.EncodeTable32;
- private static readonly uint[][] JaggedDecodeTable32 = Base58BitcoinTables.DecodeTable32;
+ private static readonly uint[][] JaggedDecodeTable32 = UntransposeToJagged(Base58BitcoinTables.DecodeTable32, Base58BitcoinTables.IntermediateSz32, Base58BitcoinTables.BinarySz32);
// Multidimensional arrays (alternative implementation)
private static readonly uint[,] MultidimensionalEncodeTable32 = ConvertToMultidimensional(Base58BitcoinTables.EncodeTable32);
- private static readonly uint[,] MultidimensionalDecodeTable32 = ConvertToMultidimensional(Base58BitcoinTables.DecodeTable32);
+ private static readonly uint[,] MultidimensionalDecodeTable32 = ConvertToMultidimensional(JaggedDecodeTable32);
private readonly ref struct FastEncodeState
{
@@ -71,6 +71,23 @@ public byte[] DecodeWithMultidimensionalArray()
return DecodeBitcoin32FastMultidimensional(_encodedBase58)!;
}
+ // Production DecodeTable32 is now transposed column-major ulong; rebuild the original
+ // row-major jagged form so this layout comparison keeps compiling.
+ private static uint[][] UntransposeToJagged(ulong[] transposed, int rows, int cols)
+ {
+ var jagged = new uint[rows][];
+ for (int i = 0; i < rows; i++)
+ {
+ jagged[i] = new uint[cols];
+ for (int j = 0; j < cols; j++)
+ {
+ jagged[i][j] = (uint)transposed[j * rows + i];
+ }
+ }
+
+ return jagged;
+ }
+
private static uint[,] ConvertToMultidimensional(uint[][] jaggedArray)
{
int rows = jaggedArray.Length;
diff --git a/src/Base58Encoding.Benchmarks/Program.cs b/src/Base58Encoding.Benchmarks/Program.cs
index a534ab4..66a87ff 100644
--- a/src/Base58Encoding.Benchmarks/Program.cs
+++ b/src/Base58Encoding.Benchmarks/Program.cs
@@ -2,6 +2,4 @@
using BenchmarkDotNet.Running;
-BenchmarkRunner.Run();
-//BenchmarkRunner.Run();
-//BenchmarkRunner.Run();
+BenchmarkSwitcher.FromAssembly(typeof(EndToEndEncodeBenchmark).Assembly).Run(args);
diff --git a/src/Base58Encoding.Tests/SimpleBaseFuzzTests.cs b/src/Base58Encoding.Tests/SimpleBaseFuzzTests.cs
new file mode 100644
index 0000000..77b0e8e
--- /dev/null
+++ b/src/Base58Encoding.Tests/SimpleBaseFuzzTests.cs
@@ -0,0 +1,199 @@
+using System.Diagnostics;
+using System.Numerics;
+using System.Text;
+
+namespace Base58Encoding.Tests;
+
+// Long-running differential fuzz. Explicit so it never runs in the normal suite. Two tests: a
+// general one over all input lengths, and one focused on the Bitcoin 32/64-byte fast paths. Run:
+// dotnet run --project src/Base58Encoding.Tests -c Release -- -explicit only
+// Override the duration (seconds) with the FUZZ_SECONDS environment variable. Add a -method filter
+// to run just one (e.g. -method "*Bitcoin_32And64*").
+//
+// Ground truth is a BigInteger oracle (the literal definition of Base58), so the fuzz validates our
+// code without trusting any third party. We also cross-check our ENCODER against SimpleBase's, but
+// intentionally do NOT assert SimpleBase.Decode(ours): SimpleBase's decoder (5.6.0 and 5.6.2) drops
+// the most-significant byte on some larger inputs (verified: our encoding matches the oracle and the
+// Python base58 library, both of which decode it correctly). Reported: ssg/SimpleBase#83
+// (https://github.com/ssg/SimpleBase/issues/83).
+public class SimpleBaseFuzzTests
+{
+ private const string Alphabet = "123456789ABCDEFGHJKLMNPQRSTUVWXYZabcdefghijkmnopqrstuvwxyz";
+
+ private readonly ITestOutputHelper _output;
+
+ public SimpleBaseFuzzTests(ITestOutputHelper output)
+ {
+ _output = output;
+ }
+
+ [Fact(Explicit = true)]
+ [Trait("Category", "Fuzz")]
+ public void EncodeDecode_MatchesOracleAndSimpleBase_UnderRandomInput()
+ {
+ int seconds = int.TryParse(Environment.GetEnvironmentVariable("FUZZ_SECONDS"), out int s) && s > 0 ? s : 600;
+ var rng = new Random(20260722);
+ var sw = Stopwatch.StartNew();
+ long iterations = 0;
+ int maxLen = 0;
+
+ while (sw.Elapsed.TotalSeconds < seconds)
+ {
+ byte[] data = NextInput(rng);
+ maxLen = Math.Max(maxLen, data.Length);
+
+ string ours = Base58.Bitcoin.Encode(data);
+
+ string oracle = OracleEncode(data);
+ if (!string.Equals(ours, oracle, StringComparison.Ordinal))
+ {
+ Assert.Fail($"Encode vs oracle mismatch @ iter {iterations}: input={Convert.ToHexString(data)}\n ours ={ours}\n oracle={oracle}");
+ }
+
+ string theirs = SimpleBase.Base58.Bitcoin.Encode(data);
+ if (!string.Equals(ours, theirs, StringComparison.Ordinal))
+ {
+ Assert.Fail($"Encode vs SimpleBase mismatch @ iter {iterations}: input={Convert.ToHexString(data)}\n ours ={ours}\n theirs={theirs}");
+ }
+
+ byte[] roundTrip = Base58.Bitcoin.Decode(ours);
+ if (!roundTrip.AsSpan().SequenceEqual(data))
+ {
+ Assert.Fail($"Round-trip mismatch @ iter {iterations}: input={Convert.ToHexString(data)} decoded={Convert.ToHexString(roundTrip)}");
+ }
+
+ byte[] oursFromTheirs = Base58.Bitcoin.Decode(theirs);
+ if (!oursFromTheirs.AsSpan().SequenceEqual(data))
+ {
+ Assert.Fail($"Cross-decode (ours <- SimpleBase) mismatch @ iter {iterations}: input={Convert.ToHexString(data)} decoded={Convert.ToHexString(oursFromTheirs)}");
+ }
+
+ iterations++;
+ }
+
+ _output.WriteLine($"Fuzz OK: {iterations:N0} iterations in {sw.Elapsed.TotalSeconds:F0}s, max input {maxLen} bytes, zero mismatches.");
+ }
+
+ // Focused fuzz on the Bitcoin 32- and 64-byte fast paths (TryDecodeBitcoin{32,64}Fast and the
+ // SIMD encode). Only 32/64-byte inputs; the MSB is kept non-zero most of the time so the encoding
+ // lands in the fast-path length window (43-44 / 87-88 chars) and Decode takes the fast path.
+ // Exercises the string and byte-span overloads of both Encode and Decode against the oracle.
+ // Run this one alone with: -explicit only -method "*Bitcoin_32And64*"
+ [Fact(Explicit = true)]
+ [Trait("Category", "Fuzz")]
+ public void Bitcoin_32And64_FastPaths_MatchOracle_UnderRandomInput()
+ {
+ int seconds = int.TryParse(Environment.GetEnvironmentVariable("FUZZ_SECONDS"), out int s) && s > 0 ? s : 600;
+ var rng = new Random(58585858);
+ var sw = Stopwatch.StartNew();
+ long iterations = 0;
+
+ Span encoded = stackalloc byte[128];
+ Span decoded = stackalloc byte[64];
+
+ while (sw.Elapsed.TotalSeconds < seconds)
+ {
+ int len = (iterations & 1) == 0 ? 32 : 64;
+ byte[] data = new byte[len];
+ rng.NextBytes(data);
+
+ if (rng.Next(8) == 0)
+ {
+ // Exercise the fast encode path's leading-zero ('1'-prefix) handling.
+ Array.Clear(data, 0, rng.Next(1, len + 1));
+ }
+ else if (data[0] == 0)
+ {
+ // Keep it in the fast-path length window so Decode hits TryDecodeBitcoin{32,64}Fast.
+ data[0] = 1;
+ }
+
+ string ours = Base58.Bitcoin.Encode(data);
+
+ string oracle = OracleEncode(data);
+ if (!string.Equals(ours, oracle, StringComparison.Ordinal))
+ {
+ Assert.Fail($"[{len}B] Encode vs oracle mismatch @ iter {iterations}: input={Convert.ToHexString(data)}\n ours ={ours}\n oracle={oracle}");
+ }
+
+ string theirs = SimpleBase.Base58.Bitcoin.Encode(data);
+ if (!string.Equals(ours, theirs, StringComparison.Ordinal))
+ {
+ Assert.Fail($"[{len}B] Encode vs SimpleBase mismatch @ iter {iterations}: input={Convert.ToHexString(data)}\n ours ={ours}\n theirs={theirs}");
+ }
+
+ // Byte-span encode overload must match the string encode.
+ int encodedLength = Base58.Bitcoin.Encode(data, encoded);
+ if (!encoded[..encodedLength].SequenceEqual(Encoding.ASCII.GetBytes(ours)))
+ {
+ Assert.Fail($"[{len}B] Encode-to-bytes mismatch @ iter {iterations}: input={Convert.ToHexString(data)}");
+ }
+
+ // Decode from chars (round-trip through the fast decode path).
+ byte[] roundTrip = Base58.Bitcoin.Decode(ours);
+ if (!roundTrip.AsSpan().SequenceEqual(data))
+ {
+ Assert.Fail($"[{len}B] Decode(string) round-trip mismatch @ iter {iterations}: input={Convert.ToHexString(data)} decoded={Convert.ToHexString(roundTrip)}");
+ }
+
+ // Decode from ASCII bytes into a destination span (byte-input fast path).
+ int decodedLength = Base58.Bitcoin.Decode(encoded[..encodedLength], decoded);
+ if (!decoded[..decodedLength].SequenceEqual(data))
+ {
+ Assert.Fail($"[{len}B] Decode(bytes) round-trip mismatch @ iter {iterations}: input={Convert.ToHexString(data)} decoded={Convert.ToHexString(decoded[..decodedLength].ToArray())}");
+ }
+
+ iterations++;
+ }
+
+ _output.WriteLine($"Bitcoin 32/64 fast-path fuzz OK: {iterations:N0} iterations in {sw.Elapsed.TotalSeconds:F0}s, zero mismatches.");
+ }
+
+ // Length distribution biased toward the 32/64-byte fast paths and short inputs; 1-in-4 inputs
+ // get a run of leading zero bytes to exercise the '1'-prefix / leading-zero handling.
+ private static byte[] NextInput(Random rng)
+ {
+ int len = rng.Next(100) switch
+ {
+ < 15 => 32,
+ < 30 => 64,
+ < 45 => rng.Next(1, 8),
+ _ => rng.Next(1, 200),
+ };
+
+ byte[] data = new byte[len];
+ rng.NextBytes(data);
+
+ if (rng.Next(4) == 0)
+ {
+ Array.Clear(data, 0, rng.Next(1, len + 1));
+ }
+
+ return data;
+ }
+
+ // The literal definition of Base58: big-endian integer, repeated divide-by-58, one '1' per leading zero.
+ private static string OracleEncode(byte[] data)
+ {
+ int zeros = 0;
+ while (zeros < data.Length && data[zeros] == 0)
+ {
+ zeros++;
+ }
+
+ var num = new BigInteger(data, isUnsigned: true, isBigEndian: true);
+ var sb = new StringBuilder();
+ while (num > 0)
+ {
+ num = BigInteger.DivRem(num, 58, out BigInteger rem);
+ sb.Insert(0, Alphabet[(int)rem]);
+ }
+
+ for (int i = 0; i < zeros; i++)
+ {
+ sb.Insert(0, '1');
+ }
+
+ return sb.ToString();
+ }
+}
diff --git a/src/Base58Encoding.Tests/VectorInstructionSetTests.cs b/src/Base58Encoding.Tests/VectorInstructionSetTests.cs
new file mode 100644
index 0000000..f34c241
--- /dev/null
+++ b/src/Base58Encoding.Tests/VectorInstructionSetTests.cs
@@ -0,0 +1,84 @@
+using System.Diagnostics;
+
+namespace Base58Encoding.Tests;
+
+// The 32/64-byte fast paths (and CountLeadingZeros) select a code path from the CPU's available
+// vector instruction sets:
+// AVX2 present -> Vector256 branch
+// AVX2 off, SSE present -> Vector128 branch
+// no hardware intrinsics -> scalar fallback
+//
+// Whether Vector256/Vector128 are "hardware accelerated" is decided by the JIT when a method is
+// compiled, from the instruction sets the runtime enabled at startup -- it cannot be toggled inside
+// a running process. So to prove the Vector128 and scalar paths are correct we launch a CHILD
+// process with an instruction set disabled via an environment variable and run the whole test suite
+// there (Base58EncodeFast / Base58DecodeFast already cover the 32/64 fast paths thoroughly). Knobs
+// (x86/x64):
+// DOTNET_EnableAVX2=0 disables AVX2 (and AVX-512), leaving SSE -> forces the Vector128 path
+// DOTNET_EnableHWIntrinsic=0 disables all hardware intrinsics -> forces the scalar path
+//
+// Explicit because it spawns processes -- opt-in, not part of every run. The child runs with the
+// default (explicit off), so this test does not re-spawn itself.
+public class VectorInstructionSetTests
+{
+ private readonly ITestOutputHelper _output;
+
+ public VectorInstructionSetTests(ITestOutputHelper output)
+ {
+ _output = output;
+ }
+
+ [Theory(Explicit = true)]
+ [InlineData("DOTNET_EnableAVX2", "0")] // disable AVX2 -> forces the Vector128 path
+ [InlineData("DOTNET_EnableHWIntrinsic", "0")] // disable all hardware intrinsics -> forces the scalar path
+ public void AllTests_Pass_WithVectorInstructionSetDisabled(string environmentVariable, string value)
+ {
+#if DEBUG
+ const string configuration = "Debug";
+#else
+ const string configuration = "Release";
+#endif
+ string testProjectDir = Path.Combine(FindSrcDir(), "Base58Encoding.Tests");
+
+ var startInfo = new ProcessStartInfo
+ {
+ FileName = "dotnet",
+ WorkingDirectory = testProjectDir,
+ RedirectStandardOutput = true,
+ RedirectStandardError = true,
+ UseShellExecute = false,
+ };
+
+ startInfo.ArgumentList.Add("run");
+ startInfo.ArgumentList.Add("-c");
+ startInfo.ArgumentList.Add(configuration);
+ startInfo.ArgumentList.Add("--no-build");
+ startInfo.ArgumentList.Add("--no-restore");
+ startInfo.Environment[environmentVariable] = value;
+
+ using var process = Process.Start(startInfo)!;
+ string stdout = process.StandardOutput.ReadToEnd();
+ string stderr = process.StandardError.ReadToEnd();
+ process.WaitForExit();
+
+ _output.WriteLine($"child with {environmentVariable}={value}:");
+ _output.WriteLine(stdout);
+ if (stderr.Length > 0)
+ {
+ _output.WriteLine("stderr: " + stderr);
+ }
+
+ Assert.Equal(0, process.ExitCode);
+ }
+
+ private static string FindSrcDir()
+ {
+ var dir = new DirectoryInfo(AppContext.BaseDirectory);
+ while (dir is not null && dir.Name != "src")
+ {
+ dir = dir.Parent;
+ }
+
+ return dir?.FullName ?? throw new InvalidOperationException("Could not locate the 'src' directory.");
+ }
+}
diff --git a/src/Base58Encoding/Base58.Decode.cs b/src/Base58Encoding/Base58.Decode.cs
index 1f0ed96..c848ff3 100644
--- a/src/Base58Encoding/Base58.Decode.cs
+++ b/src/Base58Encoding/Base58.Decode.cs
@@ -2,6 +2,8 @@
using System.Buffers.Binary;
using System.Numerics;
using System.Runtime.CompilerServices;
+using System.Runtime.InteropServices;
+using System.Runtime.Intrinsics;
namespace Base58Encoding;
@@ -304,12 +306,7 @@ internal static int TryDecodeBitcoin32Fast(ReadOnlySpan encoded, S
for (int j = 0; j < Base58BitcoinTables.BinarySz32; j++)
{
- ulong acc = 0UL;
- for (int i = 0; i < Base58BitcoinTables.IntermediateSz32; i++)
- {
- acc += intermediate[i] * Base58BitcoinTables.DecodeTable32[i][j];
- }
- binary[j] = acc;
+ binary[j] = TensorDot(intermediate, Base58BitcoinTables.DecodeTable32.AsSpan(j * Base58BitcoinTables.IntermediateSz32, Base58BitcoinTables.IntermediateSz32));
}
// Reduce each term to less than 2^32
@@ -413,12 +410,7 @@ internal static int TryDecodeBitcoin64Fast(ReadOnlySpan encoded, S
for (int j = 0; j < Base58BitcoinTables.BinarySz64; j++)
{
- ulong acc = 0UL;
- for (int i = 0; i < Base58BitcoinTables.IntermediateSz64; i++)
- {
- acc += intermediate[i] * Base58BitcoinTables.DecodeTable64[i][j];
- }
- binary[j] = acc;
+ binary[j] = TensorDot(intermediate, Base58BitcoinTables.DecodeTable64.AsSpan(j * Base58BitcoinTables.IntermediateSz64, Base58BitcoinTables.IntermediateSz64));
}
// Reduce each term to less than 2^32
@@ -487,4 +479,48 @@ internal static int TryDecodeBitcoin64Fast(ReadOnlySpan encoded, S
int r = TryDecodeBitcoin64Fast(encoded, buffer);
return r < 0 ? null : buffer.ToArray();
}
+
+ // Vectorized dot product of two equal-length ulong spans: sum(x[i] * y[i]).
+ // A focused, dependency-free stand-in for TensorPrimitives.Dot, tuned for the small
+ // fixed-length decode columns (IntermediateSz32/64). Widest available width first, then a
+ // scalar tail / fallback. The ulong multiply-accumulate wraps identically to the scalar loop,
+ // so results are bit-for-bit the same.
+ private static ulong TensorDot(ReadOnlySpan x, ReadOnlySpan y)
+ {
+ ref ulong xr = ref MemoryMarshal.GetReference(x);
+ ref ulong yr = ref MemoryMarshal.GetReference(y);
+ int len = x.Length;
+ int i = 0;
+ ulong sum = 0UL;
+
+ if (Vector256.IsHardwareAccelerated && len >= Vector256.Count)
+ {
+ Vector256 acc = Vector256.Zero;
+ int upper = len - Vector256.Count;
+ for (; i <= upper; i += Vector256.Count)
+ {
+ acc += Vector256.LoadUnsafe(ref xr, (nuint)i) * Vector256.LoadUnsafe(ref yr, (nuint)i);
+ }
+
+ sum += Vector256.Sum(acc);
+ }
+ else if (Vector128.IsHardwareAccelerated && len >= Vector128.Count)
+ {
+ Vector128 acc = Vector128.Zero;
+ int upper = len - Vector128.Count;
+ for (; i <= upper; i += Vector128.Count)
+ {
+ acc += Vector128.LoadUnsafe(ref xr, (nuint)i) * Vector128.LoadUnsafe(ref yr, (nuint)i);
+ }
+
+ sum += Vector128.Sum(acc);
+ }
+
+ for (; i < len; i++)
+ {
+ sum += Unsafe.Add(ref xr, i) * Unsafe.Add(ref yr, i);
+ }
+
+ return sum;
+ }
}
diff --git a/src/Base58Encoding/Base58.Encode.cs b/src/Base58Encoding/Base58.Encode.cs
index 0a23262..6e78209 100644
--- a/src/Base58Encoding/Base58.Encode.cs
+++ b/src/Base58Encoding/Base58.Encode.cs
@@ -3,6 +3,8 @@
using System.Diagnostics;
using System.Numerics;
using System.Runtime.CompilerServices;
+using System.Runtime.InteropServices;
+using System.Runtime.Intrinsics;
namespace Base58Encoding;
@@ -255,13 +257,12 @@ private static int ComputeBitcoin32FastRaw(ReadOnlySpan data, Span r
Span intermediate = stackalloc ulong[Base58BitcoinTables.IntermediateSz32];
intermediate.Clear();
- // Matrix multiplication: intermediate = binary * EncodeTable32
+ // Matrix multiplication: for each source limb, add binary[i] * row into intermediate[1..].
+ int rowLen = Base58BitcoinTables.IntermediateSz32 - 1;
+ Span acc = intermediate.Slice(1, rowLen);
for (int i = 0; i < Base58BitcoinTables.BinarySz32; i++)
{
- for (int j = 0; j < Base58BitcoinTables.IntermediateSz32 - 1; j++)
- {
- intermediate[j + 1] += (ulong)binary[i] * Base58BitcoinTables.EncodeTable32[i][j];
- }
+ TensorMultiplyAdd(Base58BitcoinTables.EncodeTable32RowMajor.AsSpan(i * rowLen, rowLen), binary[i], acc);
}
// Reduce each term to be less than 58^5
@@ -366,25 +367,20 @@ private static int ComputeBitcoin64FastRaw(ReadOnlySpan data, Span r
Span intermediate = stackalloc ulong[Base58BitcoinTables.IntermediateSz64];
intermediate.Clear();
+ int rowLen = Base58BitcoinTables.IntermediateSz64 - 1;
+ Span acc = intermediate.Slice(1, rowLen);
for (int i = 0; i < 8; i++)
{
- for (int j = 0; j < Base58BitcoinTables.IntermediateSz64 - 1; j++)
- {
- intermediate[j + 1] += (ulong)binary[i] * Base58BitcoinTables.EncodeTable64[i][j];
- }
+ TensorMultiplyAdd(Base58BitcoinTables.EncodeTable64RowMajor.AsSpan(i * rowLen, rowLen), binary[i], acc);
}
// Mini-reduction to prevent overflow (like Firedancer)
intermediate[15] += intermediate[16] / Base58BitcoinTables.R1Div;
intermediate[16] %= Base58BitcoinTables.R1Div;
- // Finish remaining iterations
for (int i = 8; i < Base58BitcoinTables.BinarySz64; i++)
{
- for (int j = 0; j < Base58BitcoinTables.IntermediateSz64 - 1; j++)
- {
- intermediate[j + 1] += (ulong)binary[i] * Base58BitcoinTables.EncodeTable64[i][j];
- }
+ TensorMultiplyAdd(Base58BitcoinTables.EncodeTable64RowMajor.AsSpan(i * rowLen, rowLen), binary[i], acc);
}
// Reduce each term to be less than 58^5
@@ -419,6 +415,48 @@ private static int ComputeBitcoin64FastRaw(ReadOnlySpan data, Span r
return rawLeadingZeros;
}
+ // acc[k] += row[k] * scale over the whole row: multiply the row by a scalar and add into the
+ // accumulator. The encode counterpart of decode's TensorDot; mirrors TensorPrimitives.MultiplyAdd
+ // (System.Numerics.Tensors) as a tiny dependency-free version tuned for the fixed-length encode
+ // rows. Widest available vector width first, then a scalar tail that also serves as the fallback
+ // when no width is hardware-accelerated. Wrapping ulong multiply-add in source-limb order, so the
+ // result is bit-identical to the scalar loop.
+ private static void TensorMultiplyAdd(ReadOnlySpan row, ulong scale, Span acc)
+ {
+ ref ulong rr = ref MemoryMarshal.GetReference(row);
+ ref ulong ar = ref MemoryMarshal.GetReference(acc);
+ int len = row.Length;
+ int i = 0;
+
+ if (Vector256.IsHardwareAccelerated && len >= Vector256.Count)
+ {
+ Vector256 s = Vector256.Create(scale);
+ int upper = len - Vector256.Count;
+ for (; i <= upper; i += Vector256.Count)
+ {
+ Vector256 a = Vector256.LoadUnsafe(ref ar, (nuint)i);
+ Vector256 r = Vector256.LoadUnsafe(ref rr, (nuint)i);
+ (a + (r * s)).StoreUnsafe(ref ar, (nuint)i);
+ }
+ }
+ else if (Vector128.IsHardwareAccelerated && len >= Vector128.Count)
+ {
+ Vector128 s = Vector128.Create(scale);
+ int upper = len - Vector128.Count;
+ for (; i <= upper; i += Vector128.Count)
+ {
+ Vector128 a = Vector128.LoadUnsafe(ref ar, (nuint)i);
+ Vector128 r = Vector128.LoadUnsafe(ref rr, (nuint)i);
+ (a + (r * s)).StoreUnsafe(ref ar, (nuint)i);
+ }
+ }
+
+ for (; i < len; i++)
+ {
+ Unsafe.Add(ref ar, i) += Unsafe.Add(ref rr, i) * scale;
+ }
+ }
+
private readonly ref struct EncodeState
where T : struct, IBase58Alphabet
{
diff --git a/src/Base58Encoding/Base58BitcoinTables.cs b/src/Base58Encoding/Base58BitcoinTables.cs
index 14ed12b..4c73b86 100644
--- a/src/Base58Encoding/Base58BitcoinTables.cs
+++ b/src/Base58Encoding/Base58BitcoinTables.cs
@@ -30,19 +30,24 @@ internal static class Base58BitcoinTables
[0U, 0U, 0U, 0U, 0U, 0U, 0U, 1U]
];
- // Contains the unique values less than 2^32 such that:
- // 58^(5*(8-j)) = sum_k DecodeTable32[j][k]*2^(32*(7-k))
- internal static readonly uint[][] DecodeTable32 =
+ // EncodeTable32 flattened to a contiguous row-major ulong[]: TensorMultiplyAdd vector-loads a
+ // whole source-limb row at once, so rows must be contiguous, and ulong matches the accumulator /
+ // vector element type (no per-element widening). Same values as the jagged table, via Flatten.
+ internal static readonly ulong[] EncodeTable32RowMajor = Flatten(EncodeTable32, IntermediateSz32 - 1);
+
+ // Decode coefficients stored column-major (transposed): DecodeTable32[j * IntermediateSz32 + i]
+ // holds the coefficient for output limb j, intermediate limb i. Each output limb's column of
+ // IntermediateSz32 coefficients is contiguous, so it can be fed straight into TensorDot.
+ internal static readonly ulong[] DecodeTable32 =
[
- [1277U, 2650397687U, 3801011509U, 2074386530U, 3248244966U, 687255411U, 2959155456U, 0U],
- [0U, 8360U, 1184754854U, 3047609191U, 3418394749U, 132556120U, 1199103528U, 0U],
- [0U, 0U, 54706U, 2996985344U, 1834629191U, 3964963911U, 485140318U, 1073741824U],
- [0U, 0U, 0U, 357981U, 1476998812U, 3337178590U, 1483338760U, 4194304000U],
- [0U, 0U, 0U, 0U, 2342503U, 3052466824U, 2595180627U, 17825792U],
- [0U, 0U, 0U, 0U, 0U, 15328518U, 1933902296U, 4063920128U],
- [0U, 0U, 0U, 0U, 0U, 0U, 100304420U, 3355157504U],
- [0U, 0U, 0U, 0U, 0U, 0U, 0U, 656356768U],
- [0U, 0U, 0U, 0U, 0U, 0U, 0U, 1U]
+ 1277U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U,
+ 2650397687U, 8360U, 0U, 0U, 0U, 0U, 0U, 0U, 0U,
+ 3801011509U, 1184754854U, 54706U, 0U, 0U, 0U, 0U, 0U, 0U,
+ 2074386530U, 3047609191U, 2996985344U, 357981U, 0U, 0U, 0U, 0U, 0U,
+ 3248244966U, 3418394749U, 1834629191U, 1476998812U, 2342503U, 0U, 0U, 0U, 0U,
+ 687255411U, 132556120U, 3964963911U, 3337178590U, 3052466824U, 15328518U, 0U, 0U, 0U,
+ 2959155456U, 1199103528U, 485140318U, 1483338760U, 2595180627U, 1933902296U, 100304420U, 0U, 0U,
+ 0U, 0U, 1073741824U, 4194304000U, 17825792U, 4063920128U, 3355157504U, 656356768U, 1U,
];
// Constants for 64-byte encoding/decoding (from Firedancer)
@@ -72,27 +77,45 @@ internal static class Base58BitcoinTables
[0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 1U]
];
- // Contains the unique values less than 2^32 such that:
- // 58^(5*(17-j)) = sum_k DecodeTable64[j][k]*2^(32*(15-k))
- internal static readonly uint[][] DecodeTable64 =
+ // EncodeTable64 flattened to a contiguous row-major ulong[]: TensorMultiplyAdd vector-loads a
+ // whole source-limb row at once, so rows must be contiguous, and ulong matches the accumulator /
+ // vector element type (no per-element widening). Same values as the jagged table, via Flatten.
+ internal static readonly ulong[] EncodeTable64RowMajor = Flatten(EncodeTable64, IntermediateSz64 - 1);
+
+ // Decode coefficients stored column-major (transposed): DecodeTable64[j * IntermediateSz64 + i]
+ // holds the coefficient for output limb j, intermediate limb i. Each output limb's column of
+ // IntermediateSz64 coefficients is contiguous, so it can be fed straight into TensorDot.
+ internal static readonly ulong[] DecodeTable64 =
[
- [249448U, 3719864065U, 173911550U, 4021557284U, 3115810883U, 2498525019U, 1035889824U, 627529458U, 3840888383U, 3728167192U, 2901437456U, 3863405776U, 1540739182U, 1570766848U, 0U, 0U],
- [0U, 1632305U, 1882780341U, 4128706713U, 1023671068U, 2618421812U, 2005415586U, 1062993857U, 3577221846U, 3960476767U, 1695615427U, 2597060712U, 669472826U, 104923136U, 0U, 0U],
- [0U, 0U, 10681231U, 1422956801U, 2406345166U, 4058671871U, 2143913881U, 4169135587U, 2414104418U, 2549553452U, 997594232U, 713340517U, 2290070198U, 1103833088U, 0U, 0U],
- [0U, 0U, 0U, 69894212U, 1038812943U, 1785020643U, 1285619000U, 2301468615U, 3492037905U, 314610629U, 2761740102U, 3410618104U, 1699516363U, 910779968U, 0U, 0U],
- [0U, 0U, 0U, 0U, 457363084U, 927569770U, 3976106370U, 1389513021U, 2107865525U, 3716679421U, 1828091393U, 2088408376U, 439156799U, 2579227194U, 0U, 0U],
- [0U, 0U, 0U, 0U, 0U, 2992822783U, 383623235U, 3862831115U, 112778334U, 339767049U, 1447250220U, 486575164U, 3495303162U, 2209946163U, 268435456U, 0U],
- [0U, 0U, 0U, 0U, 0U, 4U, 2404108010U, 2962826229U, 3998086794U, 1893006839U, 2266258239U, 1429430446U, 307953032U, 2361423716U, 176160768U, 0U],
- [0U, 0U, 0U, 0U, 0U, 0U, 29U, 3596590989U, 3044036677U, 1332209423U, 1014420882U, 868688145U, 4264082837U, 3688771808U, 2485387264U, 0U],
- [0U, 0U, 0U, 0U, 0U, 0U, 0U, 195U, 1054003707U, 3711696540U, 582574436U, 3549229270U, 1088536814U, 2338440092U, 1468637184U, 0U],
- [0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 1277U, 2650397687U, 3801011509U, 2074386530U, 3248244966U, 687255411U, 2959155456U, 0U],
- [0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 8360U, 1184754854U, 3047609191U, 3418394749U, 132556120U, 1199103528U, 0U],
- [0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 54706U, 2996985344U, 1834629191U, 3964963911U, 485140318U, 1073741824U],
- [0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 357981U, 1476998812U, 3337178590U, 1483338760U, 4194304000U],
- [0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 2342503U, 3052466824U, 2595180627U, 17825792U],
- [0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 15328518U, 1933902296U, 4063920128U],
- [0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 100304420U, 3355157504U],
- [0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 656356768U],
- [0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 1U]
+ 249448U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U,
+ 3719864065U, 1632305U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U,
+ 173911550U, 1882780341U, 10681231U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U,
+ 4021557284U, 4128706713U, 1422956801U, 69894212U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U,
+ 3115810883U, 1023671068U, 2406345166U, 1038812943U, 457363084U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U,
+ 2498525019U, 2618421812U, 4058671871U, 1785020643U, 927569770U, 2992822783U, 4U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U,
+ 1035889824U, 2005415586U, 2143913881U, 1285619000U, 3976106370U, 383623235U, 2404108010U, 29U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U,
+ 627529458U, 1062993857U, 4169135587U, 2301468615U, 1389513021U, 3862831115U, 2962826229U, 3596590989U, 195U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U,
+ 3840888383U, 3577221846U, 2414104418U, 3492037905U, 2107865525U, 112778334U, 3998086794U, 3044036677U, 1054003707U, 1277U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U,
+ 3728167192U, 3960476767U, 2549553452U, 314610629U, 3716679421U, 339767049U, 1893006839U, 1332209423U, 3711696540U, 2650397687U, 8360U, 0U, 0U, 0U, 0U, 0U, 0U, 0U,
+ 2901437456U, 1695615427U, 997594232U, 2761740102U, 1828091393U, 1447250220U, 2266258239U, 1014420882U, 582574436U, 3801011509U, 1184754854U, 54706U, 0U, 0U, 0U, 0U, 0U, 0U,
+ 3863405776U, 2597060712U, 713340517U, 3410618104U, 2088408376U, 486575164U, 1429430446U, 868688145U, 3549229270U, 2074386530U, 3047609191U, 2996985344U, 357981U, 0U, 0U, 0U, 0U, 0U,
+ 1540739182U, 669472826U, 2290070198U, 1699516363U, 439156799U, 3495303162U, 307953032U, 4264082837U, 1088536814U, 3248244966U, 3418394749U, 1834629191U, 1476998812U, 2342503U, 0U, 0U, 0U, 0U,
+ 1570766848U, 104923136U, 1103833088U, 910779968U, 2579227194U, 2209946163U, 2361423716U, 3688771808U, 2338440092U, 687255411U, 132556120U, 3964963911U, 3337178590U, 3052466824U, 15328518U, 0U, 0U, 0U,
+ 0U, 0U, 0U, 0U, 0U, 268435456U, 176160768U, 2485387264U, 1468637184U, 2959155456U, 1199103528U, 485140318U, 1483338760U, 2595180627U, 1933902296U, 100304420U, 0U, 0U,
+ 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 0U, 1073741824U, 4194304000U, 17825792U, 4063920128U, 3355157504U, 656356768U, 1U,
];
+
+ private static ulong[] Flatten(uint[][] rows, int rowLength)
+ {
+ var flat = new ulong[rows.Length * rowLength];
+ for (int i = 0; i < rows.Length; i++)
+ {
+ for (int j = 0; j < rowLength; j++)
+ {
+ flat[i * rowLength + j] = rows[i][j];
+ }
+ }
+
+ return flat;
+ }
}
diff --git a/src/Directory.Packages.props b/src/Directory.Packages.props
index 67b44f0..f1c207b 100644
--- a/src/Directory.Packages.props
+++ b/src/Directory.Packages.props
@@ -1,14 +1,12 @@
-
-
- true
-
-
-
-
-
-
-
-
-
-
+
+ true
+
+
+
+
+
+
+
+
+
\ No newline at end of file
diff --git a/src/PACKAGE.md b/src/PACKAGE.md
index 98b38e8..eac5e5c 100644
--- a/src/PACKAGE.md
+++ b/src/PACKAGE.md
@@ -2,13 +2,6 @@
A high-performance .NET 10 Base58 encoding and decoding library with support for multiple alphabet variants.
-## Features
-
-- **Multiple Alphabets**: Built-in support for Bitcoin (IPFS/Sui/Solana), Ripple, and Flickr alphabets
-- **Optimized Hot Paths**: Firedancer-based fast paths for 32-byte and 64-byte inputs (up to 15x faster than SimpleBase)
-- **Zero-allocation API**: Encode and decode directly into caller-owned buffers
-- **SIMD**: Uses `Vector256` for counting leading zeros
-
## Alphabets
| Property | First Character | Used By |
@@ -94,7 +87,3 @@ int written = Base58.Bitcoin.Encode(data, encodedBytes);
Span decodedBytes = stackalloc byte[Base58.GetTypicalDecodedLength(written)];
int decodedLen = Base58.Bitcoin.Decode(encodedBytes[..written], decodedBytes);
```
-
-## License
-
-MIT