Two hash functions, both blessed by serious cryptographers, both still shipping in production code in October 2026, and both solving the same basic problem in almost opposite ways. SHA-3 is the NIST-standardized sponge that nobody was forced to adopt but everybody was told to keep around. BLAKE3 is the upstart tree-hash that broke speed records in 2020 and never looked back. Neither one is going away, and picking between them depends entirely on what you’re protecting and who’s going to audit your choice.
This comparison pulls apart the construction differences, lines up throughput numbers from three independent benchmark sources, and answers the question developers keep typing into search bars: is BLAKE3 actually fast enough to replace SHA-3 in a real system, or is that speed gap irrelevant once you factor in certification requirements? We’ll also cover where each one is deployed right now, what migrating between them actually costs in engineering time, and which one you should reach for depending on your stack.
What Is SHA-3, and Why Does It Still Matter in 2026?
SHA-3 is the hash function family the National Institute of Standards and Technology finalized in FIPS 202 back in August 2015, built on the Keccak sponge construction that won NIST’s public hash function competition. Unlike SHA-256, which processes data through a Merkle-Damgård chain of compression steps, SHA-3 absorbs input into a 1,600-bit internal state and squeezes output back out through the same permutation. That sponge design gives SHA-3 something SHA-256 structurally lacks: native resistance to length-extension attacks without needing a workaround like HMAC.
FIPS 202 defines four fixed-output members (SHA3-224, SHA3-256, SHA3-384, SHA3-512) plus two extendable-output functions, SHAKE128 and SHAKE256, that can produce a digest of any length you ask for. NIST isn’t done tuning the family either. On March 12, 2025, the agency announced it would update FIPS 202 and revise SP 800-185, the companion standard covering cSHAKE, KMAC, TupleHash, and ParallelHash, and a draft revision of SP 800-185 went out for public comment on October 8, 2026. None of that activity signals a break in the algorithm. It’s standards maintenance, not a recall.
Here’s the detail that trips people up in casual conversation: Ethereum does not use standardized SHA3-256. It uses Keccak-256, the pre-finalization version of the algorithm with the original padding scheme rather than the FIPS 202 padding. The two share the same Keccak-f[1600] permutation internally but produce different digests for identical input. If you’ve ever heard someone say “Ethereum runs on SHA-3,” that’s technically imprecise, and it matters if you’re implementing anything that has to match Ethereum’s address derivation or trie hashing byte for byte.
What Is BLAKE3, and Where Did It Come From?
BLAKE3 shipped in January 2020 as the third generation of a hash function lineage that traces back to the 2008 SHA-3 competition, where the original BLAKE was a finalist that lost to Keccak. The design team, which includes Jean-Philippe Aumasson and Zooko Wilcox-O’Hearn, rebuilt the algorithm around a binary Merkle tree of chunks instead of a sequential chain. That structural choice is the entire reason BLAKE3 is fast: each 1,024-byte chunk of input can be hashed independently and in parallel, then the partial results get combined up the tree, so a multi-core machine (or a single core with wide SIMD) can chew through large files far quicker than a function that has to process bytes in strict sequence.
Under the hood, BLAKE3 reuses the same core permutation as BLAKE2s, itself a descendant of the ChaCha stream cipher, but cuts the number of compression rounds from BLAKE2’s 10 down to 7. It also drops BLAKE2’s configurable salt and personalization parameters in favor of a simpler, single keyed mode. The result is a hash function that’s also a built-in extendable-output function, a key derivation function, and a MAC, all from one 32-byte core design. BLAKE3 is published as an open specification with a reference implementation on GitHub, but critically, it has never gone through FIPS validation and isn’t on NIST’s approved algorithm list for federal systems.
Sponge vs Merkle Tree: The Core Architectural Difference
The fastest way to understand why these two functions behave so differently under load is to look at how they move data through their internal state. SHA-3’s sponge construction is inherently sequential: you absorb a block, permute the full 1,600-bit state, absorb the next block, permute again, and so on until you’ve consumed the whole message, then you squeeze output blocks out the same way. There’s no way to parallelize a single Keccak-256 computation across cores; the only parallelism available is hashing multiple independent messages at once.
BLAKE3 flips that constraint on its head by design. Input gets split into 1,024-byte chunks, each chunk is hashed independently using the same ChaCha-derived compression function, and the chunk outputs are combined pairwise up a binary tree until a single root hash emerges. For anything over roughly 1 KiB, this means a single hash operation can legitimately be split across threads or vectorized across a wide SIMD register, which is the direct mechanical reason BLAKE3’s large-file throughput numbers look so different from SHA-3’s. For small inputs under a kilobyte or so, the tree structure provides no benefit and BLAKE3’s performance edge narrows significantly, something worth remembering if your workload is dominated by hashing short tokens rather than large files.
Both constructions resist length-extension attacks, though they get there differently. SHA-3’s sponge design makes the attack structurally impossible because the internal state isn’t simply the previous output. BLAKE3 achieves the same resistance through its tree finalization step, which domain-separates the root node from internal nodes using a flag bit, so an attacker can’t extend a known digest into a valid hash of a longer message. Classic SHA-256, by contrast, remains vulnerable to length-extension unless you wrap it in HMAC, which is a real operational difference that’s pushed some systems toward the newer designs.
Full Specification Comparison
| Attribute | SHA-3 (SHA3-256 / Keccak) | BLAKE3 |
|---|---|---|
| Standard body | NIST FIPS 202 | Open spec, no NIST/FIPS status |
| Finalized | August 2015 | January 2020 |
| Core construction | Sponge (absorb/squeeze) | Binary Merkle tree |
| Internal permutation | Keccak-f[1600] | ChaCha-derived (from BLAKE2s) |
| Compression rounds | 24 rounds of Keccak-f | 7 rounds per chunk (vs. BLAKE2’s 10) |
| Default output size | 256 bits (SHA3-256 variant) | 256 bits, extendable to any length |
| Native XOF support | Yes, via SHAKE128/SHAKE256 | Yes, built into the core design |
| Native keyed/MAC mode | Via KMAC (SP 800-185) | Built-in keyed hashing mode |
| Native key derivation | Via cSHAKE/KMAC constructions | Built-in derive_key mode |
| Parallelizable single-hash | No, strictly sequential | Yes, via tree structure and SIMD |
| Length-extension resistant | Yes, structurally | Yes, via tree finalization flags |
| Dedicated CPU instructions | None widely deployed | None; relies on general SIMD (AVX2/AVX-512/NEON) |
| FIPS 140-3 module eligible | Yes | No |
| Typical use today | Keccak-256 in Ethereum, federal systems, SPHINCS+ | Build systems, content-addressed storage, game assets |
Benchmarks: Three Independent Sources, Compared
Throughput numbers for hash functions swing wildly depending on input size, CPU microarchitecture, compiler flags, and whether the implementation uses hand-tuned assembly. To avoid cherry-picking a single flattering benchmark, here’s what three separate measurement efforts found when they put SHA3-256 and BLAKE3 on the same hardware.
The first data point comes from a 2025 throughput benchmark run on an AMD EPYC 4245P (Zen 5) processor measuring large-buffer performance. At 1 MiB input sizes, that test recorded SHA-256 at 2,373 MB/s, SHA3-256 at 686 MB/s, and BLAKE3 at 13,196 MB/s, putting BLAKE3 roughly 19 times ahead of SHA3-256 at that specific buffer size. A second, independently run benchmark on the same EPYC 4245P chip using a 10 MiB buffer told a more conservative story: SHA-256 at 1,772 MB/s, SHA3-256 at 509 MB/s, and BLAKE3 at 6,121 MB/s, a roughly 12x gap over SHA3-256. The spread between those two results on identical silicon is itself the lesson: don’t trust a single MB/s figure without knowing the exact test harness behind it.
The third source is cycle-accurate rather than throughput-based. Thomas Pornin’s SHA-3 candidate speed report, hosted at bolet.org, and independent technical writeups like kerkour.com’s SHA-3 breakdown put reference-style SHA3-256 implementations around 12.6 cycles per byte, dropping to roughly 6.4 cycles per byte with AVX-512VL optimization. SHA-256 with dedicated hardware instructions can run as low as 1 to 3 cycles per byte on modern silicon, which is the real reason SHA-256 still beats SHA3-256 even before BLAKE3 enters the picture.
| Benchmark source | Hardware | SHA3-256 | BLAKE3 | BLAKE3 advantage |
|---|---|---|---|---|
| EPYC 2025 run, 1 MiB buffers | AMD EPYC 4245P (Zen 5) | 686 MB/s | 13,196 MB/s | ~19x |
| EPYC 2025 run, 10 MiB buffers | AMD EPYC 4245P (Zen 5) | 509 MB/s | 6,121 MB/s | ~12x |
| BLAKE3 team’s own 16 KiB benchmark | Intel Cascade Lake-SP 8275CL | Not directly tested | ~12x SHA-256 single-thread | ~12x vs. SHA-256 |
| Pornin cycle report / kerkour.com | General x86-64, AVX-512 | 6.4-12.6 cycles/byte | Sub-1 cycle/byte, large files | 6-12x on cycle count |
Taking the most conservative of the three independent figures, BLAKE3 is roughly 12 times faster than SHA3-256 on large-buffer workloads on modern server-class CPUs. That’s the number worth repeating in a planning meeting, not the flashier 19x figure that shows up under slightly different test conditions. Either way, the direction of the gap is consistent across every source: BLAKE3 wins on raw throughput, and it wins by more than an order of magnitude once file sizes get large enough for its tree parallelism to kick in.
Hardware Acceleration: Why SHA-3 Doesn’t Get the SHA-NI Treatment
Part of why SHA-256 still outruns SHA3-256 in most software stacks has nothing to do with algorithm design and everything to do with silicon. Intel’s SHA Extensions, a set of seven SSE-based instructions documented in Intel’s own technical whitepaper, accelerate SHA-1 and SHA-256 directly in hardware. ARMv8-A chips carry equivalent cryptographic extensions for the same two algorithms. Neither Intel nor AMD nor ARM ships a comparable dedicated instruction set for SHA-3’s Keccak permutation, so SHA3-256 implementations are stuck relying on general-purpose integer or SIMD instructions, which is a meaningful part of why it trails SHA-256 in nearly every throughput test.
BLAKE3 is in the same boat as SHA-3 in terms of lacking a dedicated instruction set, but it compensates by being designed from the start to exploit whatever wide SIMD a CPU already has, AVX2, AVX-512, or NEON, plus genuine multi-core parallelism for large inputs. That’s a software-engineering advantage rather than a hardware one, and it’s portable: BLAKE3 gets faster on newer CPUs with wider vector units without anyone having to add a new instruction, whereas SHA-3 is permanently waiting on a chipmaker that has shown no sign of prioritizing it the way SHA-256 was prioritized back in 2013.
Security Margins and the State of Cryptanalysis
Neither algorithm has a known practical collision or preimage attack against its full, standardized form as of late 2026. SHA-3’s internal state is 1,600 bits wide against a 256-bit output, leaving an enormous security margin that’s been the subject of continuous academic cryptanalysis since the original Keccak submission, with published results limited to reduced-round variants that don’t threaten the full 24-round design. BLAKE3 inherits most of its security analysis from BLAKE2 and, further back, from the SHA-3 competition process itself, since the original BLAKE was a finalist. Reducing the round count from BLAKE2’s 10 to BLAKE3’s 7 did shrink that margin, and the design team has published their own analysis arguing 7 rounds still leaves adequate distance from the best known attacks, but it’s a smaller cushion by design than either SHA-3 or BLAKE2 carry.
The more realistic risk for both algorithms isn’t a theoretical cryptanalytic break, it’s implementation mistakes: incorrect domain separation, unsafe output truncation, or using a raw hash where a dedicated MAC construction was actually needed. That risk profile is identical for SHA-3 and BLAKE3, and it’s a bigger practical concern than which permutation is running underneath.
Quantum Resistance: Output Length Matters More Than Construction
A common misconception is that SHA-3, as the newer NIST standard, is automatically more quantum-resistant than older or less formal designs. That’s not accurate for either SHA-3 or BLAKE3. Quantum attacks against hash functions are generic: Grover’s algorithm cuts preimage search roughly in half in bit-strength terms regardless of the underlying construction, and generic quantum collision-search techniques reduce effective collision resistance to roughly a third of the classical bit-length. Because SHA3-256 and BLAKE3 both default to 256-bit output, they land in the same place under quantum analysis: approximately 128 bits of quantum preimage resistance and roughly 85 bits of quantum collision resistance, give or take depending on the exact attack model.
If quantum collision resistance is the actual design requirement, the lever to pull is output length, not algorithm family. SHA3-512 and a 512-bit BLAKE3 digest both buy a larger quantum margin than their 256-bit counterparts. Neither SHA-3 nor BLAKE3 is a post-quantum signature or key-exchange scheme; they’re hash primitives, and their quantum story is governed by arithmetic on the output size rather than which team designed the compression function.
Standards and Certification: The Line That Actually Decides Adoption
This is where the comparison stops being about speed and starts being about procurement paperwork. SHA-3 sits inside FIPS 202, which means it can go into a FIPS 140-3 validated cryptographic module, the kind of certification that federal agencies, many financial institutions, and healthcare systems bound by compliance regimes either require or strongly prefer. BLAKE3 has no FIPS status, no NIST validation path currently underway, and no realistic timeline for getting one, because it was never submitted to a NIST competition the way SHA-3’s Keccak was.
That single fact filters almost every deployment decision downstream. A bank’s data-integrity pipeline that has to pass a FIPS audit is not going to swap in BLAKE3 no matter how many MB/s it gains, because the auditor’s checklist asks for an approved algorithm, not a fast one. Conversely, a build system, a game engine, or a content-addressed storage layer with no compliance obligation has zero reason to accept SHA3-256’s throughput penalty just to get a certification nobody is asking for.
Real-World Adoption: Where Each One Actually Runs Today
BLAKE3’s adoption pattern is concentrated in systems that care about raw throughput on large data and don’t carry regulatory baggage. Rust’s Cargo package manager uses BLAKE3 for package checksums. Google’s Bazel build system lists BLAKE3 as a supported file digest function for its build graph. Cloudflare uses BLAKE3 for hashing request content inside Cloudflare Pages, and it’s also supported in Cloudflare Pipelines. The OpenZFS filesystem added native BLAKE3 support for checksumming filesystem blocks, trading some CPU-instruction acceleration for raw parallel throughput on modern multi-core storage servers. The Mesa 3D graphics library adopted BLAKE3 for faster Vulkan shader hashing, a change that also touched performance in games like Tekken 8. Chia Network built its entire proof-of-space-and-time consensus mechanism around BLAKE3’s speed. ClickHouse, the columnar analytics database, offers BLAKE3 as a checksum option for its storage engine.
SHA-3’s footprint looks almost opposite: regulatory, blockchain, and research-adjacent rather than throughput-adjacent. Ethereum’s entire account and state model runs on Keccak-256 for address derivation, transaction hashing, and the Merkle-Patricia trie, plus the EVM’s native KECCAK256 opcode that every Solidity contract can call directly. The IETF formalized SHAKE128 and SHAKE256 for digital signatures in TLS through RFC 8702, though mainstream TLS stack adoption of that option remains limited next to SHA-256-based suites. In October 2025, the IETF published RFC 9861, standardizing KangarooTwelve and TurboSHAKE, both built on the same Keccak permutation but with half the rounds for roughly double the speed, aimed at narrowing exactly the throughput gap that pushes developers toward BLAKE3 in the first place. SPHINCS+, the stateless hash-based signature scheme finalized in FIPS 205, uses SHA-3/SHAKE as one of its two standard parameterizations, and Ethereum researchers have published proposals adapting SPHINCS+-style signatures for EVM verification using the chain’s native KECCAK256 opcode to keep gas costs in the 127,000 to 150,000 range, a research direction rather than a shipped feature as of this writing.
Library and Language Ecosystem
Neither algorithm is hard to reach from a modern language toolchain, but the maturity and performance of the bindings differ. SHA-3 benefits from nearly two decades of integration work: OpenSSL has shipped SHA3-224 through SHA3-512 and SHAKE128/256 since version 1.1.1, which means every language with OpenSSL bindings (Python’s hashlib, Ruby’s openssl gem, Go’s crypto/sha3 package, Java’s built-in MessageDigest) gets SHA-3 support essentially for free. libsodium also exposes a generic hash interface that can route to SHA-3-family functions on platforms where it’s compiled in.
BLAKE3’s ecosystem is younger but moved fast specifically because the reference implementation was written with portable SIMD in mind from day one. The official Rust crate remains the reference implementation and the fastest in most benchmarks, since Rust is what the BLAKE3 team builds and tests against first. Official or team-endorsed bindings exist for Python, Node.js, Go, C, and Java, and the algorithm has also been reimplemented independently in pure Go (notably in the lukechampine.com/blake3 package referenced in several open-source build tools) for projects that want to avoid a C dependency entirely. The tradeoff is that third-party or pure-language reimplementations don’t always match the hand-tuned SIMD assembly in the official releases, so a benchmark run against a pure-Go BLAKE3 port can look meaningfully slower than the headline numbers published by the BLAKE3 team itself.
| Language/Runtime | SHA-3 support | BLAKE3 support |
|---|---|---|
| Python | Built into hashlib (sha3_256, shake_128, etc.) | Official blake3 PyPI package |
| Rust | sha3 crate (RustCrypto) | Official blake3 crate (reference implementation) |
| Go | Standard library golang.org/x/crypto/sha3 | Third-party ports (e.g., lukechampine.com/blake3) |
| Node.js / JavaScript | Built into Node’s crypto module | Official blake3 npm package (native bindings) |
| Java | Built into java.security.MessageDigest (JDK 9+) | Third-party libraries only, no JDK-native support |
| C / C++ | OpenSSL, libsodium, Botan | Official reference C implementation |
Energy Use on Constrained and Embedded Devices
Most of the benchmarking in this comparison targets server and desktop CPUs, but hash function choice also matters on battery-powered and embedded hardware, where energy per bit hashed is often a bigger concern than raw throughput. Research evaluating SHA-3-family functions on resource-constrained IoT devices has measured KangarooTwelve, the reduced-round Keccak variant, achieving energy efficiencies in the range of 0.9 to 1.2 gigabits per watt under constrained test conditions, a meaningful improvement over full-round SHA-3 on the same class of hardware. That makes the reduced-round Keccak derivatives a more realistic fit for battery-sensitive sensor networks than either full SHA3-256 or BLAKE3.
BLAKE3’s parallel tree design, by contrast, is optimized around desktop and server assumptions: multiple cores, wide SIMD registers, and large input buffers where the tree structure has room to pay off. A single-core microcontroller hashing small sensor readings a few bytes at a time doesn’t get much benefit from BLAKE3’s parallelism and may do just as well, or better on energy, with a purpose-built lightweight construction. This is a case where neither SHA-3 nor BLAKE3 in their standard form is clearly the right default; it’s a reminder that “faster on a server benchmark” and “more energy-efficient on a coin-cell battery” are different engineering problems that happen to share the word “hashing.”
Pricing and Infrastructure Cost Comparison
Both algorithms are implemented in free, open-source libraries, so there’s no license fee attached to either one. The real cost difference shows up in compute time at scale, and in whether your compliance requirements force you into certified hardware. Using AWS’s published on-demand rate for a c7a.xlarge instance (AMD EPYC-based, 4 vCPU) in US East at $0.2053 per hour, and the conservative 12 MiB-scale throughput figures from the benchmark table above, here’s what hashing a large, steady stream of data actually costs in raw compute time on otherwise identical infrastructure.
| Cost factor | SHA3-256 | BLAKE3 |
|---|---|---|
| Library licensing | Free (OpenSSL, libsodium, etc.) | Free (MIT/Apache-2.0, official Rust/C/Python bindings) |
| Throughput per vCPU-hour (10 MiB buffers) | ~1.75 TB/hour | ~21.2 TB/hour |
| Approx. compute cost to hash 1 PB on c7a.xlarge ($0.2053/hr) | ~$117 | ~$9.70 |
| FIPS 140-3 validated module required for compliance | Available (SHA-3 is FIPS-approved) | Not available at any price; no FIPS path exists |
| AWS CloudHSM, if FIPS-certified hardware is mandated | ~$1.45-$1.80/hour per HSM, region-dependent | Not applicable; BLAKE3 can’t run inside a FIPS-scoped HSM workflow |
Read that table carefully before treating the compute-cost gap as the deciding factor. If your workload genuinely needs certified hardware, the CloudHSM line item doesn’t disappear just because BLAKE3 is cheaper per terabyte; it simply isn’t an option on the table at all. For workloads with no compliance constraint, the compute savings compound fast at scale, which is exactly the economic case that pushed companies like Cloudflare and Chia Network toward BLAKE3 in the first place.
Pros and Cons
SHA-3 / SHA3-256
- FIPS 202 approved, eligible for FIPS 140-3 validated modules and government/financial compliance work
- Structurally immune to length-extension attacks with a huge internal-state security margin (1,600 bits vs. 256-bit output)
- Native XOF support via SHAKE128/SHAKE256, plus KMAC, cSHAKE, and TupleHash for keyed and tuple hashing under SP 800-185
- Underpins Ethereum’s Keccak-256 and SPHINCS+’s SHA-3 parameterization, so it’s already load-bearing in blockchain and post-quantum signature infrastructure
- No dedicated CPU instruction set on any mainstream architecture, so software throughput trails SHA-256 and BLAKE3 by a wide margin
- Mainstream TLS stacks still default to SHA-256/SHA-384, limiting SHA-3’s real-world reach in the protocol that matters most for everyday internet traffic
BLAKE3
- Roughly 12x the throughput of SHA3-256 on large buffers, per the most conservative of three independent benchmarks
- Genuinely parallelizable at the algorithm level via its Merkle tree structure, scaling with both SIMD width and core count
- One design covers hashing, XOF, keyed MAC, and key derivation without bolting on separate constructions
- Proven in production at Cargo, Bazel, Cloudflare Pages, OpenZFS, Mesa, Chia Network, and ClickHouse since 2020
- No FIPS validation path exists, closing the door on federal, many financial, and healthcare compliance use cases
- Smaller security margin than SHA-3 by design (7 rounds vs. Keccak’s 24), and its speed advantage shrinks substantially on short inputs under roughly 1 KiB
Migration Guide: Moving Between SHA-3 and BLAKE3
Teams usually migrate in one of two directions: toward BLAKE3 for throughput on an internal system with no compliance exposure, or toward SHA-3 because an auditor flagged an unapproved algorithm. Both paths follow a similar shape.
- Inventory every call site. Grep your codebase for the hash library imports, not just the function names, since SHA3-256 and Keccak-256 calls often live under the same library namespace but produce different digests.
- Classify by compliance exposure. Separate call sites that touch regulated data, signed artifacts, or anything subject to a FIPS, SOC 2, or PCI audit from internal-only checksums like build caches or CDN content hashes.
- Freeze a digest format version. Add an explicit algorithm tag to any stored hash (a version byte or field) so old digests computed under the previous algorithm remain distinguishable from new ones during a phased rollout.
- Swap the library, not the interface. Both algorithms have mature bindings for most languages (OpenSSL and libsodium for SHA-3, the official
blake3crate and its Python/Node bindings for BLAKE3); keep your internal hashing interface stable and swap the implementation behind it. - Re-benchmark on your actual hardware. The 12x-to-19x spread documented above means you should run your own throughput test on your production instance type rather than trusting any single published number.
- Dual-write during transition. For anything with long-lived stored hashes (file integrity databases, deduplication indexes), compute both old and new digests in parallel until every consumer has migrated to reading the new format.
- Update compliance documentation last. If you’re moving toward SHA-3 for audit reasons, the algorithm swap is the easy part; updating your module’s FIPS 140-3 documentation and getting it re-validated is the part that takes months, not days.
- Decommission the old path deliberately. Set a hard cutover date for removing the legacy hashing code, rather than letting both implementations linger indefinitely and double your maintenance surface.
A minimal code-level comparison helps illustrate how close the two implementations actually are at the API level, once you’ve picked a library:
# SHA3-256 via Python's hashlib (OpenSSL-backed)
import hashlib
digest = hashlib.sha3_256(data).hexdigest()
# BLAKE3 via the official Python binding
import blake3
digest = blake3.blake3(data).hexdigest()
The interface difference is trivial. The decision that actually matters happens before you write either line: does this code path need to survive a compliance audit, and does it process enough data for BLAKE3’s parallelism to pay off.
Use-Case Recommendations
- Build systems and package managers: BLAKE3. Cargo and Bazel already made this call; checksum verification on every build is exactly the high-volume, no-compliance workload BLAKE3 was designed for.
- Government, financial, or healthcare data integrity pipelines: SHA-3 (or SHA-256). If a FIPS 140-3 validated module is a procurement requirement, BLAKE3 is not on the table regardless of its speed advantage.
- Ethereum or EVM-compatible smart contract development: Keccak-256, not generic SHA-3 and not BLAKE3. Matching the EVM’s native
KECCAK256opcode exactly is non-negotiable for address derivation and trie consistency. - Content-addressed storage, CDN edge hashing, or deduplication at scale: BLAKE3. Cloudflare Pages and OpenZFS both chose it for exactly this throughput-bound scenario.
- Post-quantum hash-based digital signatures: SHA-3/SHAKE, since FIPS 205’s SPHINCS+ standard uses it as one of its two approved parameterizations; BLAKE3 has no equivalent standardized signature scheme built on top of it yet.
- Game engines and real-time asset pipelines: BLAKE3, following Mesa’s adoption for Vulkan shader hashing, where low-latency, parallel hashing of frequently changing assets matters more than certification status.
- General-purpose TLS, certificates, and everyday API authentication: Neither, in practice. Mainstream SHA-256/SHA-384 remains the default for TLS 1.3 cipher suites and HKDF, and that isn’t likely to change just because SHA-3 or BLAKE3 exist as alternatives.
The Verdict
There’s no single winner here, and that’s not a dodge, it’s the actual shape of the data. On throughput, BLAKE3 wins decisively: roughly 12x faster than SHA3-256 on large buffers using the most conservative of three independent benchmark sources, scaling further with available CPU cores and SIMD width. On certification, SHA-3 wins just as decisively, because FIPS 202 approval is a binary gate that no amount of speed can unlock, and BLAKE3 sits on the wrong side of it with no path forward. On security margin, SHA-3’s 1,600-bit sponge state against a 256-bit output gives it more theoretical headroom than BLAKE3’s 7-round design, though neither has a known practical break as of October 2026.
If you’re choosing for a system with no compliance exposure and workloads dominated by files larger than a few kilobytes, BLAKE3’s 12x throughput advantage is hard to argue against, and the list of production deployments at Cargo, Bazel, Cloudflare, OpenZFS, Mesa, Chia Network, and ClickHouse shows it holds up under real load. If you’re choosing for a system that has to answer to an auditor, a regulator, or Ethereum’s EVM, the question isn’t close: SHA-3, in its appropriate form (standardized SHA3-256, or Keccak-256 specifically for EVM compatibility), is the only option that satisfies the requirement, full stop.
Frequently Asked Questions
Is BLAKE3 actually faster than SHA-3 in every scenario?
No. BLAKE3’s speed advantage comes from its parallel Merkle tree structure, which only pays off once input size clears roughly 1 KiB. For short inputs like tokens, small JSON payloads, or individual database keys, the gap narrows substantially, and in some short-input benchmarks both BLAKE3 and SHA-3 perform similarly because fixed per-call overhead dominates the measurement rather than raw throughput.
Can BLAKE3 ever become FIPS-approved?
There’s no active NIST submission or validation track for BLAKE3 as of late 2026, and getting a new primitive into a FIPS standard typically requires a formal competition process similar to what produced SHA-3 itself. Nothing is technically impossible, but there’s no indication this is happening, and teams with compliance requirements should not plan around it becoming available.
Is Ethereum’s hashing the same as standard SHA-3?
No, and this trips up a lot of implementers. Ethereum uses Keccak-256, the pre-FIPS-202 version of the algorithm with the original Keccak padding rather than the padding NIST finalized for SHA3-256. Both share the same underlying Keccak-f[1600] permutation, but they produce different digests for the same input, so you cannot substitute a standard SHA3-256 library call where Ethereum’s KECCAK256 opcode behavior is required.
Does SHA-3 have any dedicated hardware acceleration at all?
Not in the way SHA-256 does. Intel, AMD, and ARM all ship dedicated instruction extensions for accelerating SHA-256 (and SHA-1), but none of the three major CPU architectures has shipped an equivalent widely deployed instruction set specifically for Keccak/SHA-3. Some specialized ASICs and FPGAs implement Keccak acceleration for niche use cases, but nothing comparable exists in mainstream consumer or server CPUs.
Is BLAKE3 less secure than SHA-3 because it uses fewer rounds?
BLAKE3 uses 7 compression rounds versus Keccak’s 24 rounds of permutation, which does mean a smaller raw security margin on paper. The BLAKE3 design team has published analysis arguing 7 rounds still maintains a healthy distance from the best published attacks, and no practical collision or preimage attack exists against full BLAKE3 as of October 2026. It’s a smaller margin by design, not a demonstrated weakness.
Which one should I use for file integrity checks on a personal backup system?
BLAKE3, in nearly every case. Personal and small-team backup systems have no compliance requirement forcing a FIPS-approved algorithm, and BLAKE3’s throughput advantage on large files (photos, videos, archives) will make checksum verification passes noticeably faster, which is exactly the use case OpenZFS built native support for.
Does KangarooTwelve close the speed gap for SHA-3-family functions?
Partially. KangarooTwelve and TurboSHAKE, standardized in RFC 9861 in October 2025, use the same Keccak permutation as SHA-3 but cut the round count in half, roughly doubling throughput compared to standard SHAKE. That narrows the gap with BLAKE3 but doesn’t close it, and real-world adoption of KangarooTwelve outside research and specialized contexts remains limited as of this writing.
Should a new project default to SHA-256 instead of either of these?
For anything touching TLS, certificates, or general-purpose API authentication, yes, SHA-256 remains the pragmatic default because of its hardware acceleration and universal library support. Reach for SHA-3 specifically when a standard or compliance requirement names it, and reach for BLAKE3 specifically when you’ve measured a real throughput bottleneck on large-input hashing that a parallel tree hash would actually fix.




