Choosing between GPU and CPU for password hash cracking is one of the most consequential hardware decisions a security professional makes — the wrong choice can mean waiting hours instead of seconds, or burning electricity for zero gain. The gap between a modern GPU like the RTX 5090 and even a 64-core server CPU is not linear: for certain algorithms it is measured in orders of magnitude, while for others the GPU advantage practically disappears. This guide breaks down the real architecture, the real tradeoffs, and the practical implications for every major hash algorithm you will encounter in a pentest or forensic engagement.
Table of Contents
- Why GPU Architecture Dominates Hash Cracking for Most Algorithms
- RTX 5090 Architecture and What It Means for Hashcat
- Algorithm-by-Algorithm: GPU vs CPU Speed Breakdown
- Watts Per Hash: Understanding Energy Efficiency in Cracking
- Practical Hashcat Configuration for GPU Cracking Campaigns
- When to Use CPU Instead of GPU (And When to Use Both)
- Further Reading and Resources
- Frequently Asked Questions
1. Why GPU Architecture Dominates Hash Cracking for Most Algorithms
Password hash cracking is an embarrassingly parallel workload: each candidate password is hashed independently, with no dependency on any other candidate. This is precisely the class of problem that Graphics Processing Units (GPUs) were architecturally built to exploit.
A modern high-end CPU — think AMD EPYC Genoa or Intel Xeon Sapphire Rapids — ships with 64 to 96 physical cores, each deeply out-of-order with large caches, branch predictors, and sophisticated execution units optimised for serial, low-latency workloads. A single NVIDIA RTX 5090, by contrast, exposes 21,760 CUDA cores [[VERIFY exact RTX 5090 shader count at launch]] operating in a massively parallel SIMD fashion across thousands of threads simultaneously.
The SIMD Advantage in Hashing
Hash functions like MD5, SHA-1, SHA-256, and NTLM are composed of bitwise operations, modular additions, and rotations — exactly the operations a GPU's shader pipeline executes at massive throughput. Hashcat, the industry-standard cracking tool, exploits this through OpenCL and CUDA kernels that dispatch billions of hash candidates per second across the GPU's streaming multiprocessors.
The CPU is not idle in this workflow: it handles rule parsing, mask expansion, wordlist streaming, and result reporting. But the hashing computation itself is almost always faster on GPU — sometimes by a factor of 100x or more for fast algorithms.
- GPU strength: throughput-bound, highly parallel, fixed compute paths
- CPU strength: latency-sensitive, memory-intensive, complex branching algorithms
- The key insight: GPU wins when hash computation is the bottleneck; CPU wins when memory bandwidth or sequential logic dominates
2. RTX 5090 Architecture and What It Means for Hashcat
The NVIDIA GeForce RTX 5090 represents the consumer flagship of NVIDIA's Blackwell architecture generation, released in early. While NVIDIA has not published final hashcat kernel benchmark sheets for the RTX 5090 at time of writing [[VERIFY official RTX 5090 hashcat benchmark availability]], the architectural improvements over the RTX 4090 are well-documented and translate directly into cracking throughput.
Key RTX 5090 Specifications Relevant to Cracking
- Memory bandwidth: Significantly increased over the RTX 4090's 1,008 GB/s — estimated at over 1,700 GB/s with GDDR7 [[VERIFY final RTX 5090 memory bandwidth spec]]
- Shader count: Increased over the RTX 4090's 16,384 CUDA cores [[VERIFY final RTX 5090 CUDA core count]]
- TDP: Reported at approximately 575W [[VERIFY RTX 5090 TDP]] — significantly higher than RTX 4090's 450W
- VRAM: 32 GB GDDR7, critical for large rule sets and hash lists
How Blackwell Improves Cracking Over Ada Lovelace (RTX 4090)
The RTX 4090 under hashcat achieved approximately 164 GH/s on MD5 and 370 GH/s on NTLM in community benchmarks. Blackwell's improvements in shader throughput and memory bandwidth suggest meaningful gains on these fast algorithms. For slow KDF algorithms (bcrypt, Argon2, scrypt), the improvement is far less dramatic because those algorithms are intentionally designed to be memory-hard or compute-sequential.
Hashcat's OpenCL/CUDA kernels must be updated to fully exploit Blackwell optimisations, so raw architecture gains do not automatically translate to cracking speed until kernel support matures.
3. Algorithm-by-Algorithm: GPU vs CPU Speed Breakdown
Not all hash algorithms benefit equally from GPU acceleration. The critical factor is whether the algorithm is compute-bound (favours GPU) or memory-hard and sequential (partially resists GPU parallelism). Below is a practical breakdown of the algorithms you will encounter in real engagements.
Fast Algorithms — GPU Dominates
These are legacy or performance-optimised algorithms where the GPU advantage is enormous:
| Algorithm | Hashcat Mode | RTX 4090 Approx. Speed | High-End CPU (e.g., 64-core EPYC) Approx. Speed | GPU Advantage |
|---|---|---|---|---|
| MD5 | 0 | ~164 GH/s | ~1-2 GH/s | ~80-160x |
| NTLM | 1000 | ~370 GH/s | ~3-4 GH/s | ~90-120x |
| SHA-1 | 100 | ~56 GH/s | ~800 MH/s | ~70x |
| SHA-256 | 1400 | ~22 GH/s | ~400 MH/s | ~55x |
| SHA-512 | 1700 | ~8 GH/s | ~250 MH/s | ~32x |
Note: All figures above are community-reported estimates for RTX 4090 single-GPU. RTX 5090 figures are not yet widely available. [[VERIFY against current hashcat benchmark threads]]
Slow/Memory-Hard Algorithms — GPU Advantage Narrows Sharply
Key Derivation Functions (KDFs) used in modern password storage are deliberately designed to resist GPU acceleration:
| Algorithm | Hashcat Mode | RTX 4090 Approx. Speed | High-End CPU Approx. Speed | GPU Advantage |
|---|---|---|---|---|
| bcrypt ($2a$, cost 10) | 3200 | ~100-110 KH/s | ~30-60 KH/s | ~2-4x |
| scrypt (N=16384) | 8900 | ~350-500 KH/s | ~50-150 KH/s | ~3-7x |
| Argon2id (default params) | 13400 | ~1-5 KH/s | ~1-4 KH/s | ~1-2x |
[[VERIFY bcrypt, scrypt, Argon2 RTX 4090 rates — community figures vary significantly by configuration]]
The take-away: if your engagement involves bcrypt or Argon2 at high cost parameters, neither GPU nor CPU will save you without an exceptional wordlist strategy or a pre-computed rainbow table approach.
4. Watts Per Hash: Understanding Energy Efficiency in Cracking
Raw hash rate figures tell only half the story. For professionals running extended cracking campaigns — or cloud environments billed by power consumption — the watts-per-hash (or its inverse, hashes-per-watt) metric determines real-world cost and feasibility.
Why Power Efficiency Matters in Practice
A cracking rig running continuously for 48 hours at 450W (RTX 4090 TDP) consumes 21.6 kWh. At an average US commercial electricity rate of approximately $0.12/kWh, that is roughly $2.59 in electricity cost per 48-hour run — negligible for a pentester, but significant at scale in a multi-GPU array running for weeks. The RTX 5090's higher TDP (~575W reported) means each unit consumes more power per session.
Hashes-Per-Watt Comparison: GPU vs CPU
Using MD5 as the benchmark for fast-algorithm comparison:
- RTX 4090 (450W TDP): ~164 GH/s ÷ 450W ≈ 364 MH/s per watt
- RTX 5090 (~575W TDP): Estimated higher total throughput, but efficiency per watt depends on final silicon behaviour [[VERIFY RTX 5090 MD5 H/s-per-watt once hashcat benchmarks available]]
- AMD EPYC 9654 (360W TDP, 96 cores): ~1.5-2 GH/s on MD5 ÷ 360W ≈ 4-5 MH/s per watt
The GPU achieves roughly 70-90x better energy efficiency than a high-end CPU for MD5. This advantage collapses for memory-hard KDFs:
- bcrypt (cost 10) on RTX 4090: ~110 KH/s ÷ 450W ≈ 244 H/s per watt
- bcrypt (cost 10) on EPYC 9654: ~50 KH/s ÷ 360W ≈ 139 H/s per watt
For bcrypt the GPU is only ~1.75x more efficient in hashes-per-watt — a far cry from the 70x advantage seen on MD5. This is why cloud cracking services like OnlineHashCrack are economically viable: GPU arrays operated at scale amortise power costs across many jobs, delivering per-hash costs that no single-node setup can match.
Cost-Per-Hash in the Cloud Context
When renting GPU compute (AWS p4d, Lambda Labs, vast.ai), the metric that matters is cost per recovered
password candidate tested. A single RTX 4090-class GPU-hour at approximately $0.35-0.50/hr on spot pricing
can test roughly 590 trillion MD5 candidates — enough to exhaust an 8-character mixed-case alphanumeric
keyspace multiple times over.
5. Practical Hashcat Configuration for GPU Cracking Campaigns
Understanding benchmarks is only useful if you can translate them into optimised hashcat configurations. Below are the key parameters, tuning principles, and command-line patterns every security professional needs when running GPU-based cracking sessions.
Essential Hashcat Flags for GPU Optimisation
-
-d 1— select GPU device (use--opencl-infoor--cuda-infoto enumerate devices) -
-w 3— workload profile: 3 is "High", 4 is "Nightmare" (may cause display lag, fine on headless rigs) -
-O— enable optimised kernels (limits max password length to 32 characters but dramatically boosts speed on fast algorithms) -
--force— bypass driver warnings (use cautiously; on bare-metal this is often needed)
Benchmarking Your Own Hardware
Before any real engagement, run a full benchmark to establish your baseline:
hashcat -b --benchmark-all
Or target a specific algorithm:
hashcat -b -m 1000
This is critical because published community benchmarks reflect specific driver versions, CUDA toolkit versions, and system configurations. Your numbers will vary.
Attack Mode Selection for Maximum Efficiency
GPU throughput is wasted if your attack strategy is poorly chosen:
-
Dictionary + rules (-a 0): Best overall coverage per hash-per-second. Use
rockyou.txtwithbest64.ruleorOneRuleToRuleThemAllas a starting point. -
Mask attacks (-a 3): Ideal for known-format passwords (e.g., 8-char alphanumeric company
passwords). Example:
hashcat -a 3 -m 1000 hashes.txt ?u?l?l?l?d?d?d?s - Combinator attacks (-a 1): Useful for passphrase-style passwords combining two wordlists.
- Prince attack (PACK/prince processor): Chain-generates candidates from a wordlist; effective against longer passwords.
Slow Algorithm Strategy
For bcrypt, Argon2id, and scrypt, raw GPU throughput cannot compensate for slow per-hash speed. Prioritise:
- A tightly curated, high-probability wordlist (breach dumps relevant to the target sector)
- Minimal rule sets — every rule multiplies candidates and time
- Known password patterns from the engagement's OSINT phase
6. When to Use CPU Instead of GPU (And When to Use Both)
Despite the overwhelming GPU advantage for fast algorithms, there are genuine scenarios where CPU cracking is preferable, complementary, or the only viable option.
Scenarios Where CPU Cracking Makes Sense
-
Argon2id with high memory parameters: Argon2id is explicitly designed to require large amounts of
memory per hash computation. At high
mparameters (e.g., 64 MB per hash), a GPU's per-core VRAM allocation quickly saturates, meaning only a fraction of CUDA cores can operate simultaneously. A CPU with direct access to 512 GB of system RAM can sometimes keep pace or exceed a single GPU. - John the Ripper CPU-only targets: Some exotic hash formats lack GPU kernel support in hashcat. John the Ripper's CPU implementation may be the only automated option.
- Availability and cost: CPU compute is universally available (every pentest laptop has one). For a quick 15-minute wordlist run on a single bcrypt hash, spinning up cloud GPU infrastructure may not be worth the overhead.
Hybrid GPU + CPU Approaches
Hashcat supports simultaneous use of CPU and GPU devices using the -d flag with multiple device IDs, or
via OpenCL on platforms that expose the CPU as a compute device. In practice, for fast algorithms, the CPU
contribution is so small relative to the GPU that the management overhead is rarely worth it. For slow KDFs,
distributing across all available compute — including CPU — can provide marginal gains.
Distributed Cracking Across Multiple GPUs
Hashcat's --keyspace and offset flags allow deterministic job splitting across multiple nodes. Tools
like Hashtopolis coordinate distributed hashcat agents across GPU fleets, useful for large
enterprise audits. Cloud GPU providers like vast.ai and Lambda Labs allow on-demand scaling — the same model
OnlineHashCrack operates at a managed-service level, removing the infrastructure burden from the analyst entirely.
- Single RTX 4090: Suitable for most SMB-scale hash audits
- 4x RTX 4090 rig: Approaches near-linear scaling on fast algorithms; diminishing returns on bcrypt
- Cloud GPU fleet (8-16 GPUs): Required for large NTDS.dit dumps or time-sensitive forensic deadlines
7. Further Reading and Resources
- Hashcat Official Wiki — hashcat.net
- NIST SP 800-63B: Digital Identity Guidelines (Password Storage) — NIST
- Password Storage Cheat Sheet — OWASP
- RFC 9106: Argon2 Memory-Hard Function — IETF/IANA
- Benchmark Hashcat RTX 5090
- Password Cracking Guide: 5 Latest Techniques
- Hash Algorithms Explained: Secure Password Storage
- How to Extract Hashes (eg: NTLM, Kerberos) from Windows Systems
- GPU Password Cracking Benchmarks: RTX vs CPUs
Run Your Hash Against Our GPU Fleet on OHC
OnlineHashCrack operates a managed fleet of GPU-accelerated cracking nodes, delivering results on MD5, NTLM, SHA-family, bcrypt, WPA, and many other formats without requiring you to provision or maintain hardware. You submit the hash; OHC handles the GPU fleet, wordlists, and rule sets.
Submit your hash — Try OHC →Frequently Asked Questions
How much faster is an RTX 5090 than a CPU for cracking NTLM hashes?
Based on RTX 4090 community benchmarks (~370 GH/s on NTLM) versus a 64-core EPYC (~3-4 GH/s), a high-end GPU is roughly 90-120x faster on NTLM. The RTX 5090 is expected to exceed this, though exact figures await verified hashcat benchmarks on production hardware. [[VERIFY once RTX 5090 hashcat benchmarks are published]]
Is it legal to run GPU password cracking on hashes I found during a penetration test?
Password cracking is legal only when performed on systems and data you are explicitly authorised to test, under a signed scope-of-work or rules of engagement. OnlineHashCrack requires users to confirm they own or have authorisation over any data submitted — cracking unauthorised credentials is a criminal offence in most jurisdictions.
Can I crack bcrypt hashes online with GPU acceleration on OnlineHashCrack?
Yes, OHC supports bcrypt (Hashcat mode 3200). However, because bcrypt is intentionally slow, recovery depends heavily on password complexity and cost factor — OHC's GPU fleet maximises your chances using curated wordlists and targeted rule sets, but there is no guarantee for high-entropy passwords.
Does Argon2id resist GPU cracking better than bcrypt?
Argon2id is generally considered more resistant to GPU cracking than bcrypt because its memory-hard design (parameterised by memory size, not just iteration count) limits the number of parallel GPU threads that can operate simultaneously — at high memory settings, a GPU's advantage over CPU nearly disappears. For both algorithms, a strong, high-entropy password remains the best defence.