Poker is a game of leveraging information and setting prices.

Every real poker decision comes back to the same question: what does my opponent likely have, what are they likely to do with it, and what price can I set to make that mistake cost them the most?

A poker solver, by contrast, is not trying to answer poker that way. A solver is not asking, “What does this specific player misunderstand?” It is not asking, “What size gets this guy to overfold?” It is not asking, “How do I punish this pool tendency?”

A solver is solving a different problem.

What does a poker solver do?

A poker solver is any software that takes a defined poker situation and computes an approximate Nash equilibrium strategy for both players.

That is the technical definition.

In plain English, a solver takes a simplified version of poker (fixed ranges, fixed bet sizes, fixed stack depths, and a fixed decision tree) and then works toward a pair of strategies where neither player can gain by changing only their own play. The measure of how close a strategy is to this ideal is called “exploitability.”

That is what people mean when they talk about “GTO.” For what an equilibrium guarantees and where that guarantee stops, see Nash equilibrium in poker.

And that is also why a solver is often misunderstood.

Many players hear “optimal” and assume a solver is telling them how to play poker in the real world. That is not what it is doing. It is telling you how to play the toy universe you gave it.

How does a poker solver work?

The algorithm behind most modern poker solvers is called Counterfactual Regret Minimization (CFR). At a high level, CFR works like this:

You start with a strategy pair. The solver evaluates the expected value of actions across the decision tree you gave it. It compares the EV of each action to the EV of the current mixed strategy at each decision point. Actions that outperform the current mix accumulate positive regret. Actions that underperform it accumulate negative or zero regret. Then the solver updates future action frequencies so that more weight goes to actions with positive regret.

Repeat that process enough times, and the average strategy approaches equilibrium.

That is the core loop.

That matters because it tells you what a solver really is. It is not intuition. It is not instinct. It is not a black-box poker soul. It is a regret-minimization engine solving the mathematical model you handed it.

The result is not a single “correct” action. It is a mixed strategy, a set of frequencies. The solver might say “bet 67% pot with this hand 40% of the time, check 60% of the time.” That mix is what makes the strategy unexploitable.

What a Solver Needs From You

A solver does not play poker. You play poker. The solver needs you to define the problem before it can solve anything:

  • Preflop ranges: what hands each player can have
  • Bet sizes: what sizing options are available (1/3 pot, 2/3 pot, pot, all-in, etc.)
  • Stack depth: how deep the stacks are relative to the pot
  • Board: the community cards

Change any of these inputs and the output changes. A solver’s answer to “should I bet?” is always conditional on the universe you defined. If your opponent’s actual range differs from what you told the solver, the output is wrong. Not wrong in theory, wrong in practice.

Why GTO Study Still Matters

If solvers give answers for simplified toy games and not real poker, why study them?

Because GTO provides a baseline. It tells you what balanced play looks like, so you can recognize when your opponents deviate from it, and punish those deviations.

A player who never studies theory does not know what “normal” looks like. They cannot identify that a villain is folding too much to river bets, because they have no reference point for how much folding is correct.

Solver study gives you that reference point. It is the map. Exploitative play is navigating the actual terrain.

Where Solvers Fall Short

Solvers assume both players are trying to play optimally. Real opponents are not.

  • A solver makes about 30% of its river bets bluffs, because that makes the opponent indifferent to calling. But if your opponent folds to river bets 70% of the time, you should bluff far more often.
  • A solver says to call with middle pair because the pot odds justify it. But if the villain never bluffs the river, folding is clearly correct.
  • A solver balances between thin value bets and bluffs. But against a calling station who never folds, you should drop every bluff and value bet relentlessly.

The solver is not wrong in these cases. It is solving for a different opponent, one who plays back optimally. Your opponent does not.

There is a second gap, and it has nothing to do with the opponent. Inside its model the solver scores every showdown knowing both hands and every fold knowing what was folded. A live player sees only the line that was played. When the villain mucks, you do not learn what they had; when they fold, you do not learn how they would have responded to a different bet; and if they mix their play, one observed action is one sample, not the frequency. The solver’s counterfactual values are exactly the thing a table never shows you, which is why the reads you build there have to come from ranges and patterns rather than from a single result.

Is there an exploitative poker solver? Best response

Equilibrium is not the only thing a solver can compute. There is a second question, older than GTO and much closer to how winning players actually think: if I know how this opponent plays, what is the most profitable strategy against them?

Game theory calls that strategy a best response. Instead of assuming a perfect opponent, you feed in a real one: their calling ranges, their fold frequencies, their sizing habits. The math then stops hedging. It no longer mixes frequencies to stay unexploitable, because there is nothing to defend against. It simply finds the line that wins the most against that specific player.

A best response is the mathematical ceiling of exploitative play. Every read you make at the table is a small human step toward it. You notice a player overfolds rivers, so you bluff more rivers. Computing a true best response takes the same idea to its limit, checking every decision against everything known about the opponent, across all 1,755 strategically distinct flops.

The two kinds of solving answer different questions. An equilibrium solver answers: what is the play no one can exploit? A best response answers: what is the play that exploits this player the most? At tables where opponents have large, repeating leaks, the second answer is worth more.

What node locking means, and why a few locks answer the wrong question

Node locking is the solver feature that lets you fix part of an opponent’s strategy and solve again. You tell the solver that this player folds too much to a river bet, and it finds the counter. It is the closest thing most solvers have to an exploitative mode, and it is genuinely useful for studying one deviation at a time.

It has three limits worth knowing before you trust the answer. One locked node cannot represent the whole opponent. Lock a flop overcall and the solver still assumes the player defends the turn and river perfectly, so it can reject a bluff that would work against the real player. GTO Wizard’s own example: locking an opponent to check back too often makes the big blind lead about 14% of the time, and adding that the opponent also overfolds to that lead changes the recommendation to leading the whole range. The solver solved the locked game correctly and answered the wrong question about the opponent.

A small input error flips the decision. A river bluff risking $150 to win $200 earns $7.50 if nine of twenty equally likely hands fold. Turn one of those nine folds into a call and the same bluff loses $10. A five-point error in the assumed fold rate reverses the play, and no amount of extra solving corrects the input.

Reproducing a real opponent by hand does not scale. A tree with 100,000 opponent decisions, at ten seconds to set or check each one, is about 278 hours of locking, and every one of those judgments has to be right for the exploit to hold.

This is why Poker Shark starts from complete opponent styles instead of asking you to lock nodes. Each style’s behavior is already specified across every street, size and runout, the same style plays the hand against you and sits under the analysis, and the best response is computed against all of it. When you want a steady reference, you can play the GTO opponent, which follows hand-written rules modeled on solver play, in the same arena.

How to study with a poker solver

The best use of a solver is not to memorize its output. It is to use solver study to build intuition about hand strength, board texture, and bet sizing, and then to deviate from that baseline when your opponent gives you a reason to. A quick place to start is the preflop ranges tool, where you can see how the GTO solver opens each seat next to how each real style does.

This is the core idea behind exploitative poker strategy: understand the theory well enough to know when and how to break it.

That second question is the one Poker Shark is built around. You practice against opponents who actually make these mistakes, each modeled on real cash game data with distinct tendencies, and you learn to find the line that punishes each one. You can still use the free Spot Calculator to check the equity math behind any decision point.

Key Takeaways

  • A poker solver computes approximate Nash equilibrium strategies using CFR
  • Solver output is conditional on the ranges, sizes, and tree you define. Change the inputs, change the answer
  • GTO is a baseline, not a playbook. It tells you what balanced looks like so you can spot when opponents deviate
  • A best response is the opposite computation: the most profitable counter-strategy to one known opponent. Exploitative play is the human version of it
  • Real profit comes from exploiting those deviations, not from memorizing solver frequencies
  • Study theory to build intuition, then apply it against real opponents who make real mistakes
  • Node locking studies one deviation at a time; everything you leave unlocked is solved as perfect play, so a few locks can recommend the wrong exploit