Calculate the missing value in the Bellman equation from reward, discount factor, and next state value to get the updated state value.

Apply the simplified one-step equation V(s) = R + ฮณV(sโ€ฒ).

Use the same reward units as the state values.

Enter a factor from 0 to 1, not a percentage.


Related Calculators

Bellman Equation Formula

The calculator uses the simplified one-step Bellman equation. It assumes one immediate reward and one next-state value.

V(s) = R + ฮณ V(s')

To solve for a missing input, the same equation is rearranged:

R = V(s) - ฮณ V(s')
ฮณ = (V(s) - R) / (V(s'))
V(s') = (V(s) - R) / (ฮณ)
  • V(s) is the updated value of the current state.
  • R is the immediate reward received after taking the action.
  • ฮณ is the discount factor, restricted to 0 through 1 in this calculator.
  • V(s’) is the value of the next state.

Select the value to solve for, enter the three displayed inputs, then select Calculate. The calculator supports Updated state value, Immediate reward, Discount factor, and Next-state value. It distinguishes conflicting inputs from cases with infinitely many solutions.

Discount Factor and Bellman Value Interpretation

Discount factor ฮณ Meaning Effect on V(s)
0 Only the immediate reward matters. V(s) = R
0.1 to 0.4 The next-state value is multiplied by 0.1 to 0.4. The actual contribution also depends on the magnitude and sign of the next-state value.
0.5 to 0.8 Immediate and future values both matter. The contribution is ฮณ times the supplied next-state value.
0.9 to 1 Future value is weighted heavily. The future-value term retains 90% to 100% of the supplied next-state value.

Selected solve mode What is calculated Required nonzero condition
Reward R R = V(s) – ฮณV(s’) None
Discount factor ฮณ ฮณ = (V(s) – R) / V(s’) V(s’) cannot be 0
Next-state value V(s’) V(s’) = (V(s) – R) / ฮณ ฮณ cannot be 0
Updated value V(s) V(s) = R + ฮณV(s’) None

Example Problems

Example 1: Calculate the updated value

Suppose the reward is 5, the discount factor is 0.9, and the next-state value is 20.

V(s) = 5 + 0.9(20)
V(s) = 5 + 18 = 23

The updated value is 23.

Example 2: Calculate the reward

Suppose the updated value is 14, the discount factor is 0.8, and the next-state value is 10.

R = 14 - 0.8(10)
R = 14 - 8 = 6

The reward is 6.

FAQ

What does the Bellman equation calculate?

The Bellman equation calculates the value of a state by combining the immediate reward with the discounted value of the next state. In this simplified one-step version, it answers: what is this state worth if you receive reward R now and then move to a next state worth V(s’)?

Why is the discount factor usually between 0 and 1?

The discount factor controls how much future value counts. A value near 0 reduces the weight on the next-state value; a value near 1 preserves most of it. Its contribution also depends on the size and sign of that value. For bounded rewards in continuing tasks, a factor below 1 supports a finite discounted return. A factor of 1 is undiscounted and requires separate assumptions, such as a terminating episode.

Can the next-state value or reward be negative?

Yes. A negative reward can represent a penalty or cost. A negative next-state value can represent a bad future state. The same formula still applies, but negative values will reduce the updated state value depending on the discount factor.