Dynamic Programming (DP) is a way to solve complex problems by breaking them into smaller, easier problems. Instead of solving the same small problems again and again, DP stores their solutions in a structure like an array, table, or map. This avoids wasting time on repeated calculations and makes the process much faster and more efficient.
The main opportunity for DP is repeated work. When different choices lead to the same subproblem, solve that state once and reuse its result. The benefit depends on having fewer distinct states than recursive calls in the naive solution.
DP needs a state definition that contains everything needed to solve a subproblem and a recurrence that combines previously solved states. Optimization problems require optimal substructure; counting and feasibility problems combine counts or Boolean answers instead. Overlapping subproblems make storing these results useful. With a valid dependency order, each state can be computed once.
Check both the recurrence and the amount of overlap before adding a cache. Reuse can save substantial work, but a large state space can still make a DP solution expensive.
This method was introduced by Richard Bellman in the 1950s and has become a valuable tool in areas like computer science, economics, and operations research. It has been used to solve problems that would otherwise take too long by turning slow, exponential-time algorithms into much faster polynomial-time solutions. DP is used in practice for tackling real-world optimization challenges.
DP appears in routing, scheduling, matching, compression, and resource allocation. In each setting, the central task is to identify which partial results can be reused without losing information needed for correctness.
Two properties explain when dynamic programming is useful, especially for optimization:
A useful mindset here is to treat DP like a story you’re building: first you define the “chapters” (subproblems), then you decide how chapters connect (transitions), and finally you make sure you never rewrite a chapter you already finished (storage).
A problem has optimal substructure when the best solution to the overall problem can be built from the best solutions to its smaller parts. In simple terms, solving the smaller pieces perfectly ensures the entire problem is solved perfectly too.
Formally, if the optimal solution
Why this matters: DP depends on trust. You’re trusting that if you make the best choice for a subproblem, you’re not accidentally ruining the final answer. If that trust holds, you can safely “lock in” sub-results. If it doesn’t, DP will confidently assemble a solution that looks logical but isn’t globally optimal.
Mathematical Representation:
Consider a problem where we want to find the optimal value
$ V(n) = \min_{1 \le k < n} f(V(k), V(n-k)), \qquad n \ge 2 $
then this is one possible form of an optimal-substructure recurrence, provided the state is sufficient and the combination accounts for every valid choice and its cost. It is not a universal DP formula: many problems need several state parameters or different transitions. Base cases must be defined separately.
Example: Shortest Path in Graphs
In the context of graph algorithms, suppose we want to find the shortest path from vertex
This example also highlights a common “do”: make sure you can clearly explain how a best solution is composed. If you can’t describe how the big answer is built from smaller best answers, your “DP solution” is probably just a recursive solution with a table taped to it.
A problem has overlapping subproblems when it can be divided into smaller problems that are solved multiple times. This happens when the same subproblem appears in different parts of the solution process. In a straightforward recursive approach, this leads to solving the same subproblem repeatedly, which wastes time and resources. Dynamic Programming addresses this by solving each subproblem once and storing the result for reuse, improving efficiency significantly.
A valid recurrence establishes correctness; repeated subproblems explain why caching helps. Without repetition, storing every intermediate result may add memory overhead without reducing the amount of computation.
Mathematical Representation:
Let
Example: Fibonacci Numbers
The recursive computation of Fibonacci numbers
Fibonacci makes the overlap easy to see: the same values are recomputed in separate recursive branches. Memoization preserves the recurrence while eliminating those repeated computations.
There are two primary methods for implementing dynamic programming algorithms:
You can think of these as two different writing styles for the same story. Memoization writes chapters only when needed (and bookmarks them). Tabulation writes every chapter in order, from page 1 onward. Both can end at the same ending; the difference is how they get there.
Memoization is a technique used to optimize recursive problem-solving by storing the results of solved subproblems in a data structure, such as a hash table or an array. When the algorithm encounters a subproblem, it first checks if the result is already stored. If it is, the algorithm simply retrieves the cached result instead of recomputing it. This approach reduces redundant calculations, making the solution process faster and more efficient.
A “do” for memoization: keep the recursion because it matches the problem’s natural definition, but make every subproblem cheap after the first time. A “don’t”: let your memo structure accidentally depend on hidden state (like mutable defaults or changing globals) in a way that makes results inconsistent.
Algorithm Steps:
- Identify the parameters that uniquely define a subproblem and create a recursive function using these parameters.
- Establish base cases to terminate recursion.
- Initialize a data structure to store computed subproblem results.
- Before computing a subproblem, check if its result is already stored.
- If not already computed, compute the subproblem's result and store it.
Example: Computing Fibonacci Numbers with Memoization
def fibonacci(n, memo=None):
if n < 0:
raise ValueError("n must be non-negative")
if memo is None:
memo = {}
if n in memo:
return memo[n]
if n <= 1:
memo[n] = n
else:
memo[n] = fibonacci(n - 1, memo) + fibonacci(n - 2, memo)
return memo[n]Analysis:
- Time Complexity:
$O(n)$ , since each number up to$n$ is computed once. - Space complexity:
$O(n)$ stored values and an$O(n)$ recursion stack. Python's recursion limit can prevent this version from handling large$n$ even when sufficient heap memory is available.
These Fibonacci examples assume integer
One practical note when learning: top-down DP is often easier to write correctly first, because you start from the real question (“solve
Tabulation is a technique that solves a problem by working from the smallest subproblems up to the larger ones in a step-by-step, iterative way. It stores the solutions to these subproblems in a table, often an array, and uses these stored values to solve bigger subproblems until the final solution is reached. Unlike memoization, which works recursively and checks if a solution is already computed, tabulation systematically builds the solution from the ground up. This method is efficient and avoids the overhead of recursive calls.
The “why” behind tabulation is control: you decide the exact order states are computed, and you avoid recursion limits and call overhead. Fill the table in dependency order. Do not fill it in a convenient order that violates dependencies, because then you’ll be reading values that aren’t valid yet.
Algorithm Steps:
- Initialize the table by setting up a structure to store the results of subproblems, and ensure that base cases are properly initialized to handle the simplest instances of the problem.
- Iterative computation is carried out using loops to fill the table, making sure that each subproblem is solved in the correct order before being used to solve larger subproblems.
- Construct the solution by referencing the filled table, using the stored values to derive the final solution to the original problem efficiently.
Example: Computing Fibonacci Numbers with Tabulation
def fibonacci(n):
if n < 0:
raise ValueError("n must be non-negative")
if n <= 1:
return n
fib_table = [0] * (n + 1)
fib_table[0], fib_table[1] = 0, 1
for i in range(2, n + 1):
fib_table[i] = fib_table[i - 1] + fib_table[i - 2]
return fib_table[n]Analysis:
- The time complexity is
$O(n)$ , since the algorithm iterates from 2 to$n$ , computing each Fibonacci number sequentially. - The space complexity is
$O(n)$ , due to the storage required for the table that holds the Fibonacci numbers up to$n$ .
A nice way to make tabulation feel natural is to ask: “What’s the smallest thing I must know before I can know the next thing?” That question basically is the loop order.
| Aspect | Memoization (Top-Down) | Tabulation (Bottom-Up) |
|---|---|---|
| Approach | Recursive | Iterative |
| Storage | Stores solutions as needed | Pre-fills table with solutions to all subproblems |
| Overhead | Function call overhead due to recursion | Minimal overhead due to iteration |
| Flexibility | May be easier to implement for complex recursive problems | May require careful ordering of computations |
| Space Efficiency | Potentially higher due to recursion stack and memoization | Can be more space-efficient with careful table design |
If you’re choosing between them, a simple rule of thumb is: start with memoization to get the recurrence right, then switch to tabulation when you want tighter performance or cleaner memory control.
Dynamic programming problems are often formulated using recurrence relations, which express the solution to a problem in terms of its subproblems.
A recurrence specifies how the answer for a state depends on other states. Define that relationship, its base cases, and a valid evaluation order before deciding how to store the results.
Example: Longest Common Subsequence (LCS)
Given two sequences
Implementation:
We can implement the LCS problem using either memoization or tabulation. With tabulation, we build a two-dimensional table
- The time complexity is
$O(mn)$ , as the algorithm processes a grid or matrix of size$m \times n$ , iterating through each cell. - The space complexity is
$O(mn)$ , due to the table storing intermediate results, but this can be reduced to$O(n)$ by optimizing the storage to only keep necessary data for the current and previous rows.
The reason LCS is such a beloved DP example is that the “choices” are easy to explain: if characters match, you take the diagonal; if they don’t, you take the best of left/top. That clear decision structure is exactly what good DP feels like: simple local rules that reliably build a global answer.
Properly defining the state is crucial for dynamic programming.
- State variables are the parameters that uniquely define each subproblem, helping to break down the problem into smaller, manageable components.
- State transition refers to the rules or formulas that describe how to move from one state to another, typically using the results of smaller subproblems to solve larger ones.
Think of state design like labeling drawers in a workshop. If the labels are precise, you can find what you need instantly and build bigger things confidently. If the labels are vague, you’ll keep opening drawers, guessing, and making mistakes, even if the math is technically “there.”
Example: 0/1 Knapsack Problem
- Assume positive integer weights, non-negative integer capacity, and at most one copy of each item. Initialize
$dp[0][w]=0$ and$dp[i][0]=0$ . - The problem statement focuses on selecting
$n$ items, each with a weight$w_i$ and value$v_i$ , while ensuring the total weight stays within the knapsack capacity$W$ , in order to maximize the total value. - In state representation,
$dp[i][w]$ represents the maximum value that can be achieved using the first$i$ items with a total weight capacity of$w$ . - State Transition:
Implementation:
We fill the table
- The time complexity is
$O(nW)$ , where$n$ is the number of items and$W$ is the capacity of the knapsack, as the algorithm iterates through both items and weights. - The space complexity is
$O(nW)$ , but this can be optimized to$O(W)$ because each row in the table depends only on the values from the previous row, allowing for space reduction.
The
A key “do” here is to interpret your state in plain language. If you can’t say what dp[i][w] means in one sentence, debugging will be painful. A key “don’t” is mixing meanings (for example, letting w sometimes mean “remaining capacity” and sometimes mean “used weight”), that’s how DP tables become nonsense.
In some cases, we can optimize space complexity by noticing dependencies between states.
This is where DP graduates from “works” to “works well.” Once the recurrence is correct, you can ask: “Do I truly need all previous states, or only a slice of them?” Many DP solutions only depend on the previous row, previous column, or a small window, so storing everything is optional.
Example: Since
def knapsack(weights, values, capacity):
best = [0] * (capacity + 1)
for weight, value in zip(weights, values):
for remaining in range(capacity, weight - 1, -1):
best[remaining] = max(
best[remaining], best[remaining - weight] + value
)
return best[capacity]Here weights and values must have the same length, weights must be positive integers, and capacity must be a non-negative integer. Keeping only values is enough to return the optimum; reconstructing the chosen items requires additional predecessor information or recomputation.
One important “do/don’t” hidden in this snippet is the reverse loop. Updating from high to low prevents using the same item multiple times in a single iteration. If you go forward, you silently switch the problem you’re solving.
If a problem has optimal substructure but does not have overlapping subproblems, it is often better to use Divide and Conquer instead of Dynamic Programming. Divide and Conquer works by breaking the problem into independent subproblems, solving each one separately, and then combining their solutions. Since there are no repeated subproblems to reuse, storing intermediate results (as in Dynamic Programming) is unnecessary, making Divide and Conquer a more suitable and efficient choice in such cases.
This distinction matters because DP isn’t “better recursion”, it’s “recursion plus reuse.” If there’s nothing to reuse, DP is extra work for no gain. Picking the right paradigm is part of writing efficient algorithms, not just correct ones.
Example: Merge Sort algorithm divides the list into halves, sorts each half, and then merges the sorted halves.
These terms show up constantly around DP because DP is really about breaking a big idea into structured pieces. The more comfortable you are with these building blocks, the faster DP problems start to feel like patterns instead of puzzles.
A process is called recursion when a function solves a problem by calling itself, either directly or indirectly, with a smaller instance of the same problem. This continues until the function reaches a base case, which is a condition that stops further recursive calls and provides a straightforward solution. Recursion is useful for problems that can naturally be divided into similar smaller subproblems.
Mathematical Perspective:
A recursive function
for some function
Example: Computing
with base case
In DP, recursion is often your first draft: it expresses the logic cleanly. Then DP adds the missing ingredient: remembering what you already solved.
For a set
Mathematical Properties:
- The total subsets of a set with
$n$ elements is$2^n$ , as each element can either be included or excluded from a subset. - The power set of a set
$S$ , denoted as$\mathcal{P}(S)$ , is the set of all possible subsets of$S$ , including the empty set and$S$ itself.
Relevance to DP:
Subsets often represent different states or configurations in combinatorial problems, such as the subset-sum problem.
The “why you should care” is right in the number
A contiguous segment of an array
Mathematical Representation:
Example:
Given
Relevance to DP:
Subarray problems include finding the maximum subarray sum (Kadane's algorithm), where dynamic programming efficiently computes optimal subarrays.
Subarrays matter in DP because contiguity gives you a natural order, perfect for transitions. When a problem is about “best segment,” “best window,” or “best range,” DP patterns show up immediately.
A contiguous sequence of characters within a string
Mathematical Representation:
A substring
with
Example:
For
Relevance to DP:
Substring problems include finding the longest palindromic substring or the longest common substring between two strings.
Strings are a DP playground because “prefixes” and “ends at i/j” are easy states to define. If you’ve ever seen a 2D DP table with characters along the top and side, you’ve met this world.
A sequence derived from another sequence by deleting zero or more elements without changing the order of the remaining elements.
Mathematical Representation:
Given sequence
where
Example:
For
Relevance to DP:
The Longest Common Subsequence (LCS) problem is a classic dynamic programming problem.
LCS Dynamic Programming Formulation:
Let
Recurrence Relation:
Implementation:
We build a two-dimensional table
Time Complexity:
Subsequences are where DP really earns its reputation, because “skip or take” decisions can branch exponentially. DP keeps that branching logically, but prevents it from becoming computational chaos.
This section is your pattern-matching toolkit. DP gets dramatically easier once you stop trying to “invent DP” every time and instead learn to recognize the signals that a table of reused sub-results will pay off.
I. If the problem asks for the number of ways to do something:
- Counting paths in a grid.
- DP avoids enumerating every route. An obstacle-free rectangular grid also has a direct binomial-coefficient formula.
II. If the task is to find the minimum or maximum value under constraints:
- Knapsack problem.
- Subset enumeration is a baseline; branch-and-bound and other methods may also apply.
III. If the same inputs appear again during recursion:
- Fibonacci numbers.
- Without DP, Fibonacci numbers would be recomputed many times.
IV. If the solution depends on both the current step and remaining resources (time, weight, money, length):
- Scheduling tasks within a time limit.
- A state can capture remaining time, although the exact scheduling constraints determine whether DP, greedy selection, or another method is appropriate.
V. If the problem works with prefixes, substrings, or subsequences:
- Longest common subsequence.
- A naive recursion explores exponentially many choices; caching prefix states removes repeated work.
VI. If choices at each step must be explored and combined carefully:
- Coin change with mixed denominations.
- DP guarantees the fewest coins for arbitrary positive integer denominations when its recurrence is correct. Exhaustive search can also guarantee optimality, while greedy works only for suitable coin systems.
VII. If the state space can be stored in a table or array:
- Problems with discrete states.
- A finite state space makes direct tabulation possible. Continuous-state problems need additional structure, a different representation, or an approximation; an array is not a requirement for every DP formulation.
A good “do” when scanning a problem is to ask: “Can I describe a state with a small number of integers (like i, j, w)?” If yes, DP is often on the table. A good “don’t” is jumping into DP just because the problem is hard, hard problems also show up in greedy, graph, and divide-and-conquer territory.
- A well-chosen state defines what each subproblem represents, while a poorly chosen one leaves the formulation incomplete; for example,
dp[i][w]in the knapsack problem captures value usingiitems and capacityw. - A correct transition connects states consistently, while skipping this leads to undefined progress; in knapsack, the choice to include or exclude an item gives the formula for moving between states.
If DP ever feels “mysterious,” it’s usually because the state meaning isn’t crisp. Once the state is clear, the transition almost writes itself: you ask what decision moves you from smaller states to bigger ones, and you encode that decision as a recurrence.
- Reducing memory usage by discarding unnecessary states makes solutions efficient, while failing to do so can waste resources; for example, knapsack space can shrink from
O(nW)toO(W)with a one-dimensional array. - Using pruning to skip impossible paths speeds up computation, while omitting it allows redundant work; in recursive search with memoization, a branch can be ignored only when a proven bound shows it cannot improve the answer. Do not cache a partially explored or pruned result as an exact state value if it depends on the current global bound.
The key “why” here is that DP gives you structure, but structure can still be expensive. Optimization is about keeping the structure while trimming what you don’t truly need, whether that’s table size, transitions, or states that can never be reached.
I. Failure to Define Proper Base Cases
- Example: In grid path counting, omitting
dp[0][0] = 1prevents any valid paths from being constructed. - Consequence: Without correct starting values, the DP table propagates errors and produces incorrect results.
A simple “do”: before filling anything, write down the smallest cases and ensure they make sense. A simple “don’t”: assume the base case is “obvious” and skip it, DP will punish that immediately.
II. Updating States in the Wrong Dependency Order
- Example: In knapsack with a 1D array, iterating weights from low to high causes items to be reused multiple times.
- Consequence: Using the wrong order inflates computed values and leads to invalid or impossible solutions.
The ordering rule is not a style preference, it’s part of correctness. If your transition reads from states you’ve already updated in the same iteration, you may be solving a different problem than you think.
III. Ignoring Special or Edge Case Inputs
- Example: In knapsack, a zero-capacity input should return zero value rather than throwing an error.
- Consequence: Overlooking edge inputs causes crashes or incorrect answers in boundary conditions.
A good “do” is to test the edges early: zeros, ones, empty inputs, minimal sizes. DP solutions can look perfect on “normal” cases while quietly breaking on the boundaries where states and base cases are most exposed.
For the sum and coin problems below, assume a non-negative integer target and positive integer choices. Zero or negative choices can create nonterminating recurrences when reuse is allowed. For string construction, use nonempty word-bank entries; words may be reused, and their order in the concatenation matters. Enumerating every construction remains output-sensitive even with memoization.
The Fibonacci sequence is a series of numbers where each number is the sum of the two preceding ones, starting with 0 and 1. This sequence is a classic example used to demonstrate recursive algorithms and dynamic programming techniques.
The Grid Traveler problem involves finding the total number of ways to traverse an m x n grid from the top-left corner to the bottom-right corner, with the constraint that you can only move right or down.
The Climbing Stairs problem requires determining the number of distinct ways to reach the top of a staircase with 'n' steps, given that you can climb either 1, 2, or 3 steps at a time.
The Can Sum problem involves determining if it is possible to achieve a target sum using any number of elements from a given list of numbers. Each number in the list can be used multiple times.
The How Sum problem extends the Can Sum problem by identifying which elements from the list can be combined to sum up to the target value.
The Best Sum problem further extends the How Sum problem by finding the smallest combination of numbers that add up to exactly the target sum.
The Can Construct problem involves determining if a target string can be constructed from a given list of substrings.
The Count Construct problem expands on the Can Construct problem by determining the number of ways a target string can be constructed using a list of substrings.
The All Constructs problem is a variation of the Count Construct problem, which identifies all the possible combinations of substrings from a list that can be used to form the target string.
The Coins problem aims to find the minimum number of coins needed to make a given value, provided an infinite supply of each coin denomination.
The Longest Common Subsequence problem involves finding the longest subsequence that two sequences have in common, where a subsequence is derived by deleting some or no elements without changing the order of the remaining elements.
The Longest Increasing Subarray problem involves identifying the longest contiguous subarray where the elements are strictly increasing.
The Knuth-Morris-Pratt (KMP) algorithm is a pattern searching algorithm that looks for occurrences of a "word" within a main "text string" using preprocessing over the pattern to achieve linear time complexity. It is primarily a string-matching algorithm; it appears here because its prefix table reuses information from shorter prefixes.
This problem involves finding the minimum number of insertions needed to transform a given string into a palindrome. The goal is to make the string read the same forwards and backwards with the fewest insertions possible.