From 6b68ff4ebdc6ea032bf8249f11156308397245cf Mon Sep 17 00:00:00 2001 From: Adam Djellouli Date: Wed, 23 Sep 2026 00:52:53 +0200 Subject: [PATCH] Clarify and reorganize algorithm documentation Refresh the repository documentation so readers can move from fundamentals to algorithmic techniques through a consistent, navigable set of notes. Normalize heading levels, list styles, emphasis, equations, tables, and code examples across the algorithm guides. Rewrite explanations of recursion, backtracking, graph traversal, and sorting to distinguish related concepts, state assumptions, and keep the narrative focused on the algorithmic invariant. Correct recurrence notation, edge-case handling, stability claims, complexity bounds, auxiliary-space descriptions, and implementation guidance where the previous text was ambiguous or misleading. Expand the sorting comparison material to cover the revised merge, quick, heap, radix, and counting sort guidance with clearer constraints and terminology. Update the root README with a progressive reading order that links directly to the notes, and remove stale presentation artifacts while preserving the existing implementation references. --- README.md | 21 +- notes/backtracking.md | 445 ++++++------- notes/basic_concepts.md | 284 ++++---- notes/brain_teasers.md | 837 ++++++++++++------------ notes/data_structures.md | 425 ++++++------ notes/dynamic_programming.md | 407 ++++++------ notes/graphs.md | 1117 +++++++++++++++----------------- notes/greedy_algorithms.md | 484 +++++++------- notes/math_set_relationship.md | 184 +++--- notes/matrices.md | 344 +++++----- notes/searching.md | 886 ++++++++++++------------- notes/sorting.md | 602 +++++++++-------- 12 files changed, 3055 insertions(+), 2981 deletions(-) diff --git a/README.md b/README.md index a7a0d9f..d8ea47b 100644 --- a/README.md +++ b/README.md @@ -125,13 +125,19 @@ This command formats all Python files in the current directory and its subdirect ## Notes -* Basic concepts. -* Data structures. -* Graph algorithms. -* Backtracking. -* Dynamic programming. -* Sorting. -* Brain teasers. +The notes build from fundamentals to algorithmic techniques. A suggested reading order is: + +1. [Basic concepts](notes/basic_concepts.md) +2. [Sets, combinations, and permutations](notes/math_set_relationship.md) +3. [Data structures](notes/data_structures.md) +4. [Searching](notes/searching.md) +5. [Sorting](notes/sorting.md) +6. [Graph algorithms](notes/graphs.md) +7. [Backtracking](notes/backtracking.md) +8. [Dynamic programming](notes/dynamic_programming.md) +9. [Greedy algorithms](notes/greedy_algorithms.md) +10. [Matrices and grids](notes/matrices.md) +11. [Programming brain teasers](notes/brain_teasers.md) ## List of projects @@ -646,4 +652,3 @@ This project is licensed under the [MIT License](LICENSE) - see the LICENSE file ## Star History [![Star History Chart](https://api.star-history.com/svg?repos=djeada/Algorithms-And-Data-Structures&type=Date)](https://star-history.com/#djeada/Algorithms-And-Data-Structures&Date) - diff --git a/notes/backtracking.md b/notes/backtracking.md index 8007893..8e3976c 100644 --- a/notes/backtracking.md +++ b/notes/backtracking.md @@ -1,32 +1,32 @@ -## Backtracking +# Backtracking -Backtracking is a method used to solve problems by building potential solutions step by step. If it becomes clear that a partial solution cannot lead to a valid final solution, the process "backtracks" by undoing the last step and trying a different path. This approach is commonly applied to **constraint satisfaction problems**, **combinatorial optimization**, and puzzles like **N-Queens** or **Sudoku**, where all possibilities need to be explored systematically while avoiding unnecessary computations. +Backtracking is a method used to solve problems by building potential solutions step by step. If it becomes clear that a partial solution cannot lead to a valid final solution, the process "backtracks" by undoing the last step and trying a different path. This approach is commonly applied to constraint satisfaction problems, combinatorial optimization, and puzzles like N-Queens or Sudoku, where all possibilities need to be explored systematically while avoiding unnecessary computations. -Think of backtracking as “explore with a rewind button.” You move forward making choices, but you’re never afraid to erase and try again. The reason this matters is simple: brute force tries *everything* blindly, but backtracking tries to be smart about quitting early. The moment a choice breaks the rules (a constraint), you stop investing time in that branch. That “stop early” habit is the real superpower, it keeps problems that *could* explode into millions of possibilities from wasting your time on guaranteed failures. +Backtracking explores choices depth first and rejects a branch as soon as its partial solution violates a constraint. Undoing the last choice restores the state for the next candidate. Early rejection reduces wasted exploration, though the worst-case search can still be exponential. -A good mental model is: **choose → test → continue or undo**. The “test” part is what turns random exploration into an algorithm with structure. +A good mental model is: choose → test → continue or undo. The “test” part is what turns random exploration into an algorithm with structure. -### Recursive Functions +## Recursive Functions Recursive functions are functions that call themselves directly or indirectly to solve a problem by breaking it down into smaller, more manageable subproblems. This concept is fundamental in computer science and mathematics, as it allows for elegant solutions to complex problems through repeated application of a simple process. -Recursion shows up here because backtracking naturally creates “nested decisions.” Every time you make a choice, you enter a smaller version of the same problem: “Okay, given what I’ve chosen so far, what choices can I make next?” Recursion fits that pattern perfectly. The big win is that recursion keeps the code focused on the **current step**, while the call stack quietly remembers how you got there, so when you backtrack, you return to the exact point where the last decision was made. +Recursion shows up here because backtracking naturally creates “nested decisions.” Every time you make a choice, you enter a smaller version of the same problem: “Okay, given what I’ve chosen so far, what choices can I make next?” Recursion fits that pattern perfectly. The big win is that recursion keeps the code focused on the current step, while the call stack quietly remembers how you got there, so when you backtrack, you return to the exact point where the last decision was made. -That said, recursion is not “magic.” It’s a disciplined way of saying: *solve the small version, then build up*. If you don’t have a clear stop condition, recursion doesn’t politely fail, it loops until you run out of stack. +Each recursive call needs a base case or a demonstrable move toward one. Without that progress, calls continue until a resource limit is reached. Main idea: -1. **Base Case (Termination Condition)** is the condition under which the recursion stops. It prevents infinite recursion by providing an explicit solution for the simplest instance of the problem. -2. **Recursive Case** is the part of the function where it calls itself with a modified parameter, moving towards the base case. +1. Base Case (Termination Condition) is the condition under which the recursion stops. It prevents infinite recursion by providing an explicit solution for the simplest instance of the problem. +2. Recursive Case is the part of the function where it calls itself with a modified parameter, moving towards the base case. A useful way to keep recursion readable is to ask two questions while writing it: -* **When do I stop?** (base case) -* **How do I make the problem smaller?** (recursive case) +- When do I stop? (base case) +- How do I make the problem smaller? (recursive case) If you can answer those cleanly, you usually get clean code. -#### Mathematical Foundation +### Mathematical Foundation Recursion closely relates to mathematical induction, where a problem is solved by assuming that the solution to a smaller instance of the problem is known and building upon it. @@ -34,25 +34,25 @@ A recursive function can often be expressed using a recurrence relation: $$ f(n) = \begin{cases} -g(n) & \text{if } n = \text{base case} \ +g(n) & \text{if } n = \text{base case} \\ h(f(n - 1), n) & \text{otherwise} \end{cases} $$ where: -* $g(n)$ is the base case value, -* $h$ is a function that defines how to build the solution from the smaller instance. +- $g(n)$ is the base case value, +- $h$ is a function that defines how to build the solution from the smaller instance. This is where recursion becomes more than a programming trick, it becomes a way to *prove* correctness. Induction and recursion share the same shape: handle the simplest case, then show how everything else reduces to it. If you ever want confidence that a recursive solution is correct, that structure is what you lean on. -#### Example: Calculating Factorial +### Example: Calculating Factorial The factorial of a non-negative integer $n$ is the product of all positive integers less than or equal to $n$. Mathematically, it is defined as: $$ n! = \begin{cases} -1 & \text{if } n = 0 \ +1 & \text{if } n = 0 \\ n \times (n - 1)! & \text{if } n > 0 \end{cases} $$ @@ -61,26 +61,28 @@ Python Implementation: ```python def factorial(n): + if n < 0: + raise ValueError("n must be non-negative") if n == 0: return 1 # Base case: 0! = 1 else: return n * factorial(n - 1) # Recursive case ``` -Factorial is the “hello world” of recursion because it’s easy to see the shrinking problem: `factorial(n)` becomes `factorial(n-1)`. And it also shows the two phases you always get with recursion: **going down** (making calls) and **coming back up** (combining results). That “coming back up” part is basically backtracking in miniature: you return from deep calls and rebuild the final answer step by step. +Factorial is the “hello world” of recursion because it’s easy to see the shrinking problem: `factorial(n)` becomes `factorial(n-1)`. And it also shows the two phases you always get with recursion: going down (making calls) and coming back up (combining results). Returning through the call stack is called unwinding. It is useful preparation for backtracking, but factorial is not a backtracking search: it does not choose among alternatives or undo tentative decisions. -##### Detailed Computation for $n = 5$ +#### Detailed Computation for $n = 5$ Let's trace the recursive calls for `factorial(5)`: -* Call `factorial(5)` and compute $5 \times factorial(4)$ because $5 \neq 0$. -* Call `factorial(4)` and compute $4 \times factorial(3)$. -* Call `factorial(3)` and compute $3 \times factorial(2)$. -* Call `factorial(2)` and compute $2 \times factorial(1)$. -* Call `factorial(1)` and compute $1 \times factorial(0)$. -* Call `factorial(0)` and return $1$ as the base case is reached. +- Call `factorial(5)` and compute $5 \times factorial(4)$ because $5 \neq 0$. +- Call `factorial(4)` and compute $4 \times factorial(3)$. +- Call `factorial(3)` and compute $3 \times factorial(2)$. +- Call `factorial(2)` and compute $2 \times factorial(1)$. +- Call `factorial(1)` and compute $1 \times factorial(0)$. +- Call `factorial(0)` and return $1$ as the base case is reached. -Now, we backtrack and compute the results: +Now the calls unwind and compute the results: 1. `factorial(1)` returns $1 \times 1 = 1$. 2. `factorial(2)` returns $2 \times 1 = 2$. @@ -90,9 +92,9 @@ Now, we backtrack and compute the results: Thus, $5! = 120$. -Notice how the “real math” happens on the way back. On the way down, you’re basically stacking up IOUs: “I’ll finish factorial(5) once I know factorial(4).” Backtracking is the same vibe: explore deep, then unwind and try something else if needed. +The downward phase creates pending multiplications; the return phase completes them. Backtracking uses the same call-stack behavior, with the additional work of undoing a choice and trying another candidate. -#### Visualization with Recursion Tree +### Visualization with Recursion Tree Each recursive call can be visualized as a node in a tree: @@ -112,15 +114,15 @@ factorial(5) The leaves represent the base case, and the tree unwinds as each recursive call returns. -**Important Considerations:** +Important Considerations: -* When using recursion, ensure **termination** by designing the recursive function such that all possible paths eventually reach a base case. This prevents infinite recursion. -* Be mindful of **stack depth**, as each recursive call adds a new frame to the call stack. Too many recursive calls, especially in deep recursion, can result in a stack overflow error. -* Consider **efficiency** when choosing a recursive approach. While recursive solutions can be elegant and clean, they may not always be optimal in terms of time and space, particularly when dealing with large input sizes or deep recursive calls. +- When using recursion, ensure termination by designing the recursive function such that all possible paths eventually reach a base case. This prevents infinite recursion. +- Be mindful of stack depth, as each recursive call adds a new frame to the call stack. Too many recursive calls, especially in deep recursion, can result in a stack overflow error. +- Consider efficiency when choosing a recursive approach. While recursive solutions can be elegant and clean, they may not always be optimal in terms of time and space, particularly when dealing with large input sizes or deep recursive calls. -These points matter even more once you move from “mathy recursion” (like factorial) to “search recursion” (like backtracking). Search trees can get deep *and* wide. That’s why good backtracking code doesn’t just recurse, it also **prunes** aggressively and keeps state management clean, so you aren’t blowing up time or memory. +These points matter even more once you move from “mathy recursion” (like factorial) to “search recursion” (like backtracking). Search trees can get deep and wide. That’s why good backtracking code doesn’t just recurse, it also prunes aggressively and keeps state management clean, so you aren’t blowing up time or memory. -### Depth-First Search (DFS) +## Depth-First Search (DFS) Depth-First Search is an algorithm for traversing or searching tree or graph data structures. It starts at a selected node and explores as far as possible along each branch before backtracking. @@ -128,20 +130,20 @@ DFS is the bridge between “recursion as a technique” and “backtracking as Main idea: -* The **traversal strategy** of Depth-First Search (DFS) involves exploring each branch of a graph or tree to its deepest point before backtracking to explore other branches. -* **Implementation** of DFS can be achieved either through recursion, which implicitly uses the call stack, or by using an explicit stack data structure to manage the nodes. -* **Applications** of DFS include tasks such as topological sorting, identifying connected components in a graph, solving puzzles like mazes, and finding paths in trees or graphs. +- The traversal strategy of Depth-First Search (DFS) involves exploring each branch of a graph or tree to its deepest point before backtracking to explore other branches. +- Implementation of DFS can be achieved either through recursion, which implicitly uses the call stack, or by using an explicit stack data structure to manage the nodes. +- Applications of DFS include tasks such as topological sorting, identifying connected components in a graph, solving puzzles like mazes, and finding paths in trees or graphs. A quick “do and don’t” that makes DFS feel less abstract: -* **Do** mark visited nodes when the graph can loop, or you’ll walk in circles forever. -* **Don’t** assume DFS finds the shortest path, it finds *a* path and explores deeply, not optimally. +- Do mark visited nodes when the graph can loop, or you’ll walk in circles forever. +- Don’t assume DFS finds the shortest path, it finds a path and explores deeply, not optimally. -#### Algorithm Steps +### Algorithm Steps -* **Start at the root node** by marking it as visited to prevent revisiting it during the traversal. -* **Explore each branch** by recursively performing DFS on each unvisited neighbor, diving deeper into the graph or tree structure. -* **Backtrack** once all neighbors of a node are visited, returning to the previous node to continue exploring other branches. +- Start at the root node by marking it as visited to prevent revisiting it during the traversal. +- Explore each branch by recursively performing DFS on each unvisited neighbor, diving deeper into the graph or tree structure. +- Backtrack once all neighbors of a node are visited, returning to the previous node to continue exploring other branches. Pseudocode: @@ -153,7 +155,7 @@ DFS(node): DFS(neighbor) ``` -#### Example: Tree Traversal +### Example: Tree Traversal Consider the following tree: @@ -168,11 +170,11 @@ Tree: Traversal using DFS starting from node 'A': -* **Visit 'A'** to begin the traversal, marking it as visited. -* **Move to 'B'**, but since 'B' has no unvisited neighbors, **backtrack to 'A'** to explore other branches. -* **Move to 'C'**, continuing the traversal to the next unvisited node. -* **Move to 'D'**, but as 'D' has no unvisited neighbors, **backtrack to 'C'**. -* **Move to 'E'**, but since 'E' also has no unvisited neighbors, **backtrack to 'C'**, and then further **backtrack to 'A'** to complete the exploration. +- Visit 'A' to begin the traversal, marking it as visited. +- Move to 'B', but since 'B' has no unvisited neighbors, backtrack to 'A' to explore other branches. +- Move to 'C', continuing the traversal to the next unvisited node. +- Move to 'D', but as 'D' has no unvisited neighbors, backtrack to 'C'. +- Move to 'E', but since 'E' also has no unvisited neighbors, backtrack to 'C', and then further backtrack to 'A' to complete the exploration. Traversal order: $A → B → C → D → E$ @@ -209,33 +211,33 @@ dfs(node_a) Analysis: -* The **time complexity** of Depth-First Search (DFS) is $O(V + E)$, where $V$ represents the number of vertices and $E$ represents the number of edges in the graph. -* The **space complexity** is $O(V)$, primarily due to the space used by the recursion stack or an explicit stack, as well as the memory required for tracking visited nodes. +- The time complexity of Depth-First Search (DFS) is $O(V + E)$, where $V$ represents the number of vertices and $E$ represents the number of edges in the graph. +- The space complexity is $O(V)$, primarily due to the space used by the recursion stack or an explicit stack, as well as the memory required for tracking visited nodes. If you care about performance, the key takeaway is: DFS is often cheap enough to be a default traversal tool, but it can still get expensive in huge graphs. Also, that `visited` flag is doing a lot of work, without it, graphs with cycles can turn DFS into an accidental infinite adventure. -#### Applications +### Applications -* **Cycle detection** in directed and undirected graphs. -* **Topological sorting** in directed acyclic graphs (DAGs). -* **Solving mazes and puzzles** by exploring all possible paths. -* Identifying **connected components** in a network or graph. +- Cycle detection in directed and undirected graphs. +- Topological sorting in directed acyclic graphs (DAGs). +- Solving mazes and puzzles by exploring all possible paths. +- Identifying connected components in a network or graph. -### Backtracking +## Backtracking Backtracking is an algorithmic technique for solving problems recursively by trying to build a solution incrementally, removing solutions that fail to satisfy the constraints at any point. -At this point, you can think of backtracking as DFS with standards. DFS explores; backtracking explores *while enforcing rules* and cutting off bad paths early. That’s why people use it for puzzles and constraint problems: you don’t just want to wander, you want to wander with a checklist, so you can confidently say “this can’t possibly work” and move on fast. +At this point, you can think of backtracking as DFS with standards. DFS explores; backtracking explores while enforcing rules and cutting off bad paths early. That’s why people use it for puzzles and constraint problems: you don’t just want to wander, you want to wander with a checklist, so you can confidently say “this can’t possibly work” and move on fast. Main Idea: -* Building solutions one piece at a time and evaluating them against the constraints. -* Early detection of invalid solutions to prune the search space. -* When a partial solution cannot be extended to a complete solution, the algorithm backtracks to try different options. +- Building solutions one piece at a time and evaluating them against the constraints. +- Early detection of invalid solutions to prune the search space. +- When a partial solution cannot be extended to a complete solution, the algorithm backtracks to try different options. The “prune the search space” is the reason backtracking is usable. Without pruning, N-Queens becomes “try everything,” and “everything” grows ridiculously fast. With pruning, you still explore possibilities, but you stop feeding time into branches that are already doomed. -#### General Algorithm Framework +### General Algorithm Framework 1. Understand the possible configurations of the solution. 2. Start with an empty solution. @@ -251,7 +253,7 @@ General Template (pseudocode) function backtrack(partial): if is_complete(partial): handle_solution(partial) - return // or continue if looking for all solutions + return for candidate in generate_candidates(partial): if is_valid(candidate, partial): @@ -262,13 +264,13 @@ function backtrack(partial): Pieces you supply per problem: -* `is_complete`: does `partial` represent a full solution? -* `handle_solution`: record/output the solution. -* `generate_candidates`: possible next choices given current partial. -* `is_valid`: pruning test to reject infeasible choices early. -* `place` / `unplace`: apply and revert the choice. +- `is_complete`: does `partial` represent a full solution? +- `handle_solution`: record/output the solution. +- `generate_candidates`: possible next choices given current partial. +- `is_valid`: pruning test to reject infeasible choices early. +- `place` / `unplace`: apply and revert the choice. -The beauty of this template is that it teaches you what backtracking really is: not one specific algorithm, but a reusable *shape*. Most backtracking problems differ only in those helper functions. If you can clearly define “what counts as valid so far,” you can solve a lot of classic puzzles with the same skeleton. +The beauty of this template is that it teaches you what backtracking really is: not one specific algorithm, but a reusable shape. Most backtracking problems differ only in those helper functions. If you can clearly define “what counts as valid so far,” you can solve a lot of classic puzzles with the same skeleton. Python-ish Generic Framework @@ -288,30 +290,30 @@ def backtrack(partial, is_complete, generate_candidates, is_valid, handle_soluti partial.pop() ``` -You can wrap those callbacks into a class or closures for stateful problems. +You can wrap those callbacks into a class or closures for stateful problems. Returning from a complete solution still allows the caller to try the next candidate, so this template enumerates all solutions. To stop at the first solution, return a success flag through every call. If `handle_solution` stores a mutable list, store a copy; otherwise later undo operations alter the recorded result. One practical “do/don’t” with this style: -* **Do** make `place/unplace` symmetrical (every change you make must be undone). -* **Don’t** mutate shared state without undoing it, or you’ll get “ghost effects” where one branch pollutes another. +- Do make `place/unplace` symmetrical (every change you make must be undone). +- Don’t mutate shared state without undoing it, or you’ll get “ghost effects” where one branch pollutes another. -#### N-Queens Problem +### N-Queens Problem The N-Queens problem is a classic puzzle in which the goal is to place $N$ queens on an $N \times N$ chessboard such that no two queens threaten each other. In chess, a queen can move any number of squares along a row, column, or diagonal. Therefore, no two queens can share the same row, column, or diagonal. Objective: -* Place $N$ queens on the board. -* Ensure that no two queens attack each other. -* Find all possible arrangements that satisfy the above conditions. +- Place $N$ queens on the board. +- Ensure that no two queens attack each other. +- Find all possible arrangements that satisfy the above conditions. -N-Queens is the poster child for backtracking because it’s easy to describe, hard to brute force, and perfect for pruning. The moment you place a queen, you instantly rule out a bunch of squares. That means you don’t need to “wait until the end” to discover failure, you can detect it *as soon as it happens*, which is exactly what backtracking wants. +N-Queens is the poster child for backtracking because it’s easy to describe, hard to brute force, and perfect for pruning. The moment you place a queen, you instantly rule out a bunch of squares. That means you don’t need to “wait until the end” to discover failure, you can detect it as soon as it happens, which is exactly what backtracking wants. -##### Visual Representation +#### Visual Representation To better understand the problem, let's visualize it using ASCII graphics. -**Empty $4 \times 4$ Chessboard:** +Empty $4 \times 4$ Chessboard: ``` 0 1 2 3 (Columns) @@ -329,34 +331,34 @@ To better understand the problem, let's visualize it using ASCII graphics. Each cell can be identified by its row and column indices ((row, column)). -**Example Solution for $N = 4$:** +Example Solution for $N = 4$: One of the possible solutions for placing 4 queens on a $4 \times 4$ chessboard is: ``` 0 1 2 3 (Columns) +---+---+---+---+ -0| Q | | | | (Queen at position (0, 0)) +0| | Q | | | (Queen at position (0, 1)) +---+---+---+---+ -1| | | Q | | (Queen at position (1, 2)) +1| | | | Q | (Queen at position (1, 3)) +---+---+---+---+ -2| | | | Q | (Queen at position (2, 3)) +2| Q | | | | (Queen at position (2, 0)) +---+---+---+---+ -3| | Q | | | (Queen at position (3, 1)) +3| | | Q | | (Queen at position (3, 2)) +---+---+---+---+ (Rows) ``` -* `Q` represents a queen. -* Blank spaces represent empty cells. +- `Q` represents a queen. +- Blank spaces represent empty cells. -##### Constraints +#### Constraints -* Only one queen per row. -* Only one queen per column. -* No two queens share the same diagonal. +- Only one queen per row. +- Only one queen per column. +- No two queens share the same diagonal. -##### Approach Using Backtracking +#### Approach Using Backtracking Backtracking is an ideal algorithmic approach for solving the N-Queens problem due to its constraint satisfaction nature. The algorithm incrementally builds the solution and backtracks when a partial solution violates the constraints. @@ -372,14 +374,16 @@ High-Level Steps: 8. When $N$ queens have been successfully placed without conflicts, record the solution. 9. Continue the process to find all possible solutions. -This flow is exactly “choose → test → recurse → undo.” The fun part is that the board doesn’t need to be fully drawn most of the time. You can represent a queen placement compactly (like “row -> column”), and then your safety check becomes pure logic. That’s a nice lesson: backtracking is often more about managing **state** than about fancy data structures. +This flow is exactly “choose → test → recurse → undo.” The fun part is that the board doesn’t need to be fully drawn most of the time. You can represent a queen placement compactly (like “row -> column”), and then your safety check becomes pure logic. That’s a nice lesson: backtracking is often more about managing state than about fancy data structures. -##### Python Implementation +#### Python Implementation Below is a Python implementation of the N-Queens problem using backtracking. ```python def solve_n_queens(N): + if N < 0: + raise ValueError("N must be non-negative") solutions = [] board = [-1] * N # board[row] = column position of queen in that row @@ -420,7 +424,7 @@ for index, sol in enumerate(solutions): print(' '.join(line)) ``` -##### Execution Flow +#### Execution Flow 1. Try placing a queen in columns `0` to `N - 1`. 2. For each valid placement, proceed to row `1`. @@ -428,11 +432,11 @@ for index, sol in enumerate(solutions): 4. If no safe column is found, backtrack to the previous row. 5. When a valid placement is found for all $N$ rows, record the solution. -##### All Solutions for $N = 4$ +#### All Solutions for $N = 4$ There are two distinct solutions for $N = 4$: -**Solution 1:** +Solution 1: ``` Board Representation: [1, 3, 0, 2] @@ -449,7 +453,7 @@ Board Representation: [1, 3, 0, 2] +---+---+---+---+ ``` -**Solution 2:** +Solution 2: ``` Board Representation: [2, 0, 3, 1] @@ -466,7 +470,7 @@ Board Representation: [2, 0, 3, 1] +---+---+---+---+ ``` -##### Output of the Program +#### Output of the Program ``` Number of solutions for N=4: 2 @@ -484,58 +488,58 @@ Q . . . . Q . . ``` -##### Visualization of the Backtracking Tree +#### Visualization of the Backtracking Tree The algorithm explores the solution space as a tree, where each node represents a partial solution (queens placed up to a certain row). The branches represent the possible positions for the next queen. -* **Nodes** represent partial solutions where a certain number of queens have already been placed in specific rows. -* **Branches** correspond to the possible positions for placing the next queen in the following row, exploring each valid option. -* **Leaves** are the complete solutions when all $N$ queens have been successfully placed on the board without conflicts. +- Nodes represent partial solutions where a certain number of queens have already been placed in specific rows. +- Branches correspond to the possible positions for placing the next queen in the following row, exploring each valid option. +- Leaves are either complete solutions or dead ends with no valid next placement. The backtracking occurs when a node has no valid branches (no safe positions in the next row), prompting the algorithm to return to the previous node and try other options. -##### Analysis +#### Analysis -I. The **time complexity** of the N-Queens problem is $O(N!)$ as the algorithm explores permutations of queen placements across rows. +I. The search has factorial growth because valid partial placements use distinct columns. The displayed implementation also tests every column and scans earlier rows in `is_safe`. A conservative bound is $O(N^2 N!)$, plus $O(SN)$ to copy $S$ solutions. Quoting $O(N!)$ alone omits this implementation's candidate-checking cost. Column and diagonal sets can make each conflict test constant time. -II. The **space complexity** is $O(N)$, where: +II. Auxiliary space is $O(N)$, excluding the $O(SN)$ stored output, where: -* The `board` array stores the positions of the $N$ queens. -* The recursion stack can go as deep as $N$ levels during the backtracking process. +- The `board` array stores the positions of the $N$ queens. +- The recursion stack can go as deep as $N$ levels during the backtracking process. This is a good moment to connect the “why should I care?” dot: many real problems look like N-Queens under the hood, scheduling, assignment, routing with constraints, configuration systems, even some parts of compiler design. You’re practicing a general way to search through choices without drowning in them. -##### Applications +#### Applications -* **Constraint satisfaction problems** often use the N-Queens problem as a classic example to study and develop solutions for placing constraints on variable assignments. -* In **algorithm design**, the N-Queens problem helps illustrate the principles of backtracking and recursive problem-solving. -* In **artificial intelligence**, it serves as a foundational example for search algorithms and optimization techniques. +- Constraint satisfaction problems often use the N-Queens problem as a classic example to study and develop solutions for placing constraints on variable assignments. +- In algorithm design, the N-Queens problem helps illustrate the principles of backtracking and recursive problem-solving. +- In artificial intelligence, it serves as a foundational example for search algorithms and optimization techniques. -##### Potential Improvements +#### Potential Improvements -* Implementing more efficient conflict detection methods. -* Using heuristics to choose the order of columns to try first. -* Converting the recursive solution to an iterative one using explicit stacks to handle larger values of $N$ without stack overflow. +- Implementing more efficient conflict detection methods. +- Using heuristics to choose the order of columns to try first. +- Converting the recursive solution to an iterative one using explicit stacks to handle larger values of $N$ without stack overflow. -#### Example: Maze Solver +### Example: Maze Solver Given a maze represented as a 2D grid, find a path from the starting point to the goal using backtracking. The maze consists of open paths and walls, and movement is allowed in four directions: up, down, left, and right (no diagonal moves). The goal is to determine a sequence of moves that leads from the start to the goal without crossing any walls. -Maze solving makes backtracking feel instantly real. You’re making choices (directions), you hit walls or loops (constraints), and you undo moves when you get stuck. Even if you never care about mazes, you *do* care about the pattern: this is the same structure as exploring possible decisions in a game AI, navigating states in a program, or searching combinations until you find one that works. +Maze solving makes backtracking feel instantly real. You’re making choices (directions), you hit walls or loops (constraints), and you undo moves when you get stuck. Even if you never care about mazes, you do care about the pattern: this is the same structure as exploring possible decisions in a game AI, navigating states in a program, or searching combinations until you find one that works. -##### Maze Representation +#### Maze Representation -**Grid Cells:** +Grid Cells: -* `.` (dot) represents an **open path**. -* `#` (hash) represents a **wall** or **obstacle**. +- `.` (dot) represents an open path. +- `#` (hash) represents a wall or obstacle. -**Allowed Moves:** +Allowed Moves: -* Up, down, left, right. -* No diagonal movement. +- Up, down, left, right. +- No diagonal movement. -##### ASCII Representation +#### ASCII Representation Let's visualize the maze using ASCII graphics to better understand the problem. @@ -583,10 +587,12 @@ Objective: Find a sequence of moves from `S` to `G`, navigating only through open paths (`.`) and avoiding walls (`#`). The path should be returned as a list of grid coordinates representing the steps from the start to the goal. -##### Python Implementation +#### Python Implementation ```python def solve_maze(maze, start, goal): + if not maze or not maze[0]: + return None rows, cols = len(maze), len(maze[0]) path = [] @@ -606,6 +612,7 @@ def solve_maze(maze, start, goal): explore(x - 1, y) or explore(x, y + 1) or explore(x, y - 1)): + maze[x][y] = '.' return True path.pop() # Backtrack maze[x][y] = '.' # Unmark visited @@ -638,77 +645,77 @@ else: print("No path found.") ``` -One small but important detail: this code’s `is_valid` only allows stepping onto `'.'`, so `S` and `G` are treated as labels in the diagram, not actual characters in the grid. In practice, you either keep the grid as dots and track `start`/`goal` separately (like this code does), or you expand `is_valid` to allow stepping onto `'G'` as well. The version here is consistent because the grid itself is all `'.'` and `'#'`, and the goal is a coordinate. +One small but important detail: this code’s `is_valid` only allows stepping onto `'.'`, so `S` and `G` are treated as labels in the diagram, not actual characters in the grid. In practice, you either keep the grid as dots and track `start`/`goal` separately (like this code does), or you expand `is_valid` to allow both `'S'` and `'G'` and restore each cell's original character after exploring it. The version here is consistent because the grid itself is all `'.'` and `'#'`, and the goal is a coordinate. -##### Recursive Function `explore(x, y)` +#### Recursive Function `explore(x, y)` -I. **Base Cases:** +I. Base Cases: -* If `(x, y)` is not valid (out of bounds, wall, or visited), return `False`. -* If `(x, y)` equals the goal position, append it to `path` and return `True`. +- If `(x, y)` is not valid (out of bounds, wall, or visited), return `False`. +- If `(x, y)` equals the goal position, append it to `path` and return `True`. -II. **Recursive Exploration:** +II. Recursive Exploration: -* Mark the current cell `(x, y)` as visited by setting `maze[x][y] = 'V'`. -* Append `(x, y)` to the `path`. -* Recursively attempt to explore neighboring cells in the following order: -* Move **Down**: `explore(x + 1, y)` -* Move **Up**: `explore(x - 1, y)` -* Move **Right**: `explore(x, y + 1)` -* Move **Left**: `explore(x, y - 1)` -* If any recursive call returns `True`, propagate the `True` value upwards. +- Mark the current cell `(x, y)` as visited by setting `maze[x][y] = 'V'`. +- Append `(x, y)` to the `path`. +- Recursively attempt to explore neighboring cells in the following order: +- Move Down: `explore(x + 1, y)` +- Move Up: `explore(x - 1, y)` +- Move Right: `explore(x, y + 1)` +- Move Left: `explore(x, y - 1)` +- If a recursive call succeeds, restore the current cell to `'.'` and propagate `True`, retaining the successful path. The input grid is restored on both successful and failed searches. -III. **Backtracking:** +III. Backtracking: -* If none of the neighboring cells lead to a solution, backtrack: -* Remove `(x, y)` from `path` using `path.pop()`. -* Unmark the cell by setting `maze[x][y] = '.'`. -* Return `False` to indicate that this path does not lead to the goal. +- If none of the neighboring cells lead to a solution, backtrack: +- Remove `(x, y)` from `path` using `path.pop()`. +- Unmark the cell by setting `maze[x][y] = '.'`. +- Return `False` to indicate that this path does not lead to the goal. -This is the heart of backtracking in one function: **mark, explore, undo**. The “undo” is what keeps the search honest. Without undoing the visited mark and the path append, you’d either get stuck or accidentally block valid routes later. +This is the heart of backtracking in one function: mark, explore, undo. The “undo” is what keeps the search honest. Here visited marks belong to the current path and are undone so alternative simple paths can reuse cells. For finding any route in a static maze, a permanent visited set is also correct and avoids repeated exploration. Path-dependent puzzles, such as word search, need more careful state tracking. -##### Execution Flow +#### Execution Flow -I. **Start at `(0, 0)`**: +I. Start at `(0, 0)`: -* The algorithm begins at the starting position. -* Marks `(0, 0)` as visited and adds it to the path. +- The algorithm begins at the starting position. +- Marks `(0, 0)` as visited and adds it to the path. -II. **Explore Neighbors**: +II. Explore Neighbors: -* Tries moving **Down** to `(1, 0)`. +- Tries moving Down to `(1, 0)`. -III. **Recursive Exploration**: +III. Recursive Exploration: -* From `(1, 0)`, continues moving **Down** to `(2, 0)`. -* From `(2, 0)`, attempts **Right** to `(2, 1)`. +- From `(1, 0)`, continues moving Down to `(2, 0)`. +- From `(2, 0)`, first explores downward through `(3, 0)` into the dead end along row 4. After undoing those moves, it tries right to `(2, 1)`. -IV. **Dead Ends and Backtracking**: +IV. Dead Ends and Backtracking: -* If a path leads to a wall or visited cell, the algorithm backtracks to the previous cell and tries a different direction. -* This process continues, exploring all possible paths recursively. +- If a path leads to a wall or visited cell, the algorithm backtracks to the previous cell and tries a different direction. +- This process continues, exploring all possible paths recursively. -V. **Reaching the Goal**: +V. Reaching the Goal: -* Eventually, the algorithm reaches the goal `(5, 5)` if a path exists. -* The function returns `True`, and the full path is constructed via the recursive calls. +- Eventually, the algorithm reaches the goal `(5, 5)` if a path exists. +- The function returns `True`, and the full path is constructed via the recursive calls. -##### Output +#### Output -* If a path is found, it prints "Path to goal:" followed by the list of coordinates in the path. -* If no path exists, it prints "No path found." +- If a path is found, it prints "Path to goal:" followed by the list of coordinates in the path. +- If no path exists, it prints "No path found." -##### Final Path Found +#### Final Path Found The path from start to goal: ``` -[(0, 0), (1, 0), (2, 0), (2, 1), (2, 2), (2, 3), - (1, 3), (1, 4), (1, 5), (2, 5), (3, 5), (4, 5), - (5, 5)] +[(0, 0), (1, 0), (2, 0), (2, 1), (2, 2), (1, 2), + (1, 3), (0, 3), (0, 4), (1, 4), (1, 5), (2, 5), + (3, 5), (4, 5), (5, 5)] ``` -##### Visual Representation of the Path +#### Visual Representation of the Path Let's overlay the path onto the maze for better visualization. We'll use `*` to indicate the path. @@ -717,11 +724,11 @@ Maze with Path: 0 1 2 3 4 5 +---+---+---+---+---+---+ -0 | * | * | # | . | . | . | +0 | * | . | # | * | * | . | +---+---+---+---+---+---+ -1 | * | # | . | * | * | * | +1 | * | # | * | * | * | * | +---+---+---+---+---+---+ -2 | * | * | * | * | # | * | +2 | * | * | * | . | # | * | +---+---+---+---+---+---+ 3 | . | # | # | # | . | * | +---+---+---+---+---+---+ @@ -736,88 +743,90 @@ Legend: . - Open path ``` -##### Advantages of Using Backtracking for Maze Solving +For a rectangular grid with $R$ rows and $C$ columns, this path-based search uses $O(RC)$ auxiliary space but can explore exponentially many simple paths. A DFS with permanent visited marks finds any reachable goal in $O(RC)$ time; BFS has the same time bound and finds a shortest path when every move has equal cost. + +#### Advantages of Using Backtracking for Maze Solving -* Ensures that all possible paths are explored until the goal is found. -* Only the current path and visited cells are stored, reducing memory usage compared to storing all possible paths. -* Recursive implementation leads to clean and understandable code. +- Ensures that all possible paths are explored until the goal is found. +- Only the current path and visited cells are stored, reducing memory usage compared to storing all possible paths. +- Recursive implementation leads to clean and understandable code. -##### Potential Improvements +#### Potential Improvements -* This algorithm finds a path but not necessarily the shortest path. -* To find the shortest path, algorithms like Breadth-First Search (BFS) are more suitable. -* Modify the code to collect all possible paths by removing early returns when the goal is found. -* Allowing diagonal movements would require adjusting the `explore` function to include additional directions. +- This algorithm finds a path but not necessarily the shortest path. +- To find the shortest path, algorithms like Breadth-First Search (BFS) are more suitable. +- Modify the code to collect all possible paths by removing early returns when the goal is found. +- Allowing diagonal movements would require adjusting the `explore` function to include additional directions. -This is where you can choose your “vibe” depending on the problem: +Choose the search method according to the required result: -* If you just need **any** solution fast and the maze is small, backtracking is perfect. -* If you need the **best** solution (shortest path), you switch strategies (like BFS). -* If the search space is massive, you start adding smarter pruning, heuristics, or even entirely different approaches. +- If you just need any solution fast and the maze is small, backtracking is perfect. +- If you need the best solution (shortest path), you switch strategies (like BFS). +- If the search space is massive, you start adding smarter pruning, heuristics, or even entirely different approaches. -Backtracking isn’t the final boss :) it’s just a foundation. Once you understand it, you start recognizing when a problem is a “choices + constraints” machine, and you’ll know exactly how to build a solver that’s not only correct, but actually fun to reason about. +The reusable pattern is to represent partial choices, check constraints, explore a candidate, and restore state. Understanding these responsibilities makes it easier to adapt the method to other constraint problems. -### List of Problems +## List of Problems -#### Permutations +### Permutations Develop an algorithm to generate all possible permutations of a given list of elements. This problem requires creating different arrangements of the elements where the order matters. -* [C++ Solution](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/backtracking/cpp/all_permutations) -* [Python Solution](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/backtracking/python/all_permutations) +- [C++ Solution](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/backtracking/cpp/all_permutations) +- [Python Solution](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/backtracking/python/all_permutations) -#### Combinations +### Combinations Design an algorithm to generate all possible combinations of 'k' elements selected from a given list of elements. This involves selecting elements where the order does not matter, but the selection size does. -* [C++ Solution](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/backtracking/cpp/all_combinations) -* [Python Solution](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/backtracking/python/all_combinations) +- [C++ Solution](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/backtracking/cpp/all_combinations) +- [Python Solution](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/backtracking/python/all_combinations) -#### String Pattern +### String Pattern -Create a solution to determine whether a given string adheres to a specified pattern, where the pattern may include letters and wildcard characters that represent any character. This problem often involves checking for matches and handling special pattern symbols. +Determine whether pattern symbols can be mapped to nonempty substrings so that concatenating their mappings reproduces the entire input string. Repeated occurrences of a symbol must use the same substring. For example, `abab` can match `redblueredblue` with `a → red` and `b → blue`. State separately whether distinct symbols must have distinct mappings; that is an additional constraint, not ordinary wildcard matching. -* [C++ Solution](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/backtracking/cpp/string_pattern) -* [Python Solution](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/backtracking/python/string_pattern) +- [C++ Solution](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/backtracking/cpp/string_pattern) +- [Python Solution](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/backtracking/python/string_pattern) -#### Generating Words +### Generating Words -Generate all possible words that can be formed from a given list of characters and match a specified pattern. The pattern can contain letters and wildcard characters, requiring the algorithm to account for flexible matching. +Given a board of letters and a dictionary, find dictionary words formed by paths through neighboring cells. This exercise allows horizontal, vertical, and diagonal moves. Track cells used in the current path, define whether reuse is allowed, and reject prefixes that cannot lead to dictionary words. -* [C++ Solution](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/backtracking/cpp/generating_words) -* [Python Solution](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/backtracking/python/generating_words) +- [C++ Solution](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/backtracking/cpp/generating_words) +- [Python Solution](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/backtracking/python/generating_words) -#### Hamiltonian Path +### Hamiltonian Path -Create an algorithm that identifies whether a simple path exists within a provided undirected or directed graph. This path should visit every vertex exactly once. Known as the "traveling salesman problem," it can be addressed using depth-first search to explore possible paths. +Create an algorithm that identifies whether a simple path exists within a provided undirected or directed graph. This path should visit every vertex exactly once. A Hamiltonian path visits every vertex once; a Hamiltonian cycle also returns to its start. The traveling salesperson problem is a related optimization problem that asks for a minimum-weight tour. Backtracking can enumerate candidate paths and reject repeated vertices. -* [C++ Solution](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/backtracking/cpp/hamiltonian_paths) -* [Python Solution](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/backtracking/python/hamiltonian_paths) +- [C++ Solution](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/backtracking/cpp/hamiltonian_paths) +- [Python Solution](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/backtracking/python/hamiltonian_paths) -#### K-Colorable Configurations +### K-Colorable Configurations Develop an algorithm to find all possible ways to color a given graph with 'k' colors such that no two adjacent vertices share the same color. This graph coloring problem requires ensuring valid color assignments for all vertices. -* [C++ Solution](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/backtracking/cpp/k_colorable_configurations) -* [Python Solution](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/backtracking/python/k_colorable_configurations) +- [C++ Solution](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/backtracking/cpp/k_colorable_configurations) +- [Python Solution](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/backtracking/python/k_colorable_configurations) -#### Knight Tour +### Knight Tour Create an algorithm to find all potential paths a knight can take on an 'n' x 'n' chessboard to visit every square exactly once. This classic chess problem involves ensuring the knight moves in an L-shape and covers all board positions. -* [C++ Solution](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/backtracking/cpp/knight_tour) -* [Python Solution](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/backtracking/python/knight_tour) +- [C++ Solution](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/backtracking/cpp/knight_tour) +- [Python Solution](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/backtracking/python/knight_tour) -#### Topological Orderings +### Topological Orderings -Determine a topological ordering of the vertices in a given directed graph, if one exists. This involves sorting the vertices such that for every directed edge UV from vertex U to vertex V, U comes before V in the ordering. +Enumerate topological orderings of a directed acyclic graph by repeatedly choosing an available zero-indegree vertex, updating indegrees, and undoing the choice. If only one ordering is needed, a standard topological sort is more efficient. This involves sorting the vertices such that for every directed edge UV from vertex U to vertex V, U comes before V in the ordering. -* [C++ Solution](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/backtracking/cpp/topological_sort) -* [Python Solution](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/backtracking/python/topological_sort) +- [C++ Solution](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/backtracking/cpp/topological_sort) +- [Python Solution](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/backtracking/python/topological_sort) -#### Tic-Tac-Toe (Minimax) +### Tic-Tac-Toe (Minimax) Develop an algorithm to determine the optimal move for a player in a game of tic-tac-toe using the minimax algorithm. This requires evaluating possible moves to find the best strategy for winning or drawing the game. -* [C++ Solution](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/backtracking/cpp/minimax) -* [Python Solution](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/backtracking/python/minimax) +- [C++ Solution](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/backtracking/cpp/minimax) +- [Python Solution](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/backtracking/python/minimax) diff --git a/notes/basic_concepts.md b/notes/basic_concepts.md index 769cee5..0451620 100644 --- a/notes/basic_concepts.md +++ b/notes/basic_concepts.md @@ -1,22 +1,22 @@ -## Introduction to Data Structures & Algorithms +# Introduction to Data Structures & Algorithms -Data structures and algorithms are fundamental concepts in computer science and they are the only way to write efficient software. +Data structures and algorithms are fundamental tools for writing efficient software. They determine how a program organizes information and how much work it performs; implementation choices, I/O, and hardware also affect performance. -Most “slow code” isn’t slow because the programmer typed badly, it’s slow because the program is *doing the wrong kind of work* for the size of the input. Data structures and algorithms are the two knobs you can turn to fix that. Once you see them as practical tools (not academic trivia), the whole topic becomes less intimidating and a lot more useful. +Software often becomes slow because its approach performs too much work as the input grows. Choosing a suitable data structure and algorithm can reduce that work and make the solution easier to reason about. -* A **data structure** specifies how data is stored and organized in memory. Examples include arrays, linked lists, stacks, queues, trees, and graphs. Choosing the right data structure can simplify solving specific problems. -* An **algorithm** is a step-by-step method for solving a problem or performing a task. Algorithms can range from simple operations like searching or sorting to complex computations in artificial intelligence or optimization. -* The combination of efficient data structures and algorithms enables developers to optimize software performance, both in terms of speed and memory usage. +- A data structure specifies how data is stored and organized in memory. Examples include arrays, linked lists, stacks, queues, trees, and graphs. Choosing the right data structure can simplify solving specific problems. +- An algorithm is a step-by-step method for solving a problem or performing a task. Algorithms can range from simple operations like searching or sorting to complex computations in artificial intelligence or optimization. +- The combination of efficient data structures and algorithms enables developers to optimize software performance, both in terms of speed and memory usage. -### Data Structures +## Data Structures -A **data structure** organizes and stores data in a way that allows efficient access, modification, and processing. The choice of the appropriate data structure depends on the specific use case and can significantly impact the performance of an application. Here are some common data structures: +A data structure organizes and stores data in a way that allows efficient access, modification, and processing. The choice of the appropriate data structure depends on the specific use case and can significantly impact the performance of an application. Here are some common data structures: -A useful way to think about data structures is: they’re not just “ways to store data,” they’re **promises**. An array promises fast indexing. A stack promises “last thing in is the first thing out.” A tree promises hierarchy. When you pick the structure whose promise matches your problem, the solution often becomes simpler and faster at the same time. +Each data structure supports a particular set of operations and ordering rules. Arrays offer constant-time indexing, stacks enforce last-in-first-out access, and trees represent hierarchy. Match those operations to the problem before choosing an implementation. -**I. Array** +### Array -Imagine an **array** as a row of lockers, each labeled with a number and capable of holding one item of the same type. Technically, arrays are blocks of memory storing elements sequentially, allowing quick access using an index. However, arrays have a fixed size, which limits their flexibility when you need to add or remove items. +Imagine an array as a row of lockers, each labeled with a number and capable of holding one item of the same type. Technically, arrays are blocks of memory storing elements sequentially, allowing quick access using an index. A fixed-size array cannot grow after allocation. Dynamic arrays, such as Python lists and C++ vectors, grow by allocating a larger buffer and moving or copying elements when their capacity is exhausted. A “do” with arrays: use them when you need quick random access by index, or when you’re scanning data in order. A “don’t”: assume insertions and deletions in the middle are cheap, shifting elements can turn a clean-looking solution into a slow one when the array is large. @@ -25,11 +25,11 @@ Indices: 0 1 2 3 Array: [A] [B] [C] [D] ``` -**II. Stack** +### Stack -Think of a **stack** like stacking plates: you always add new plates on top (push), and remove them from the top as well (pop). This structure follows the Last-In, First-Out (LIFO) approach, meaning the most recently added item is removed first. Stacks are particularly helpful in managing function calls (like in the call stack of a program) or enabling "undo" operations in applications. +Think of a stack like stacking plates: you always add new plates on top (push), and remove them from the top as well (pop). This structure follows the Last-In, First-Out (LIFO) approach, meaning the most recently added item is removed first. Stacks are particularly helpful in managing function calls (like in the call stack of a program) or enabling "undo" operations in applications. -Stacks shine when your problem has a “most recent context” feel: undo/redo, parsing, backtracking, evaluating expressions, matching parentheses. The “do” is to lean on the LIFO rule to avoid complicated bookkeeping. The “don’t” is to use a stack when you actually need “oldest first”, that’s a queue. +Stacks shine when your problem has a “most recent context” feel: undo/redo, parsing, backtracking, evaluating expressions, matching parentheses. Lean on the LIFO rule to avoid complicated bookkeeping. Do not use a stack when you actually need “oldest first”, that’s a queue. ``` Top @@ -43,9 +43,9 @@ Top Bottom ``` -**III. Queue** +### Queue -A **queue** is similar to a line at the grocery store checkout. People join at the end (enqueue) and leave from the front (dequeue), adhering to the First-In, First-Out (FIFO) principle. This ensures the first person (or item) that arrives is also the first to leave. Queues work great for handling tasks or events in the exact order they occur, like scheduling print jobs or processing messages. +A queue is similar to a line at the grocery store checkout. People join at the end (enqueue) and leave from the front (dequeue), adhering to the First-In, First-Out (FIFO) principle. This ensures the first person (or item) that arrives is also the first to leave. Queues work great for handling tasks or events in the exact order they occur, like scheduling print jobs or processing messages. Queues are your go-to when fairness and order matter: job scheduling, buffering, breadth-first search, producer/consumer pipelines. A solid “do” is to use a queue when the problem sounds like “process things in the order they arrived.” A common “don’t” is implementing a queue with an array in a way that forces shifting on every dequeue, use a deque or circular buffer instead. @@ -54,21 +54,21 @@ Front → [A] → [B] → [C] → [D] ← Rear (dequeue) (enqueue) ``` -**IV. Linked List** +### Linked List -You can picture a **linked list** as a treasure hunt, where each clue leads you to the next one. Each clue, or node, holds data and a pointer directing you to the next node. Because nodes can be added or removed without shifting other elements around, linked lists offer dynamic and flexible management of data at any position. +You can picture a linked list as a treasure hunt, where each clue leads you to the next one. Each clue, or node, holds data and a pointer directing you to the next node. Because nodes can be added or removed without shifting other elements around, linked lists offer dynamic and flexible management of data at any position. -Linked lists are great when you need frequent insertions or deletions *once you already have the node*, but they’re not great at random access. The “do” is to use them when you’re walking forward and rewiring pointers. The “don’t” is to treat them like arrays, finding “the 10,000th element” is slow because you have to traverse there step by step. +Linked lists are great when you need frequent insertions or deletions once you already have the node, but they’re not great at random access. Use them when you’re walking forward and rewiring pointers. Do not treat them like arrays, finding “the 10,000th element” is slow because you have to traverse there step by step. ``` Head -> [A] -> [B] -> [C] -> NULL ``` -**V. Tree** +### Tree -A **tree** resembles a family tree, starting from one ancestor (the root) and branching out into multiple descendants (nodes), each of which can have their own children. Formally, trees are hierarchical structures organized across various levels. They’re excellent for showing hierarchical relationships, such as organizing files on your computer or visualizing company structures. +A tree resembles a family tree, starting from one ancestor (the root) and branching out into multiple descendants (nodes), each of which can have their own children. Formally, trees are hierarchical structures organized across various levels. They’re excellent for showing hierarchical relationships, such as organizing files on your computer or visualizing company structures. -Trees matter because many real problems are hierarchical even when they don’t look like it at first: file systems, DOM structures, organization charts, decision processes, indexes. The “do” is to use tree traversals (pre/in/post/level-order) so your logic stays systematic. The “don’t” is to forget that tree shape affects performance, an unbalanced tree can quietly turn fast operations into slow ones. +Trees matter because many real problems are hierarchical even when they don’t look like it at first: file systems, DOM structures, organization charts, decision processes, indexes. Use tree traversals (pre/in/post/level-order) so your logic stays systematic. Do not forget that tree shape affects performance, an unbalanced tree can quietly turn fast operations into slow ones. ``` # Tree @@ -79,11 +79,11 @@ Trees matter because many real problems are hierarchical even when they don’t (LL) (LR) (RR) ``` -**VI. Graph** +### Graph -Consider a **graph** like a network of cities connected by roads. Each city represents a node, and the roads connecting them are edges, which can either be one-way (directed) or two-way (undirected). Graphs effectively illustrate complex relationships and networks, such as social media connections, website link structures, or even mapping transportation routes. +Consider a graph like a network of cities connected by roads. Each city represents a node, and the roads connecting them are edges, which can either be one-way (directed) or two-way (undirected). Graphs effectively illustrate complex relationships and networks, such as social media connections, website link structures, or even mapping transportation routes. -Graphs are the “everything is connected” structure. The moment you see relationships that aren’t strictly hierarchical, friends, links, routes, dependencies, you’re in graph territory. The “do” is to ask early: is this graph directed or undirected, weighted or unweighted? That choice decides whether BFS, DFS, Dijkstra, or something else is the right move. The “don’t” is to ignore direction or weights and then wonder why the result feels wrong. +Graphs are the “everything is connected” structure. The moment you see relationships that aren’t strictly hierarchical, friends, links, routes, dependencies, you’re in graph territory. Ask early: is this graph directed or undirected, weighted or unweighted? That choice decides whether BFS, DFS, Dijkstra, or something else is the right move. Do not ignore direction or weights and then wonder why the result feels wrong. ``` (A) ↔ (B) @@ -95,29 +95,29 @@ Graphs are the “everything is connected” structure. The moment you see relat ![image](https://github.com/user-attachments/assets/f1962de7-aa28-4348-9933-07e49c737cd9) -### Algorithms +## Algorithms -An **algorithm** is like a clear and detailed set of instructions or steps for solving a specific problem or performing a particular task. Think of it like following a precise recipe when cooking: +An algorithm is like a clear and detailed set of instructions or steps for solving a specific problem or performing a particular task. Think of it like following a precise recipe when cooking: -The reason this “recipe” framing sticks is that good algorithms are repeatable and dependable: the same inputs give the same output, and the steps don’t depend on magic. When you’re debugging or optimizing, that predictability is everything, it’s what lets you reason about behavior instead of guessing. +An algorithm defines its steps precisely enough to reason about correctness. A deterministic algorithm produces the same output for the same input. A randomized algorithm also uses random choices, so its execution or output can vary; its correctness and performance guarantees must account for that randomness. -* The ingredients needed represent the **input**, which are the data or information the algorithm uses to begin its work. -* The finished dish is the **output**, or the final result the algorithm provides after processing the input. -* Each instruction in an algorithm must be clear and precise. This clarity is known as **definiteness**, ensuring anyone following the steps reaches the same result without confusion. -* Algorithms must have a definite end-point, known as **termination**, meaning they can’t run indefinitely and must eventually finish with a result. -* Every step in an algorithm must be practical and achievable, known as **effectiveness**, ensuring the instructions can realistically be carried out to achieve the desired outcome. +- The ingredients needed represent the input, which are the data or information the algorithm uses to begin its work. +- The finished dish is the output, or the final result the algorithm provides after processing the input. +- Each instruction in an algorithm must be clear and precise. This clarity is known as definiteness, ensuring anyone following the steps reaches the same result without confusion. +- Algorithms must have a definite end-point, known as termination, meaning they can’t run indefinitely and must eventually finish with a result. +- Every step in an algorithm must be practical and achievable, known as effectiveness, ensuring the instructions can realistically be carried out to achieve the desired outcome. To evaluate how good an algorithm is, we often look at its efficiency in terms of time complexity (how long it takes to run) and space complexity (how much memory it uses). We will discuss it in greater detail in later sections. A quick “do and don’t” here: do focus on clarity first, an algorithm you can’t explain is hard to trust. Don’t confuse a clever trick with an algorithmic improvement; sometimes the biggest win is choosing a better approach, not writing tighter code. -#### Algorithms vs. Programs +### Algorithms vs. Programs -An **algorithm** is a high-level blueprint for solving a specific problem. It is abstract, language-independent, and specifies a clear sequence of steps without relying on any particular programming syntax. An algorithm can be thought of as a recipe or method for solving a problem and can be represented in multiple forms, such as plain text or a flowchart. +An algorithm is a high-level blueprint for solving a specific problem. It is abstract, language-independent, and specifies a clear sequence of steps without relying on any particular programming syntax. An algorithm can be thought of as a recipe or method for solving a problem and can be represented in multiple forms, such as plain text or a flowchart. -The distinction matters because you can evaluate an algorithm *before* you write code: does it terminate, does it cover edge cases, what’s its complexity, how does it scale? That’s a superpower in interviews and in real projects, design first, implement second. +Before implementing an algorithm, check whether it terminates, covers edge cases, and has acceptable time and space costs. This separates questions about the method from mistakes in its implementation. -**Example:** Algorithm for adding two numbers: +Example: Algorithm for adding two numbers: ``` Step 1: Start @@ -161,7 +161,7 @@ This algorithm can also be visualized using a flowchart: -------------------------- ``` -In contrast, a **program** is a concrete implementation of an algorithm. It is language-dependent and adheres to the specific syntax and rules of a programming language. For example, the above algorithm can be implemented in Python as: +In contrast, a program is a concrete implementation of an algorithm. It is language-dependent and adheres to the specific syntax and rules of a programming language. For example, the above algorithm can be implemented in Python as: ```python num1 = int(input("Enter first number: ")) @@ -172,17 +172,17 @@ print("The sum is", sum) Programs may sometimes run indefinitely or until an external action stops them. For instance, an operating system is a program designed to run continuously until explicitly terminated. -A good practical habit: when you write a program, keep the algorithm “visible” in the structure of the code, clear function names, logical steps, clean invariants. That makes debugging and optimization feel like adjusting a plan, not untangling a knot. +Keep the algorithm recognizable in the program through clear names, explicit steps, and invariants. This makes it easier to locate mistakes and measure the effect of an optimization. -#### Types of Algorithms +### Types of Algorithms Algorithms can be classified into various types based on the problems they solve and the strategies they use. Here are some common categories with consistent explanations and examples: -This classification helps because it turns “a million problems” into “a few families.” When you recognize the family, you can reuse known techniques instead of reinventing solutions. The “do” is to build pattern recognition. The “don’t” is to memorize implementations without understanding what problem shape they fit. +This classification helps because it turns “a million problems” into “a few families.” When you recognize the family, you can reuse known techniques instead of reinventing solutions. Build pattern recognition. Do not memorize implementations without understanding what problem shape they fit. -I. **Sorting Algorithms** arrange data in a specific order, such as ascending or descending. Examples include bubble sort, insertion sort, selection sort, and merge sort. +I. Sorting Algorithms arrange data in a specific order, such as ascending or descending. Examples include bubble sort, insertion sort, selection sort, and merge sort. -Sorting is often the sneaky first step that makes everything else easier: once data is ordered, you can use binary search, two pointers, sweeping scans, and duplicate skipping. The “do” is to ask: am I allowed to reorder the data? The “don’t” is to sort blindly when order matters or when a hash-based approach would be cheaper. +Sorting is often the sneaky first step that makes everything else easier: once data is ordered, you can use binary search, two pointers, sweeping scans, and duplicate skipping. Ask: am I allowed to reorder the data? Do not sort blindly when order matters or when a hash-based approach would be cheaper. Example: Bubble Sort @@ -200,9 +200,9 @@ After 3rd Pass: [3, 2, 4, 5, 8] After 4th Pass: [2, 3, 4, 5, 8] (Sorted) ``` -II. **Search Algorithms** are designed to find a specific item or value within a collection of data. Examples include linear search, binary search, and depth-first search. +II. Search Algorithms are designed to find a specific item or value within a collection of data. Examples include linear search, binary search, and depth-first search. -Searching is about trade-offs: linear search is simple and often fine for small inputs; binary search is fast but needs sorted data; DFS/BFS are “search” across relationships rather than lists. The “do” is to match the search method to the structure. The “don’t” is to use binary search without guaranteeing sortedness. +Searching is about trade-offs: linear search is simple and often fine for small inputs; binary search is fast but needs sorted data; DFS/BFS are “search” across relationships rather than lists. Match the search method to the structure. Do not use binary search without guaranteeing sortedness. Example: Binary Search @@ -226,20 +226,20 @@ New mid element: 11 The remaining element is 33, which is the target. ``` -III. **Graph Algorithms** address problems related to graphs, such as finding the shortest path between nodes or determining if a graph is connected. Examples include Dijkstra's algorithm and the Floyd-Warshall algorithm. +III. Graph Algorithms address problems related to graphs, such as finding the shortest path between nodes or determining if a graph is connected. Examples include Dijkstra's algorithm and the Floyd-Warshall algorithm. -Graph algorithms matter because graphs model real systems: routes, networks, dependencies, recommendations. The “do” is to pin down the graph type first (directed? weighted? cyclic?). The “don’t” is to treat all shortest paths the same, unweighted shortest path is BFS, but weighted shortest path often needs Dijkstra (or Bellman-Ford if negative edges exist). +Graph algorithms matter because graphs model real systems: routes, networks, dependencies, recommendations. Pin down the graph type first (directed? Weighted? Cyclic?). Do not treat all shortest paths the same, unweighted shortest path is BFS, but weighted shortest path often needs Dijkstra (or Bellman-Ford if negative edges exist). Example: Dijkstra's Algorithm -Given a graph with weighted edges, find the shortest path from a starting node to all other nodes. +Given a graph with non-negative edge weights, find the shortest distances from a starting node to all reachable nodes. Dijkstra’s algorithm relies on non-negative weights to finalize distances safely. Steps: 1. Initialize the starting node with a distance of 0 and all other nodes with infinity. 2. Visit the unvisited node with the smallest known distance. 3. Update the distances of its neighboring nodes. -4. Repeat until all nodes have been visited. +4. Repeat until no reachable unvisited node remains. With a priority queue, skip stale entries whose distance no longer matches the best known distance. Example Graph: @@ -251,11 +251,11 @@ B -> D (5) C -> D (1) ``` -Trace Table +Trace Table: | Iter | Extracted Node (u) | PQ before extraction | dist[A,B,C,D] | prev[A,B,C,D] | Visited | Comments / Updates | | ---- | ------------------ | ---------------------------------- | ------------- | ------------- | --------- | -------------------------------------------------------------------------------------- | -| 0 | , (initial) | (0, A) | [0, ∞, ∞, ∞] | [-, -, -, -] | {} | Initialization: A=0, others ∞ | +| 0 |: (initial) | (0, A) | [0, ∞, ∞, ∞] | [-, -, -, -] | {} | Initialization: A=0, others ∞ | | 1 | A (0) | (0, A) | [0, 1, 4, ∞] | [-, A, A, -] | {A} | Relax A→B (1), A→C (4); push (1,B), (4,C) | | 2 | B (1) | (1, B), (4, C) | [0, 1, 3, 6] | [-, A, B, B] | {A, B} | Relax B→C: alt=3 <4 ⇒ update C; B→D: dist[D]=6; push (3,C), (6,D). (4,C) becomes stale | | 3 | C (3) | (3, C), (4, C) stale, (6, D) | [0, 1, 3, 4] | [-, A, B, C] | {A, B, C} | Relax C→D: alt=4 <6 ⇒ update D; push (4,D). (6,D) becomes stale | @@ -263,20 +263,20 @@ Trace Table Legend: -* `dist[X]`: current best known distance from A to X -* `prev[X]`: predecessor of X on that best path -* PQ: min-heap of (tentative distance, node); stale entries (superseded by better distance) are shown in parentheses -* Visited: nodes whose shortest distance is finalized +- `dist[X]`: current best known distance from A to X +- `prev[X]`: predecessor of X on that best path +- PQ: min-heap of (tentative distance, node); stale entries (superseded by better distance) are shown in parentheses +- Visited: nodes whose shortest distance is finalized Starting from A: -* Shortest path to B: A -> B (1) -* Shortest path to C: A -> B -> C (3) -* Shortest path to D: A -> B -> C -> D (4) +- Shortest path to B: A -> B (1) +- Shortest path to C: A -> B -> C (3) +- Shortest path to D: A -> B -> C -> D (4) -IV. **String Algorithms** deal with problems related to strings, such as finding patterns or matching sequences. Examples include the Knuth-Morris-Pratt (KMP) algorithm and the Boyer-Moore algorithm. +IV. String Algorithms deal with problems related to strings, such as finding patterns or matching sequences. Examples include the Knuth-Morris-Pratt (KMP) algorithm and the Boyer-Moore algorithm. -String algorithms are common because text is everywhere: search bars, logs, DNA sequences, code, messages. The “do” is to watch for repeated work, naive substring checks can be painfully slow. The “don’t” is to reach for regex as a hammer for every nail; it’s powerful, but certain patterns can be unexpectedly expensive. +String algorithms are common because text is everywhere: search bars, logs, DNA sequences, code, messages. Watch for repeated work, naive substring checks can be painfully slow. Do not reach for regex as a hammer for every nail; it’s powerful, but certain patterns can be unexpectedly expensive. Example: Boyer-Moore Algorithm @@ -288,94 +288,96 @@ Pattern: "ABABCABAB" Steps: 1. Compare the pattern from right to left. -2. If a mismatch occurs, use the bad character and good suffix heuristics to skip alignments. +2. If a mismatch occurs, shift the pattern using the bad-character rule. Full Boyer–Moore also uses a good-suffix rule; the trace below shows only the bad-character variant. 3. Repeat until the pattern is found or the text is exhausted. | Iter | Start | Text window | Mismatch (pattern vs text) | Shift applied | Next Start | Result | | ---- | ----- | ----------- | ----------------------------------------- | -------------------------------------------------- | ---------- | --------------- | -| 1 | 0 | `ABABDABAC` | pattern[8]=B vs text[8]=C | bad char C → last in pattern at idx4 ⇒ 8−4 = **4** | 4 | no match | -| 2 | 4 | `DABACDABA` | pattern[8]=B vs text[12]=A | bad char A → last at idx7 ⇒ 8−7 = **1** | 5 | no match | -| 3 | 5 | `ABACDABAB` | pattern[4]=C vs text[9]=D | D not in pattern ⇒ 4−(−1)= **5** | 10 | no match | -| 4 | 10 | `ABABCABAB` | full right-to-left comparison → **match** | , | , | **found** at 10 | +| 1 | 0 | `ABABDABAC` | pattern[8]=B vs text[8]=C | bad char C → last in pattern at idx4 ⇒ 8−4 = 4 | 4 | no match | +| 2 | 4 | `DABACDABA` | pattern[8]=B vs text[12]=A | bad char A → last at idx7 ⇒ 8−7 = 1 | 5 | no match | +| 3 | 5 | `ABACDABAB` | pattern[4]=C vs text[9]=D | D not in pattern ⇒ 4−(−1)= 5 | 10 | no match | +| 4 | 10 | `ABABCABAB` | full right-to-left comparison → match |: |: | found at 10 | Pattern matched starting at index 10 in the text. -#### Important Algorithms for Software Engineers +### Important Algorithms for Software Engineers -* As a software engineer, it is not necessary to **master every algorithm**. Instead, knowing how to use libraries and packages that implement widely-used algorithms is more practical. -* The important skill is the ability to **select the right algorithm** for a task by considering factors such as its efficiency, the problem’s requirements, and any specific constraints. -* Learning **algorithms** during the early stages of programming enhances problem-solving skills. It builds a solid foundation in logical thinking, introduces various problem-solving strategies, and helps in understanding how to approach complex issues. -* Once the **fundamentals of algorithms** are understood, the focus often shifts to utilizing pre-built libraries and tools for solving real-world problems, as writing algorithms from scratch is rarely needed in practice. +- As a software engineer, it is not necessary to master every algorithm. Instead, knowing how to use libraries and packages that implement widely-used algorithms is more practical. +- The important skill is the ability to select the right algorithm for a task by considering factors such as its efficiency, the problem’s requirements, and any specific constraints. +- Learning algorithms during the early stages of programming enhances problem-solving skills. It builds a solid foundation in logical thinking, introduces various problem-solving strategies, and helps in understanding how to approach complex issues. +- Once the fundamentals of algorithms are understood, the focus often shifts to utilizing pre-built libraries and tools for solving real-world problems, as writing algorithms from scratch is rarely needed in practice. -This is the part many people miss: in real work, you’re rarely rewarded for reimplementing Dijkstra from memory. You’re rewarded for knowing *that shortest paths is the right framing*, picking the right variant, estimating cost, and using a reliable implementation. The “do” is to become fluent in selection and trade-offs. The “don’t” is to confuse “I can code it from scratch” with “I can solve the real problem.” +This is the part many people miss: in real work, you’re rarely rewarded for reimplementing Dijkstra from memory. You’re rewarded for knowing that shortest paths is the right framing, picking the right variant, estimating cost, and using a reliable implementation. Become fluent in selection and trade-offs. Do not confuse “I can code it from scratch” with “I can solve the real problem.” -Real Life Story: +Illustrative scenario: -``` When Zara landed her first job at a logistics-tech startup, her assignment was to route delivery vans through a sprawling city in under a second, something she’d never tackled before. She remembered the semester she’d wrestled with graph theory and Dijkstra’s algorithm purely for practice, so instead of hand-coding the logic she opened the company’s Python stack and pulled in NetworkX, benchmarking its built-in shortest-path routines against the map’s size and the firm’s latency budget. The initial results were sluggish, so she compared A* with Dijkstra, toggling heuristics until the run time dipped below 500 ms, well under the one-second target. Her teammates were impressed not because she reinvented an algorithm, but because she knew which one to choose, how to reason about its complexity, and where to find a rock-solid library implementation. Later, in a sprint retrospective, Zara admitted that mastering algorithms in college hadn’t been about memorizing code, it had trained her to dissect problems, weigh trade-offs, and plug in the right tool when every millisecond and memory block counted. -``` -### Understanding Algorithmic Complexity +This example concerns shortest routes between individual locations. Choosing the order of many deliveries is a separate routing problem. Any A* heuristic used to preserve shortest-path guarantees must satisfy the conditions explained in the graph notes. + +## Understanding Algorithmic Complexity Algorithmic complexity helps us understand the computational resources (time or space) an algorithm needs as the input size increases. Here’s a breakdown of different types of complexity: -Complexity is the “budgeting” system for software. You don’t need exact microseconds to make good decisions, you need to know how costs grow when data grows. The “do” is to think: *if input doubles, what happens?* The “don’t” is to be fooled by small tests; many algorithms look fine on tiny inputs and collapse at scale. +Complexity is the “budgeting” system for software. You don’t need exact microseconds to make good decisions, you need to know how costs grow when data grows. Think: if input doubles, what happens? Do not be fooled by small tests; many algorithms look fine on tiny inputs and collapse at scale. -* In an ideal input scenario, *best-case complexity* shows the minimum work an algorithm will do; include it to set expectations for quick interactions, omit it and you may overlook fast paths that are useful for user experience, as when insertion sort finishes almost immediately on a nearly sorted list. -* When you ask what to expect most of the time, *average-case complexity* estimates typical running time; include it to make useful forecasts under normal workloads, omit it and designs can seem fine in tests but lag on common inputs, as with randomly ordered customer IDs that need $O(n log n)$ sorting. -* By establishing an upper bound, *worst-case complexity* tells you the maximum time or space an algorithm might need; include it to ensure predictable behavior, omit it and peak loads can surprise you, as when quicksort degrades to $O(n^2)$ on already sorted input without careful pivot selection. -* On memory-limited devices, *space complexity* measures how much extra storage an algorithm requires; include it to fit within available RAM, omit it and an otherwise fast solution may crash or swap, as when merge sort’s $O(n)$ auxiliary array overwhelms a phone with little free memory. -* As your dataset scales, *time complexity* describes how running time expands with input size; include it to choose faster approaches, omit it and performance can degrade sharply, as when an $O(n^2)$ deduplication routine turns a minute-long job into hours after a customer list doubles. +- In an ideal input scenario, best-case complexity shows the minimum work an algorithm will do; include it to set expectations for quick interactions, omit it and you may overlook fast paths that are useful for user experience, as when insertion sort finishes almost immediately on a nearly sorted list. +- When you ask what to expect most of the time, average-case complexity is the expected running time under a stated distribution of inputs; include it to make useful forecasts under normal workloads, omit it and designs can seem fine in tests but lag on common inputs, as with randomly ordered customer IDs that need $O(n log n)$ sorting. +- By establishing an upper bound, worst-case complexity tells you the maximum time or space an algorithm might need; include it to ensure predictable behavior, omit it and peak loads can surprise you, as when quicksort degrades to $O(n^2)$ on already sorted input without careful pivot selection. +- On memory-limited devices, space complexity measures storage use. Total space includes the input; auxiliary space counts only additional storage. State which convention is used; include it to fit within available RAM, omit it and an otherwise fast solution may crash or swap, as when merge sort’s $O(n)$ auxiliary array overwhelms a phone with little free memory. +- As your dataset scales, time complexity describes how running time expands with input size; include it to choose faster approaches, omit it and performance can degrade sharply, as when doubling the input to a quadratic deduplication routine produces approximately four times the work. -#### Analyzing Algorithm Growth Rates +### Analyzing Algorithm Growth Rates Understanding how the running time or space complexity of an algorithm scales with increasing input size is pivotal in algorithm analysis. To describe this rate of growth, we employ several mathematical notations that offer insights into the algorithm's efficiency under different conditions. These notations are less about fancy math and more about communicating clearly. When someone says “this is $O(n \log n)$,” they’re telling you how it behaves as you scale, and whether it’s likely to stay usable when today’s dataset becomes tomorrow’s dataset. -##### Big O Notation (O-notation) +#### Big O Notation (O-notation) -The Big O notation represents an asymptotic upper bound, indicating the worst-case scenario for an algorithm's time or space complexity. Essentially, it signifies an upper limit on the growth of a function. +Big O notation gives an asymptotic upper bound on a function. It does not mean “worst case”: best-case, average-case, and worst-case running times are separate functions, and each can have an upper bound. -If we designate $f(n)$ as the actual complexity and $g(n)$ as the function in Big O notation, stating $f(n) = O(g(n))$ implies that $f(n)$, the time or space complexity of the algorithm, grows no faster than $g(n)$. +For non-negative functions, $f(n) = O(g(n))$ means there are constants $c > 0$ and $n_0$ such that $f(n) \le c g(n)$ for every $n \ge n_0$. Constant factors and a finite number of small inputs do not affect this bound. For instance, if an algorithm has a time complexity of $O(n)$, it signifies that the algorithm's running time does not grow more rapidly than a linear function of the input size, in the worst-case scenario. -0902bace-952d-4c80-9533-5706e28ef3e9 +![Asymptotic growth bound](https://github.com/user-attachments/assets/152fe1b7-3e0b-4a6d-b2d1-abf248ca90cf) -##### Big Omega Notation (Ω-notation) +#### Big Omega Notation (Ω-notation) -The Big Omega notation provides an asymptotic lower bound that expresses the best-case scenario for the time or space complexity of an algorithm. +Big Omega notation provides an asymptotic lower bound. It applies to whichever function is being analyzed, including a worst-case running-time function; it does not mean “best case”. -If $f(n) = Ω(g(n))$, this means that $f(n)$ grows at a rate that is at least as fast as $g(n)$. In other words, $f(n)$ does not grow slower than $g(n)$. +For non-negative functions, $f(n) = \Omega(g(n))$ means there are constants $c > 0$ and $n_0$ such that $f(n) \ge c g(n)$ for every $n \ge n_0$. -For example, if an algorithm has a time complexity of $Ω(n)$, it implies that the running time is at the bare minimum proportional to the input size in the best-case scenario. +For example, finding the maximum of an arbitrary unsorted array requires examining every element. Its running time is $\Omega(n)$ even on its best inputs in the usual comparison model. -d189ece7-e9c2-4797-8e0d-720336c4ba4a +![Asymptotic growth bound](https://github.com/user-attachments/assets/9984cad4-e131-4d52-bcad-8206b03e625f) -##### Theta Notation (Θ-notation) +#### Theta Notation (Θ-notation) -Theta notation offers a representation of the average-case scenario for an algorithm's time or space complexity. It sets an asymptotically tight bound, implying that the function grows neither more rapidly nor slower than the bound. +Theta notation gives an asymptotically tight bound: both an upper bound and a lower bound of the same order. It does not mean “average case”. -Stating $f(n) = Θ(g(n))$ signifies that $f(n)$ grows at the same rate as $g(n)$ under average circumstances. This indicates the time or space complexity is both at most and at least a linear function of the input size. +Stating $f(n) = \Theta(g(n))$ means that, for sufficiently large $n$, positive constants $c_1$ and $c_2$ satisfy $c_1g(n) \le f(n) \le c_2g(n)$. For example, $3n^2 + 2n + 1 = \Theta(n^2)$; the bound need not be linear. -ef39373a-8e6a-4e5b-832f-698b4dde7c7e +![Asymptotic growth bound](https://github.com/user-attachments/assets/bb11e34a-da8f-45a6-9eab-cbc05676a334) These notations primarily address the growth rate as the input size becomes significantly large. While they offer a high-level comprehension of an algorithm's performance, the actual running time in practice can differ based on various factors, such as the specific input data, the hardware or environment where the algorithm is operating, and the precise way the algorithm is implemented in the code. -#### Diving into Big O Notation Examples +### Diving into Big O Notation Examples -Big O notation is a practical tool for comparing the worst-case scenario of algorithm complexities. Here are examples of various complexities: +Big O notation is a practical tool for comparing growth bounds. The examples below use worst-case time unless an average or expected bound is stated. Here are examples of various complexities: -The point of this list isn’t to memorize it like a chant, it’s to build intuition. When you see nested loops, your brain should whisper “quadratic?” When you see repeated halving, it should whisper “logarithmic?” That intuition is what helps you spot performance problems early. +Use these growth rates to explain the work an algorithm performs. Nested loops may be quadratic, but their actual iteration bounds matter; repeated halving usually introduces a logarithm. Derive the bound from the work rather than from the visual shape of the code. -* The time complexity **$O(1)$**, known as constant time complexity, means that regardless of the input size, the algorithm performs its task in a fixed amount of time. A common example of this is retrieving an item by its index from an array or accessing a key-value pair in a hash map. -* When an algorithm has **$O(log n)$** time complexity, it operates logarithmically, meaning the time taken increases logarithmically with input size. As the input size doubles, the time taken only increases marginally. Binary search and operations on balanced binary trees are typical examples. -* An algorithm with **$O(n)$** time complexity exhibits linear behavior, where the running time scales directly with the input size. This is seen in simple, single-pass processes like iterating over an array or a linked list. -* In cases of **$O(n log n)$** time complexity, also called log-linear complexity, the running time grows both linearly and logarithmically with the input size. Sorting algorithms such as QuickSort, MergeSort, and HeapSort are prime examples of this complexity. -* With **$O(n^2)$** time complexity, the running time increases quadratically, often due to nested loops. Algorithms like Bubble Sort and Insertion Sort fall into this category. -* When an algorithm has **$O(n^3)$** time complexity, its running time scales cubically with the input size. This is common in algorithms involving three nested loops, such as naive matrix multiplication. -* **$O(2^n)$** represents exponential time complexity, where the running time doubles with each additional unit of input size. This is typical in brute-force algorithms like generating all subsets of a set or solving the Travelling Salesman Problem using a naive approach. +- The time complexity $O(1)$, known as constant time complexity, means that regardless of the input size, the algorithm performs its task in a fixed amount of time. A common example of this is retrieving an item by its index from an array or, under suitable hashing and load assumptions, an expected constant-time hash-table lookup. +- When an algorithm has $O(log n)$ time complexity, it operates logarithmically, meaning the time taken increases logarithmically with input size. As the input size doubles, the time taken only increases marginally. Binary search and operations on balanced binary trees are typical examples. +- An algorithm with $O(n)$ time complexity exhibits linear behavior, where the running time scales directly with the input size. This is seen in simple, single-pass processes like iterating over an array or a linked list. +- In cases of $O(n log n)$ time complexity, also called log-linear complexity, the running time grows both linearly and logarithmically with the input size. Merge sort and heap sort have this worst-case bound; randomized quicksort has this expected bound but can take quadratic time in the worst case. +- With $O(n^2)$ time complexity, the running time increases quadratically, often due to nested loops. Algorithms like Bubble Sort and Insertion Sort fall into this category. +- When an algorithm has $O(n^3)$ time complexity, its running time scales cubically with the input size. This is common in algorithms involving three nested loops, such as naive matrix multiplication. +- $O(2^n)$ represents exponential time complexity, where the running time doubles with each additional unit of input size. A set of $n$ elements has $2^n$ subsets. Visiting their inclusion/exclusion states takes $\Theta(2^n)$ work, while copying every subset explicitly takes $\Theta(n2^n)$ total time. Brute-force travelling salesperson search enumerates permutations and has factorial growth. + +When explicitly outputting all permutations or subsets, include the cost of writing every element of every result. Counting candidates alone can omit a factor of $n$. The graph below illustrates the growth of these different time complexities: @@ -387,44 +389,46 @@ Here is a summary cheat sheet: | Notation | Name | Meaning | Common Examples | | ------------- | ----------------- | ------------------------------------------------------- | --------------------------------------------------- | -| $O(1)$ | Constant time | Running time does not depend on input size $n$. | Array indexing, hash‐table lookup | +| $O(1)$ | Constant time | Running time does not depend on input size $n$. | Array indexing, expected hash-table lookup | | $O(\log n)$ | Logarithmic time | Time grows proportionally to the logarithm of $n$. | Binary search, operations on balanced BSTs | | $O(n)$ | Linear time | Time grows linearly with $n$. | Single loop over array, scanning for max/min | | $O(n \log n)$ | Linearithmic time | Combination of linear and logarithmic growth. | Merge sort, heap sort, FFT | | $O(n^2)$ | Quadratic time | Time grows proportional to the square of $n$. | Bubble sort, selection sort, nested loops | | $O(n^3)$ | Cubic time | Time grows proportional to the cube of $n$. | Naïve matrix multiplication (3 nested loops) | | $O(2^n)$ | Exponential time | Time doubles with each additional element in the input. | Recursive Fibonacci, brute‐force subset enumeration | -| $O(n!)$ | Factorial time | Time grows factorially with $n$. | Brute‐force permutation generation, TSP brute‐force | +| $O(n!)$ | Factorial time | Time grows factorially with $n$. | Permutation search leaves, TSP candidate tours | -#### Interpreting Big O Notation +### Interpreting Big O Notation -* We focus on the rate of growth rather than the exact number of operations, which is why constant factors are typically ignored. For example, the function $5n$ is expressed as **$O(n)$**, neglecting the constant factor of 5. -* When an algorithm has multiple terms, only the term with the fastest growth rate is considered important. For example, if the running time is $n^2 + n$, the time complexity simplifies to **$O(n^2)$**, since $n^2$ grows faster than $n$. -* Big O notation describes an upper limit on the growth rate of a function, meaning that if an algorithm has a time complexity of **$O(n)$**, it can also be described as $O(n^2)$ or higher. However, an algorithm with **$O(n^2)$** complexity cannot be described as **$O(n)$**, because Big O does not imply a lower bound on growth. -* Terms that grow as fast as or faster than **$n$** or **$log n$** dominate constant terms. For example, in the complexity **$O(n + k)$**, the term **$n$** dominates, simplifying the overall complexity to **$O(n)$**. +- We focus on the rate of growth rather than the exact number of operations, which is why constant factors are typically ignored. For example, the function $5n$ is expressed as $O(n)$, neglecting the constant factor of 5. +- When an algorithm has multiple terms, only the term with the fastest growth rate is considered important. For example, if the running time is $n^2 + n$, the time complexity simplifies to $O(n^2)$, since $n^2$ grows faster than $n$. +- Big O notation describes an upper limit on the growth rate of a function, meaning that if an algorithm has a time complexity of $O(n)$, it can also be described as $O(n^2)$ or higher. An $O(n^2)$ upper bound alone does not rule out $O(n)$; a function such as $n$ satisfies both. A tight $\Theta(n^2)$ bound does rule out $O(n)$. +- Terms that grow as fast as or faster than $n$ or $log n$ dominate constant terms. For example, $O(n + k)$ simplifies to $O(n)$ if $k$ is constant or $k = O(n)$. If $k$ is an independent input size, retain both parameters. -The main “do” here is to treat Big O as a *communication tool*. It lets you compare approaches and explain choices to others. The main “don’t” is to weaponize it, “this is $O(n)$” doesn’t automatically beat “this is $O(n \log n)$” if constants, data sizes, or real constraints tell a different story. +The main “do” here is to treat Big O as a communication tool. It lets you compare approaches and explain choices to others. Do not weaponize it, “this is $O(n)$” doesn’t automatically beat “this is $O(n \log n)$” if constants, data sizes, or real constraints tell a different story. -#### Can every problem have an O(1) algorithm? +### Can every problem have an O(1) algorithm? -* Not every problem has an algorithm that can solve it, irrespective of the complexity. For instance, the Halting Problem is undecidable, no algorithm can accurately predict whether a given program will halt or run indefinitely on every possible input. -* Sometimes, we can create an illusion of $O(1)$ complexity by precomputing the results for all possible inputs and storing them in a lookup table (like a hash table). Then, we can solve the problem in constant time by directly retrieving the result from the table. This approach, known as memoization or caching, is limited by memory constraints and is only practical when the number of distinct inputs is small and manageable. -* Often, the lower bound complexity for a class of problems is $O(n)$ or $O(nlogn)$. This bound represents problems where you at least have to examine each element once (as in the case of $O(n)$ ) or perform a more complex operation on every input (as in $O(nlogn)$ ), like sorting. Under certain conditions or assumptions, a more efficient algorithm might be achievable. +- Not every problem has an algorithm that can solve it, irrespective of the complexity. For instance, the Halting Problem is undecidable, no algorithm can accurately predict whether a given program will halt or run indefinitely on every possible input. +- Sometimes, we can create an illusion of $O(1)$ complexity by precomputing the results for all possible inputs and storing them in a lookup table (like a hash table). Then, we can solve the problem in constant time by directly retrieving the result from the table. Precomputation is limited by its construction cost and memory use, and is practical only for a manageable input domain. Memoization instead caches results on demand. Lookup is not automatically constant time if hashing or comparing a variable-length key is expensive. +- Lower bounds use $\Omega$: reading all $n$ input elements requires $\Omega(n)$ work, and comparison sorting requires $\Omega(n\log n)$ comparisons in the worst case. These bounds depend on the computational model. Counting sort can avoid the comparison-sorting lower bound by exploiting a bounded integer key range. -There is a common beginner fantasy: “surely there’s always a constant-time trick.” Sometimes there isn’t, and that’s not a failure, it’s a reality of computation. The “do” is to accept lower bounds and design within them. The “don’t” is to chase miraculous optimizations when the problem inherently requires reading the input. +Some tasks cannot be reduced to constant time because they inherently require examining the input or producing many outputs. Identify the applicable lower bound before trying to optimize beyond it. -### Recognising $O(\log n)$ and $O(n \log n)$ running-times +## Recognising $O(\log n)$ and $O(n \log n)$ running-times -The growth rate of an algorithm almost always comes from **how quickly the remaining work shrinks** as the algorithm executes. Two common patterns are: +The number of iterations often depends on how quickly the remaining work shrinks. The following patterns distinguish linear, logarithmic, and combined growth. This is one of the most useful instincts you can build: look at the loop, ask what variable is changing, and ask whether it’s shrinking by subtraction (linear) or division (logarithmic). If you can do that, you can “feel” complexity without formal proofs. | Pattern | Typical loop behaviour | Resulting time-complexity | | ------------------------------------------------------------------------- | ---------------------------------------------------- | ------------------------- | -| *Halve (or otherwise divide) the problem each step* | $n \to n/2 \to n/4 \dots$ | $\Theta(\log n)$ | -| *Do a linear amount of work, but each unit of work is itself logarithmic* | outer loop counts down one by one, inner loop halves | $\Theta(n \log n)$ | +| Halve (or otherwise divide) the problem each step | $n \to n/2 \to n/4 \dots$ | $\Theta(\log n)$ | +| Do a linear amount of work, but each unit of work is itself logarithmic | outer loop counts down one by one, inner loop halves | $\Theta(n \log n)$ | + +Assume non-negative integer inputs and constant-time loop bodies. The logarithmic formulas apply for $n \ge 1$; at $n = 0$, these loops perform no iterations. -Below are four miniature algorithms written in language-neutral *pseudocode* (no Python syntax), followed by the intuition behind each bound. +Below are four miniature algorithms written in language-neutral pseudocode (no Python syntax), followed by the intuition behind each bound. I. Linear - $\Theta(n)$ @@ -436,7 +440,7 @@ procedure Linear(n) end procedure ``` -*Work left* drops by **1** each pass, so the loop executes exactly $n$ times. +Work left drops by 1 each pass, so the loop executes exactly $n$ times. II. Logarithmic - $\Theta(\log n)$ @@ -450,7 +454,7 @@ end procedure Each pass discards half of the remaining input, so only $\lfloor\log_2 n\rfloor + 1$ iterations are needed. -*Common real examples: binary search, finding the height of a complete binary tree.* +Common real examples: binary search, finding the height of a complete binary tree. III. Linear-logarithmic - $\Theta(n \log n)$ @@ -467,9 +471,9 @@ procedure LinearLogarithmic(n) end procedure ``` -* **Outer loop:** $n$ iterations. -* **Inner loop:** $\lfloor\log_2 n\rfloor + 1$ iterations for each outer pass. -* Total work $\approx n \cdot \log n$. +- Outer loop: $n$ iterations. +- Inner loop: $\lfloor\log_2 n\rfloor + 1$ iterations for each outer pass. +- Total work $\approx n \cdot \log n$. Classic real-world instances: mergesort, heapsort, many divide-and-conquer algorithms, building a heap then doing $n$ delete-min operations. @@ -488,23 +492,23 @@ procedure LogSquared(n) end procedure ``` -Both loops cut their control variable in half, so each contributes a $\log n$ factor, giving $\log^2 n$. Such bounds appear in some advanced data-structures (e.g., range trees) where *two* independent logarithmic dimensions are traversed. +Both loops cut their control variable in half, so each contributes a $\log n$ factor, giving $\log^2 n$. Such bounds appear in some advanced data-structures (e.g., range trees) where two independent logarithmic dimensions are traversed. Rules of thumb: -1. **Log factors come from repeatedly shrinking a quantity by a constant factor.** Any loop of the form `while x > 1: x \gets x / c` (for constant $c > 1$) takes $\Theta(\log x)$ steps. -2. **Multiplying two independent loops multiplies their costs.** An outer loop that counts $n$ times and an inner loop that counts $\log n$ times gives $n \cdot \log n$. -3. **Divide-and-conquer often yields $n \log n$.** Splitting the problem into a constant number of sub-problems of half size and doing $\Theta(n)$ work to combine them recurs to the *Master Theorem* case $T(n) = 2,T\bigl(n/2\bigr) + \Theta(n) = \Theta(n \log n).$ -4. **Nested logarithmic loops stack.** Two independent halving loops give $\log^2 n$; three give $\log^3 n$, and so on. +1. Log factors come from repeatedly shrinking a quantity by a constant factor. Any loop of the form `while x > 1: x \gets x / c` (for constant $c > 1$) takes $\Theta(\log x)$ steps. +2. Multiplying two independent loops multiplies their costs. An outer loop that counts $n$ times and an inner loop that counts $\log n$ times gives $n \cdot \log n$. +3. Divide-and-conquer often yields $n \log n$. Splitting the problem into two subproblems of half the size and doing $\Theta(n)$ work to combine them recurs to the Master Theorem case $T(n) = 2T\bigl(n/2\bigr) + \Theta(n) = \Theta(n \log n).$ +4. Nested logarithmic loops stack. Two independent halving loops give $\log^2 n$; three give $\log^3 n$, and so on. A “do” for interviews and real design reviews: walk through this reasoning out loud. It shows you understand growth, not just symbols. A “don’t” is to overfit to the notation, always tie the bound back to the loop behavior. -### Misconceptions +## Misconceptions -These misconceptions are worth calling out because they keep people from learning the *useful* parts. The goal isn’t to become a theoretician; it’s to become someone who can write software that keeps working as the world scales. That’s why a little complexity intuition pays off so heavily. +These misconceptions are worth calling out because they keep people from learning the useful parts. The goal isn’t to become a theoretician; it’s to become someone who can write software that keeps working as the world scales. That’s why a little complexity intuition pays off so heavily. -* Formal proof of Big O complexity is rarely necessary in everyday programming or software engineering. However, having a fundamental understanding of theoretical complexity is important when selecting appropriate algorithms, especially when solving complex problems. It aids in understanding the trade-offs between different solutions and predicting the algorithm's performance. -* It's not required to assign Big O complexity for every single function or chunk of code you write. However, if you're dealing with large datasets or performance-critical applications, understanding the time and space complexity of your algorithms and data structures can help you make informed decisions about scalability and efficiency. -* Big O notation is not a predictor of an algorithm's precise running time for a given input size. Instead, it provides an upper bound on the growth rate of the algorithm's running time or space usage as the input size increases. It's a tool to compare the scalability of different algorithms, ignoring implementation details and specific characteristics of the input data. -* In real-world scenarios, the actual running time of an algorithm can be influenced by various factors, including the specific characteristics of the input data, the efficiency of the implementation, and the hardware and software environment in which it runs. Big O notation doesn't account for these factors. -* While it's good to consider performance, it shouldn't come at the cost of code readability and maintainability. Clear, simple code is often more valuable than highly optimized code, especially if the optimizations complicate the code without offering substantial performance improvements. Instead of optimizing every detail, focus on identifying and addressing the actual bottlenecks in your code, as these are the areas where optimizations can make a significant difference. +- Formal proof of Big O complexity is rarely necessary in everyday programming or software engineering. However, having a fundamental understanding of theoretical complexity is important when selecting appropriate algorithms, especially when solving complex problems. It aids in understanding the trade-offs between different solutions and predicting the algorithm's performance. +- It's not required to assign Big O complexity for every single function or chunk of code you write. However, if you're dealing with large datasets or performance-critical applications, understanding the time and space complexity of your algorithms and data structures can help you make informed decisions about scalability and efficiency. +- Big O notation is not a predictor of an algorithm's precise running time for a given input size. Instead, it provides an upper bound on the growth rate of the algorithm's running time or space usage as the input size increases. It's a tool to compare the scalability of different algorithms, ignoring implementation details and specific characteristics of the input data. +- In real-world scenarios, the actual running time of an algorithm can be influenced by various factors, including the specific characteristics of the input data, the efficiency of the implementation, and the hardware and software environment in which it runs. Big O notation doesn't account for these factors. +- While it's good to consider performance, it shouldn't come at the cost of code readability and maintainability. Clear, simple code is often more valuable than highly optimized code, especially if the optimizations complicate the code without offering substantial performance improvements. Instead of optimizing every detail, focus on identifying and addressing the actual bottlenecks in your code, as these are the areas where optimizations can make a significant difference. diff --git a/notes/brain_teasers.md b/notes/brain_teasers.md index af25a2b..b05fe0f 100644 --- a/notes/brain_teasers.md +++ b/notes/brain_teasers.md @@ -1,234 +1,236 @@ -## Solving Programming Brain Teasers +# Solving Programming Brain Teasers Programming puzzles and brain teasers are a fun way to sharpen your coding and problem-solving skills. You’ll often see them in technical interviews, where they’re used to test how you think, analyze problems, and come up with efficient solutions. To do well, it helps to practice and build solid strategies for tackling these kinds of challenges. -The hidden goal of brain teasers is rarely “get the answer.” It’s “show your process.” Interviewers (and your future self) want to see that you can move from confusion to clarity: restate the problem, choose a direction, test assumptions, and iterate. If you treat each puzzle like a tiny engineering project, understand, design, validate, optimize, you’ll come across as calm, capable, and methodical. +A useful solution explains both the answer and the reasoning. Restate the problem, identify constraints, choose an approach, test its assumptions, and then improve it. The sections below connect common problem shapes with the data structures and algorithms that support them. -### General Strategies +## General Strategies When tackling programming puzzles, consider the following strategies: -These strategies work best when you use them in the right order. A common trap is trying to be clever too early, jumping straight into optimization, fancy data structures, or tricky math. The “do” is to build a correct baseline, then improve it deliberately. The “don’t” is to optimize a solution you don’t fully understand yet. +These strategies work best when you use them in the right order. A common trap is trying to be clever too early, jumping straight into optimization, fancy data structures, or tricky math. Build a correct baseline, then improve it deliberately. Do not optimize a solution you don’t fully understand yet. -* Start with a *simple solution* to get a clear grasp of the problem, this often highlights what really needs optimizing later. -* Use *unit tests* to make sure your code works across different inputs, especially tricky edge cases. -* Think about your algorithm’s *time and space complexity* to understand how efficient it is and where it can be improved. -* Pick the *right data structure* (like an array, hash map, or tree) to make your solution faster and easier to reason about. -* Break the problem into *smaller pieces* so it feels less overwhelming and easier to solve step by step. -* Try both *recursive* and *iterative* approaches, you’ll often find one fits the problem more naturally. -* Always keep *edge cases* and *constraints* in mind so your solution doesn’t break with unusual inputs. -* Apply *targeted optimization* only where it matters, improve performance without making the code messy. +- Start with a simple solution to get a clear grasp of the problem, this often highlights what really needs optimizing later. +- Use unit tests to make sure your code works across different inputs, especially tricky edge cases. +- Think about your algorithm’s time and space complexity to understand how efficient it is and where it can be improved. +- Pick the right data structure (like an array, hash map, or tree) to make your solution faster and easier to reason about. +- Break the problem into smaller pieces so it feels less overwhelming and easier to solve step by step. +- Try both recursive and iterative approaches, you’ll often find one fits the problem more naturally. +- Always keep edge cases and constraints in mind so your solution doesn’t break with unusual inputs. +- Apply targeted optimization only where it matters, improve performance without making the code messy. -A practical mindset that ties all of these together: always be able to answer two questions at any moment, **“What do I know is true?”** and **“What am I trying next?”** That keeps you moving forward, even when the puzzle feels unfamiliar. +A practical mindset that ties all of these together: always be able to answer two questions at any moment, “What do I know is true?” and “What am I trying next?” That keeps you moving forward, even when the puzzle feels unfamiliar. -### Data Structures +## Data Structures A solid grasp of data structures is helpful for effective programming. Below are some practical strategies and tips to help you use them more confidently and efficiently. -Data structures are where puzzles become manageable. Most brain teasers aren’t about memorizing rare tricks, they’re about recognizing a pattern and picking a container that makes the pattern easy. A good “do” is to ask: *What operations am I doing most, lookup, insert, delete, min/max, traversal?* The right structure makes those operations feel natural. - -#### Working with Arrays - -Arrays are basic data structures that store elements in a continuous block of memory, making it easy to access any element quickly. - -Arrays are the default starting point because they’re simple and fast. The “why” behind many array techniques is that arrays give you **indexes**, and indexes give you powerful structure: order, boundaries, and the ability to use two pointers, binary search, and prefix computations. The main “don’t” is forgetting that many operations look cheap but hide expensive shifting or copying. - -* Sorting an array can often simplify many problems, with algorithms like Quick Sort and Merge Sort offering efficient $O(n \log n)$ time complexity. For nearly sorted or small arrays, *Insertion Sort* might be a better option due to its simplicity and efficiency in such cases. -* In sorted arrays, *binary search* provides a fast way to find elements or their positions, working in $O(\log n)$. Be cautious with mid-point calculations in languages that may experience integer overflow due to fixed-size integer types. -* The *two-pointer* technique uses two indices, typically starting from opposite ends of the array, to solve problems involving pairs or triplets, like finding two numbers that sum to a target. It helps optimize both time and space efficiency. -* The *sliding window* technique is effective for solving subarray or substring problems, such as finding the longest substring without repeating characters. It maintains a dynamic subset of the array while iterating, improving overall efficiency. -* *Prefix sums* enable fast range sum queries after preprocessing the array in $O(n)$. Likewise, difference arrays allow efficient range updates without the need to modify individual elements one by one. -* In-place operations modify the array directly without using extra memory. This method saves space but requires careful handling to avoid unintended side effects on other parts of the program. -* When dealing with duplicates, it’s important to adjust the algorithm to handle them appropriately. For example, the two-pointer technique may need to skip duplicates to prevent redundant results or errors. -* When working with large arrays, it’s important to be mindful of memory usage, as they can consume a lot of space. To optimize, try to minimize the space complexity by using more memory-efficient data structures or algorithms. For instance, instead of storing a full array of values, consider using a *sliding window* or *in-place modifications* to avoid extra memory allocation. Additionally, analyze the space complexity of your solution and check for operations that create large intermediate data structures, which can lead to excessive memory consumption. In constrained environments, tools like memory profiling or checking the space usage of your program (e.g., using Python’s `sys.getsizeof()`) can help you identify areas for improvement. -* When using dynamic arrays, it’s helpful to allow automatic resizing, which lets the array expand or shrink based on the data size. This avoids the need for manual memory management and improves flexibility. -* Resizing arrays frequently can be costly in terms of time complexity. A more efficient approach is to resize the array exponentially, such as doubling its size, rather than resizing it by a fixed amount each time. -* To avoid unnecessary memory usage, it's important to pass arrays by reference (or using pointers in some languages) when possible, instead of copying the entire array for each function call. -* For arrays with many zero or null values, using sparse arrays or hash maps can be useful. This allows you to store only non-zero values, saving memory when dealing with large arrays that contain mostly empty data. -* When dealing with multi-dimensional arrays, flattening them into a one-dimensional array can make it easier to perform operations, but be aware that this can temporarily increase memory usage. -* To improve performance, accessing memory in contiguous blocks is important. Random access patterns may lead to cache misses, which can slow down operations, so try to access array elements sequentially when possible. -* The `bisect` module helps maintain sorted order in a list by finding the appropriate index for inserting an element or by performing binary searches. -* Use `bisect.insort()` to insert elements into a sorted list while keeping it ordered. -* Use `bisect.bisect_left()` or `bisect.bisect_right()` to find the index where an element should be inserted. -* Don’t use on unsorted lists or when frequent updates are needed, as maintaining order can be inefficient. -* Binary search operations like `bisect_left()` are `O(log n)`, but `insort()` can be `O(n)` due to shifting elements. +Data structures are where puzzles become manageable. Most brain teasers aren’t about memorizing rare tricks, they’re about recognizing a pattern and picking a container that makes the pattern easy. A good “do” is to ask: What operations am I doing most, lookup, insert, delete, min/max, traversal? The right structure makes those operations feel natural. + +### Working with Arrays + +Arrays are basic data structures that store elements in a contiguous block of memory, making it easy to access any element quickly. + +Arrays are the default starting point because they’re simple and fast. The “why” behind many array techniques is that arrays give you indexes, and indexes give you powerful structure: order, boundaries, and the ability to use two pointers, binary search, and prefix computations. The main “don’t” is forgetting that many operations look cheap but hide expensive shifting or copying. + +- Sorting an array can often simplify many problems, with merge sort guaranteeing $O(n\log n)$ time and randomized quicksort providing that expected bound, with a quadratic worst case. For nearly sorted or small arrays, Insertion Sort might be a better option due to its simplicity and efficiency in such cases. +- In sorted arrays, binary search provides a fast way to find elements or their positions, working in $O(\log n)$. Be cautious with mid-point calculations in languages that may experience integer overflow due to fixed-size integer types. +- The two-pointer technique uses two indices, typically starting from opposite ends of the array, to solve problems involving pairs or triplets, like finding two numbers that sum to a target. It helps optimize both time and space efficiency. +- The sliding window technique is effective for solving subarray or substring problems, such as finding the longest substring without repeating characters. It maintains a contiguous window of the array while iterating, improving overall efficiency. +- Prefix sums enable fast range sum queries after preprocessing the array in $O(n)$. Likewise, difference arrays allow efficient range updates without the need to modify individual elements one by one. +- In-place operations modify the array directly, usually with constant auxiliary storage. Some algorithms described as in-place still use a recursion stack. This method saves space but requires careful handling to avoid unintended side effects on other parts of the program. +- When dealing with duplicates, it’s important to adjust the algorithm to handle them appropriately. For example, the two-pointer technique may need to skip duplicates to prevent redundant results or errors. +- When working with large arrays, it’s important to be mindful of memory usage, as they can consume a lot of space. To optimize, try to minimize the space complexity by using more memory-efficient data structures or algorithms. For instance, instead of storing a full array of values, consider using a sliding window or in-place modifications to avoid extra memory allocation. Additionally, analyze the space complexity of your solution and check for operations that create large intermediate data structures, which can lead to excessive memory consumption. In constrained environments, tools like memory profiling or checking the space usage of your program (e.g., using Python’s `sys.getsizeof()`) can help you identify areas for improvement. +- When using dynamic arrays, it’s helpful to allow automatic resizing, which lets the array expand or shrink based on the data size. This avoids the need for manual memory management and improves flexibility. +- Resizing arrays frequently can be costly in terms of time complexity. A more efficient approach is to resize the array exponentially, such as doubling its size, rather than resizing it by a fixed amount each time. +- To avoid unnecessary memory usage, it's important to pass arrays by reference (or using pointers in some languages) when possible, instead of copying the entire array for each function call. +- For arrays with many zero or null values, using sparse arrays or hash maps can be useful. This allows you to store only non-zero values, saving memory when dealing with large arrays that contain mostly empty data. +- When dealing with multi-dimensional arrays, flattening them into a one-dimensional array can make it easier to perform operations, but be aware that this can temporarily increase memory usage. +- To improve performance, accessing memory in contiguous blocks is important. Random access patterns may lead to cache misses, which can slow down operations, so try to access array elements sequentially when possible. +- The `bisect` module helps maintain sorted order in a list by finding the appropriate index for inserting an element or by performing binary searches. +- Use `bisect.insort()` to insert elements into a sorted list while keeping it ordered. +- Use `bisect.bisect_left()` or `bisect.bisect_right()` to find the index where an element should be inserted. +- Don’t use on unsorted lists or when frequent updates are needed, as maintaining order can be inefficient. +- Binary search operations like `bisect_left()` are `O(log n)`, but `insort()` can be `O(n)` due to shifting elements. A small but high-impact “do”: always sanity-check whether sorting is allowed. Sorting often unlocks elegant solutions, but it changes order. If the problem cares about original positions, keep track of indices or consider a hash-based approach instead. -#### Working with Strings +### Working with Strings Strings, as sequences of characters, often require special handling due to their immutable nature in some languages and the variety of operations performed on them. -String problems often look like array problems with extra rules, immutability, encoding, and more expensive slicing. The “why” here is performance surprises: what looks like a tiny operation (like concatenation or slicing) can secretly allocate lots of memory. A good “do” is to think in terms of *streams* and *indexes* when strings get large. - -* In languages with *immutable strings*, concatenating inside a loop with `+` can be very slow, $O(n^2)$ time. Instead, use tools like `StringBuilder` in Java or `''.join(list_of_strings)` in Python to bring it down to $O(n)$. -* For *string searching*, algorithms like Knuth-Morris-Pratt (KMP), Rabin-Karp, or Boyer-Moore are much faster than the naive $O(nm)$ approach. -* *Tries (prefix trees)* let you store and retrieve strings with shared prefixes efficiently. They’re common in autocomplete, spell-checkers, and IP routing. -* *Regular expressions* are powerful for pattern matching, but they need careful design, badly written patterns can cause exponential slowdowns. -* Always keep *Unicode and character encoding* (UTF-8, ASCII, etc.) in mind when handling strings, especially with international text or external data. It prevents bugs and ensures proper compatibility. -* For *anagram checks*, compare character counts with arrays or hash maps. For *palindrome checks*, the two-pointer method or comparing against the reversed string gives a simple, efficient solution. -* *Substring and slicing operations* can vary in cost depending on the language. In some (like Python), slices create new strings, while in others (like Java, newer versions), substrings copy data. Always be mindful of hidden overhead. -* When dealing with *very large text processing*, consider using *streaming* or *buffered reading* instead of loading everything into memory at once. -* *Hashing strings* (e.g., rolling hash) is useful for fast lookups, substring checks, or detecting duplicates, especially in algorithms like Rabin-Karp or for building hash-based data structures. -* *Suffix arrays* and *suffix trees* are powerful for advanced text processing tasks like pattern matching, longest common substring, and text compression. -* *String interning* (pooling identical strings) can save memory and speed up equality checks in languages like Java, but overuse can increase GC pressure. -* For *case-insensitive comparisons*, normalize strings first (e.g., `.lower()` in Python, `.toLowerCase()` in Java), but remember locale-specific rules (e.g., Turkish dotted/dotless "i"). -* Use *string formatting libraries* (e.g., Python f-strings, Java’s `String.format`) instead of manual concatenation for cleaner, safer, and often more efficient code. -* Be cautious with *regex backtracking traps*, even professional developers can accidentally write patterns that blow up on certain inputs. Tools like regex debuggers can help test patterns. -* *Memory trade-offs* matter: a trie might be faster than a hash map for prefix searches but can use much more memory. -* For *streaming algorithms* (like real-time log processing), sliding window techniques combined with hash maps or deques can efficiently handle substring or frequency problems. +String problems often look like array problems with extra rules, immutability, encoding, and more expensive slicing. The “why” here is performance surprises: what looks like a tiny operation (like concatenation or slicing) can secretly allocate lots of memory. A good “do” is to think in terms of streams and indexes when strings get large. + +- In languages with immutable strings, concatenating inside a loop with `+` can be very slow, $O(n^2)$ time. Instead, use tools like `StringBuilder` in Java or `''.join(list_of_strings)` in Python to bring it down to $O(n)$. +- For string searching, KMP guarantees $O(n+m)$ time. Rabin–Karp and Boyer–Moore can also avoid much of the work of naive matching, but their bounds depend on hashing assumptions and the chosen variant. Small inputs may favor a simpler approach. +- Tries (prefix trees) let you store and retrieve strings with shared prefixes efficiently. They’re common in autocomplete, spell-checkers, and IP routing. +- Regular expressions are powerful for pattern matching, but they need careful design, badly written patterns can cause exponential slowdowns. +- Always keep Unicode and character encoding (UTF-8, ASCII, etc.) in mind when handling strings, especially with international text or external data. It prevents bugs and ensures proper compatibility. +- For anagram checks, compare character counts with arrays or hash maps. For palindrome checks, the two-pointer method or comparing against the reversed string gives a simple, efficient solution. +- Substring and slicing operations can vary in cost depending on the language. In some (like Python), slices create new strings, while in others (like Java, newer versions), substrings copy data. Always be mindful of hidden overhead. +- When dealing with very large text processing, consider using streaming or buffered reading instead of loading everything into memory at once. +- Hashing strings (e.g., rolling hash) is useful for fast lookups, substring checks, or detecting duplicates, especially in algorithms like Rabin-Karp or for building hash-based data structures. +- Suffix arrays and suffix trees are powerful for advanced text processing tasks like pattern matching, longest common substring, and text compression. +- String interning (pooling identical strings) can save memory and speed up equality checks in languages like Java, but overuse can increase GC pressure. +- For case-insensitive comparisons, normalize strings first (e.g., `.lower()` in Python, `.toLowerCase()` in Java), but remember locale-specific rules (e.g., Turkish dotted/dotless "i"). +- Use string formatting libraries (e.g., Python f-strings, Java’s `String.format`) instead of manual concatenation for cleaner, safer, and often more efficient code. +- Be cautious with regex backtracking traps, even professional developers can accidentally write patterns that blow up on certain inputs. Tools like regex debuggers can help test patterns. +- Memory trade-offs matter: a trie might be faster than a hash map for prefix searches but can use much more memory. +- For streaming algorithms (like real-time log processing), sliding window techniques combined with hash maps or deques can efficiently handle substring or frequency problems. One useful “don’t” with strings: don’t assume “characters” are always one byte or one visible symbol. If Unicode matters, be explicit about whether you mean bytes, code points, or grapheme clusters, many bugs come from mixing those levels. -#### Working with Linked Lists +### Working with Linked Lists Linked lists are dynamic data structures consisting of nodes that contain data and references to the next (and possibly previous) nodes. -Linked lists show up in puzzles because they force pointer thinking: you can’t jump around by index, so you learn to solve problems with *structure* instead of random access. The “do” is to rely on pointer patterns (fast/slow, dummy head, two-list merge). The “don’t” is to treat a list like an array, if you need frequent random access, you probably chose the wrong structure. +Linked lists show up in puzzles because they force pointer thinking: you can’t jump around by index, so you learn to solve problems with structure instead of random access. Rely on pointer patterns (fast/slow, dummy head, two-list merge). Do not treat a list like an array, if you need frequent random access, you probably chose the wrong structure. -* Choosing the *right type of linked list* depends on the problem requirements. A singly linked list is more memory-efficient and works well when traversal is only needed in one direction. A doubly linked list, on the other hand, is better suited when you need to move both forward and backward or when frequent insertions and deletions at both ends are required. -* *Insertions and deletions* in a linked list are efficient if the position is already known, especially at the head or tail, where they can be done in constant time, $O(1)$. When deleting, it’s crucial to correctly update the previous node’s pointer so the list structure remains intact. -* *Reversing a singly linked list* is done by iterating through the nodes and redirecting each node’s `next` pointer to its predecessor. This takes $O(n)$ time and requires only $O(1)$ extra space, making it a very efficient operation. -* *Cycle detection* can be handled with Floyd’s Tortoise and Hare algorithm. By moving two pointers at different speeds, one slow and one fast, any cycle will eventually cause the two pointers to meet. -* Handling *edge cases* is essential. For example, empty lists, single-node lists, and lists that contain cycles all require special care. Always check for null pointers to avoid runtime errors, especially in languages without built-in safety checks. -* In environments with *manual memory management*, don’t forget to explicitly free memory for nodes that are no longer in use. This prevents memory leaks and helps keep resource usage efficient. -* The *fast and slow pointer technique* (also called the two-pointer method) is widely used with linked lists. Besides cycle detection, it can be applied to find the middle of a list in a single pass: the slow pointer moves one step at a time, while the fast pointer moves two steps. When the fast pointer reaches the end, the slow pointer will be at the middle. -* *Finding the kth node from the end* can also be solved with two pointers. By moving the fast pointer $k$ steps ahead, then advancing both fast and slow pointers together, the slow pointer will land exactly on the kth node from the end. This avoids the need to first compute the length of the list. -* *Merging two sorted linked lists* can be done efficiently using pointers to walk through both lists simultaneously, always choosing the smaller current node to attach next. This is commonly used as part of merge sort for linked lists. -* *Splitting a linked list* into halves is another use case of the fast/slow pointer method. When the fast pointer reaches the end, the slow pointer indicates the midpoint, allowing the list to be divided for algorithms like merge sort. -* *Removing duplicates* from a linked list may require extra memory (such as a hash set) if the list is unsorted. For a sorted linked list, duplicates can be removed in-place by checking adjacent nodes. -* *Palindromic linked lists* can be checked by finding the middle of the list (fast/slow pointer), reversing the second half, and then comparing both halves node by node. -* *Intersection detection* between two linked lists can be solved by advancing two pointers through both lists. If one pointer reaches the end, redirect it to the start of the other list. If the lists intersect, the pointers will eventually meet at the intersection node. -* *Rotating a linked list* by $k$ positions requires computing the length, connecting the tail to the head to form a cycle, and then breaking the cycle at the right position to create the rotated list. -* For very large datasets, *linked lists vs arrays* trade-offs become important. Linked lists excel at dynamic memory usage and frequent insertions/deletions, but they have poorer cache locality and slower random access compared to arrays. +- Choosing the right type of linked list depends on the problem requirements. A singly linked list is more memory-efficient and works well when traversal is only needed in one direction. A doubly linked list, on the other hand, is better suited when you need to move both forward and backward or when frequent insertions and deletions at both ends are required. +- Insertion and deletion take $O(1)$ when all required links are available. A singly linked list needs the predecessor to remove an arbitrary node; removing its tail normally takes $O(n)$ even with a tail pointer. A doubly linked list can remove a known node in constant time. Update neighboring links and head/tail pointers consistently. +- Reversing a singly linked list is done by iterating through the nodes and redirecting each node’s `next` pointer to its predecessor. This takes $O(n)$ time and requires only $O(1)$ extra space, making it a very efficient operation. +- Cycle detection can be handled with Floyd’s Tortoise and Hare algorithm. By moving two pointers at different speeds, one slow and one fast, any cycle will eventually cause the two pointers to meet. +- Handling edge cases is essential. For example, empty lists, single-node lists, and lists that contain cycles all require special care. Always check for null pointers to avoid runtime errors, especially in languages without built-in safety checks. +- In environments with manual memory management, don’t forget to explicitly free memory for nodes that are no longer in use. This prevents memory leaks and helps keep resource usage efficient. +- The fast and slow pointer technique (also called the two-pointer method) is widely used with linked lists. Besides cycle detection, it can be applied to find the middle of a list in a single pass: the slow pointer moves one step at a time, while the fast pointer moves two steps. When the fast pointer reaches the end, the slow pointer will be at the middle. +- Finding the kth node from the end can also be solved with two pointers. By moving the fast pointer $k$ steps ahead, then advancing both fast and slow pointers together, the slow pointer will land exactly on the kth node from the end. This avoids the need to first compute the length of the list. +- Merging two sorted linked lists can be done efficiently using pointers to walk through both lists simultaneously, always choosing the smaller current node to attach next. This is commonly used as part of merge sort for linked lists. +- Splitting a linked list into halves is another use case of the fast/slow pointer method. When the fast pointer reaches the end, the slow pointer indicates the midpoint, allowing the list to be divided for algorithms like merge sort. +- Removing duplicates from a linked list may require extra memory (such as a hash set) if the list is unsorted. For a sorted linked list, duplicates can be removed in-place by checking adjacent nodes. +- Palindromic linked lists can be checked by finding the middle of the list (fast/slow pointer), reversing the second half, and then comparing both halves node by node. +- Intersection detection between two linked lists can be solved by advancing two pointers through both lists. If one pointer reaches the end, redirect it to the start of the other list. If the lists intersect, the pointers will eventually meet at the intersection node. +- Rotating a linked list by $k$ positions requires computing the length, connecting the tail to the head to form a cycle, and then breaking the cycle at the right position to create the rotated list. +- For very large datasets, linked lists vs arrays trade-offs become important. Linked lists excel at dynamic memory usage and frequent insertions/deletions, but they have poorer cache locality and slower random access compared to arrays. -A small “do” that prevents a lot of bugs: when manipulating linked lists, use a **dummy head** for operations near the front. It simplifies edge cases like deleting the first node or building a new list. +A small “do” that prevents a lot of bugs: when manipulating linked lists, use a dummy head for operations near the front. It simplifies edge cases like deleting the first node or building a new list. -#### Working with Heaps +### Working with Heaps Heaps tend to appear when the problem sounds like: “repeatedly get the best next thing.” If your loop is “pick minimum/maximum, update, repeat,” a heap often turns something expensive into something smooth. The “don’t” is using a heap when you actually need fast membership checks or deletion by value, heaps aren’t designed for that without extra indexing. -* When you frequently need the next smallest or largest item, a *priority queue* offers O(1) peek and O(log n) updates, whereas rescanning a list is O(n) per step; e.g., choosing the earliest-deadline task. -* For practical implementations, an array-backed *binary heap* is cache-friendly with O(n) build and O(log n) push/pop, whereas pointer trees add overhead; e.g., worker queues in backend services. -* If the requirement is to repeatedly extract the minimum, a *min-heap* returns it quickly as new items arrive, whereas resorting after each insertion wastes time; e.g., streaming tasks ordered by due time. -* When tracking the current maximum continuously, a *max-heap* maintains the top element efficiently, whereas maintaining a full sorted list is unnecessary; e.g., live high-bid tracking in auctions. -* For top-k results on a stream, a *fixed-size heap* of k items yields O(log k) updates and bounded memory, whereas storing everything and sorting is slower and larger; e.g., top 100 error types today. -* To keep a running median, two *heaps* (max-heap lower half, min-heap upper) give O(log n) updates and O(1) median queries, whereas recomputing from all values is O(n); e.g., live latency dashboards. -* In shortest-path search, an *indexed priority queue* supports decrease-key in O(log n), whereas a plain heap needs lazy deletes that add overhead; e.g., Dijkstra on road networks. -* During k-way merging, a *min-heap* holding one head per list outputs the next item in O(log k), whereas pairwise merges inflate complexity; e.g., merging sorted SSTables in storage engines. -* For time-ordered processing, an event *priority queue* schedules the next timestamped event correctly, whereas FIFO or LIFO queues break ordering; e.g., discrete-event network simulators. -* In heuristic search like A*, a *priority queue* ordered by f=g+h explores promising nodes sooner, whereas BFS/DFS explore many irrelevant nodes; e.g., game map pathfinding. -* To initialize from bulk data, using *heapify* builds a heap in O(n), whereas inserting one item at a time costs O(n log n); e.g., preparing a backlog for scheduling. -* When you must sort in place with bounded memory, *heapsort* provides O(n log n) time and O(1) extra space, whereas mergesort needs additional buffers; e.g., firmware sorting tables. -* If priorities change dynamically, supporting *decrease-key* (or increase-key) avoids duplicates and stale entries, whereas delete-and-reinsert inflates size and time; e.g., reprioritizing tickets as SLAs age. -* Recognize heap-friendly patterns when the loop is “extract best, then insert new,” as a *greedy workflow* matches heap operations well, whereas balanced trees maintain unused total order; e.g., job schedulers. -* Avoid heaps when the task is fast membership checks or deletes by value, because a *hash table* handles contains/remove efficiently, whereas heaps need extra indexing; e.g., deduplicating queued emails. -* When priorities are small integers, a *bucket queue* (array of buckets) gives near O(1) operations, whereas a binary heap remains O(log n); e.g., Dijkstra with weights 0–100. - -#### Working with Trees and Binary Trees - -Trees show up everywhere because they model hierarchy, and many puzzles quietly hide a tree even if they don’t call it one (ranges, prefixes, decisions, ancestors). The “do” is to lean on traversal patterns and invariants (BST order, heap property, balance). The “don’t” is assuming every tree is balanced, shape matters, and it changes complexity. +- When you frequently need the next smallest or largest item, a priority queue offers O(1) peek and O(log n) updates, whereas rescanning a list is O(n) per step; e.g., choosing the earliest-deadline task. +- For practical implementations, an array-backed binary heap is cache-friendly with O(n) build and O(log n) push/pop, whereas pointer trees add overhead; e.g., worker queues in backend services. +- If the requirement is to repeatedly extract the minimum, a min-heap returns it quickly as new items arrive, whereas resorting after each insertion wastes time; e.g., streaming tasks ordered by due time. +- When tracking the current maximum continuously, a max-heap maintains the top element efficiently, whereas maintaining a full sorted list is unnecessary; e.g., live high-bid tracking in auctions. +- For top-k results on a stream, a fixed-size heap of k items yields O(log k) updates and bounded memory, whereas storing everything and sorting is slower and larger; e.g., top 100 error types today. +- To keep a running median, two heaps (max-heap lower half, min-heap upper) give O(log n) updates and O(1) median queries, whereas recomputing from all values is O(n); e.g., live latency dashboards. +- In shortest-path search, an indexed priority queue supports decrease-key in O(log n), whereas a plain heap needs lazy deletes that add overhead; e.g., Dijkstra on road networks. +- During k-way merging, a min-heap holding one head per list outputs the next item in O(log k), while balanced pairwise merging also achieves $O(N\log k)$ total time for $N$ elements; e.g., merging sorted SSTables in storage engines. +- For time-ordered processing, an event priority queue schedules the next timestamped event correctly, whereas FIFO or LIFO queues break ordering; e.g., discrete-event network simulators. +- In heuristic search like A*, a priority queue ordered by f=g+h explores promising nodes sooner, whereas BFS/DFS explore many irrelevant nodes; e.g., game map pathfinding. +- To initialize from bulk data, using heapify builds a heap in O(n), whereas inserting one item at a time costs O(n log n); e.g., preparing a backlog for scheduling. +- When you must sort in place with bounded memory, heapsort provides O(n log n) time and O(1) extra space, whereas mergesort needs additional buffers; e.g., firmware sorting tables. +- If priorities change dynamically, supporting decrease-key (or increase-key) avoids duplicates and stale entries, whereas delete-and-reinsert inflates size and time; e.g., reprioritizing tickets as SLAs age. +- Recognize heap-friendly patterns when the loop is “extract best, then insert new,” as a greedy workflow matches heap operations well, whereas balanced trees maintain unused total order; e.g., job schedulers. +- Avoid heaps when the task is fast membership checks or deletes by value, because a hash table handles contains/remove efficiently, whereas heaps need extra indexing; e.g., deduplicating queued emails. +- Bounded integer priorities can make bucket queues useful. Inserting into a known bucket is constant time, but locating the next nonempty bucket adds scanning cost. For Dijkstra with integer edge weights in `0..C`, a conventional bucket implementation has an $O(E+VC)$ bound; the value of $C$ matters. + +### Working with Trees and Binary Trees + +Trees show up everywhere because they model hierarchy, and many puzzles quietly hide a tree even if they don’t call it one (ranges, prefixes, decisions, ancestors). Lean on traversal patterns and invariants (BST order, heap property, balance). The “don’t” is assuming every tree is balanced, shape matters, and it changes complexity. Trees are hierarchical data structures with a root node and child nodes. Binary trees are a specific type where each node has at most two children. -* To visit every node systematically, applying *tree traversals* ensures predictable orders (in-, pre-, post-order) and stack-friendly iterations, whereas skipping them leads to ad-hoc scans that miss nodes; e.g., in a BST, in-order prints sorted IDs. -* Searching large key sets is more helpful with a *binary search tree* because ordered left/right children give O(log n) searches in balanced cases, while omitting this structure can degrade to linear scans; for instance, finding a user ID among thousands. -* To keep operations predictable, maintaining a *balanced BST* uses rotations to preserve logarithmic height, whereas ignoring balance can turn inserts and lookups into near-linear time; e.g., repeatedly adding sorted timestamps. -* For always retrieving the highest-priority item first, a *heap* gives O(1) peek and O(log n) updates, but without it you may rescan a list each time; think of scheduling the next job on a build server. -* Storing priorities in an array-backed *binary heap* enables O(n) build-heap and in-place heapsort, while not using it can increase memory use and sort time; e.g., turning a backlog into a sorted release order. -* When processing a tree by layers, using *level-order traversal* with a queue visits nodes breadth-first, whereas skipping it complicates tasks like shortest unweighted paths; for example, printing org-chart levels. -* Measuring *tree height* clarifies depth-related performance and balance, while omitting it hides skew that slows operations; e.g., monitoring API routing trees for growth after deployments. -* For pairwise relationships, computing the *lowest common ancestor* answers ancestor queries quickly via parent pointers or preprocessing, whereas not doing so forces repeated climbs; e.g., finding a shared manager in an org chart. -* Speeding prefix queries becomes easier with a *trie* that stores characters along edges for O(length) lookups, while alternatives may scan many strings; e.g., autocomplete for product codes. -* Handling many range updates and queries benefits from a *segment tree* that delivers O(log n) operations, whereas naive recomputation per query is slower; for example, maintaining rolling sales sums per day range. -* When you only need prefix sums, a *Fenwick tree* offers compact O(log n) updates and queries, while skipping it can force O(n) recomputation; e.g., live leaderboards that update scores. -* Representing hierarchical data with an *N-ary tree* models multiple children per node so traversals remain organized, whereas flattening loses parent–child context; e.g., folder structures with many subfolders. -* Viewing a structure as a *tree (graph)* highlights it is a connected acyclic graph with n−1 edges, while ignoring this view can block reuse of helpful DFS/BFS tools; e.g., validating a dependency tree has no cycles. -* To store or transmit structures, applying *serialization* via preorder with nulls or level-order arrays preserves shape, while omitting it risks reconstruction errors; e.g., saving a quiz logic tree between sessions. -* On disk-backed indexes, choosing a *B-Tree* (or B+ Tree) keeps nodes wide to reduce I/O so searches stay near logarithmic, whereas simple binary trees cause many random reads; e.g., database primary-key lookups. - -#### Working with Graphs +- To visit every node systematically, applying tree traversals ensures predictable orders (in-, pre-, post-order) and stack-friendly iterations, whereas skipping them leads to ad-hoc scans that miss nodes; e.g., in a BST, in-order prints sorted IDs. +- Searching large key sets is more helpful with a binary search tree because ordered left/right children give O(log n) searches in balanced cases, while omitting this structure can degrade to linear scans; for instance, finding a user ID among thousands. +- To keep operations predictable, maintaining a balanced BST uses rotations to preserve logarithmic height, whereas ignoring balance can turn inserts and lookups into near-linear time; e.g., repeatedly adding sorted timestamps. +- For always retrieving the highest-priority item first, a heap gives O(1) peek and O(log n) updates, but without it you may rescan a list each time; think of scheduling the next job on a build server. +- Storing priorities in an array-backed binary heap enables O(n) build-heap and in-place heapsort, while not using it can increase memory use and sort time; e.g., turning a backlog into a sorted release order. +- When processing a tree by layers, using level-order traversal with a queue visits nodes breadth-first, whereas skipping it complicates tasks like shortest unweighted paths; for example, printing org-chart levels. +- Measuring tree height clarifies depth-related performance and balance, while omitting it hides skew that slows operations; e.g., monitoring API routing trees for growth after deployments. +- For pairwise relationships, computing the lowest common ancestor answers ancestor queries quickly via parent pointers or preprocessing, whereas not doing so forces repeated climbs; e.g., finding a shared manager in an org chart. +- Speeding prefix queries becomes easier with a trie that stores characters along edges for O(length) lookups, while alternatives may scan many strings; e.g., autocomplete for product codes. +- Handling many range updates and queries benefits from a segment tree that delivers O(log n) operations, whereas naive recomputation per query is slower; for example, maintaining rolling sales sums per day range. +- When you only need prefix sums, a Fenwick tree offers compact O(log n) updates and queries, while skipping it can force O(n) recomputation; e.g., live leaderboards that update scores. +- Representing hierarchical data with an N-ary tree models multiple children per node so traversals remain organized, whereas flattening loses parent–child context; e.g., folder structures with many subfolders. +- Viewing a structure as a tree (graph) highlights it is a connected acyclic graph with n−1 edges, while ignoring this view can block reuse of helpful DFS/BFS tools; e.g., validating a dependency tree has no cycles. +- To store or transmit structures, applying serialization via preorder with nulls or level-order arrays preserves shape, while omitting it risks reconstruction errors; e.g., saving a quiz logic tree between sessions. +- On disk-backed indexes, choosing a B-Tree (or B+ Tree) keeps nodes wide to reduce I/O so searches stay near logarithmic, whereas simple binary trees cause many random reads; e.g., database primary-key lookups. + +### Working with Graphs Graphs consist of vertices (nodes) and edges connecting them, used to represent complex relationships. -Graph puzzles feel intimidating until you realize most of them are built from a small set of moves: represent the graph well, traverse it correctly, and keep the right bookkeeping (visited sets, distances, parents). The “do” is to translate the story into edges and nodes early. The “don’t” is to hand-wave graph direction or weights, those details decide the algorithm. +Graph puzzles feel intimidating until you realize most of them are built from a small set of moves: represent the graph well, traverse it correctly, and keep the right bookkeeping (visited sets, distances, parents). Translate the story into edges and nodes early. Do not hand-wave graph direction or weights, those details decide the algorithm. -I. *Graph Representations*: +I. Graph Representations: -* On sparse graphs, using an *adjacency list* reduces memory and speeds neighbor iteration, whereas choosing a dense structure wastes space and slows scans; for example, a road map where towns link to few highways. -* In dense networks, choosing an *adjacency matrix* enables O(1) edge checks, whereas relying on lists requires linear searches per check; for example, a tightly meshed data-center fabric. +- On sparse graphs, using an adjacency list reduces memory and speeds neighbor iteration, whereas choosing a dense structure wastes space and slows scans; for example, a road map where towns link to few highways. +- In dense networks, choosing an adjacency matrix enables O(1) edge checks, whereas relying on lists requires linear searches per check; for example, a tightly meshed data-center fabric. -II. *Graph Traversal Algorithms*: +II. Graph Traversal Algorithms: -* When exploring deep structures, applying *depth-first search* follows long paths with backtracking, whereas unguided exploration may revisit nodes and miss structure; for example, solving a maze with dead ends. -* For fewest-hop queries, running *breadth-first search* visits nodes level by level to return shortest unweighted paths, whereas depth-first order can yield longer routes; for example, finding minimal connections between users in a social network. +- When exploring deep structures, applying depth-first search follows long paths with backtracking, whereas unguided exploration may revisit nodes and miss structure; for example, solving a maze with dead ends. +- For fewest-hop queries, running breadth-first search visits nodes level by level to return shortest unweighted paths, whereas depth-first order can yield longer routes; for example, finding minimal connections between users in a social network. -III. *Cycle Detection*: +III. Cycle Detection: -* In directed graphs, implementing *cycle detection* via DFS with a recursion stack flags back edges, whereas skipping the stack conflates exploration with cycles; for example, validating package import orders. -* On undirected graphs, performing *cycle detection* by checking revisits to a non-parent neighbor reveals loops, whereas ignoring the parent test raises false alarms; for example, auditing redundant links in a LAN. +- In directed graphs, implementing cycle detection via DFS with a recursion stack flags back edges, whereas skipping the stack conflates exploration with cycles; for example, validating package import orders. +- On undirected graphs, performing cycle detection by checking revisits to a non-parent neighbor reveals loops, whereas ignoring the parent test raises false alarms; for example, auditing redundant links in a LAN. -IV. *Shortest Path Algorithms*: +IV. Shortest Path Algorithms: -* With non-negative weights, using *Dijkstra’s algorithm* and a priority queue yields shortest paths efficiently, whereas arbitrary relaxation can be slower or incorrect; for example, computing fastest driving times. -* When negative edges are possible, selecting *Bellman-Ford* both finds shortest paths and reports negative cycles, whereas Dijkstra may return wrong distances; for example, detecting currency-exchange arbitrage. -* For all-pairs needs on moderately dense graphs, applying *Floyd–Warshall* precomputes every distance in O(n³), whereas repeating many single-source runs adds overhead; for example, planning shipping costs between warehouses and stores. +- With non-negative weights, using Dijkstra’s algorithm and a priority queue yields shortest paths efficiently, whereas arbitrary relaxation can be slower or incorrect; for example, computing fastest driving times. +- When negative edges are possible, selecting Bellman-Ford both finds shortest paths and reports negative cycles, whereas Dijkstra may return wrong distances; for example, detecting currency-exchange arbitrage. +- For all-pairs needs on moderately dense graphs, applying Floyd–Warshall precomputes every distance in O(n³), whereas repeating many single-source runs adds overhead; for example, planning shipping costs between warehouses and stores. -V. *Minimum Spanning Trees (MST)*: +V. Minimum Spanning Trees (MST): -* To grow a low-cost backbone, running *Prim’s algorithm* adds the cheapest edge to the current tree, whereas unguided choices can inflate total weight; for example, wiring campus buildings with minimal cable. -* When edges are easy to sort, employing *Kruskal’s algorithm* accepts lightest edges while Union-Find blocks cycles, whereas naive addition can create loops; for example, designing a minimal road network among towns. +- To grow a low-cost backbone, running Prim’s algorithm adds the cheapest edge to the current tree, whereas unguided choices can inflate total weight; for example, wiring campus buildings with minimal cable. +- When edges are easy to sort, employing Kruskal’s algorithm accepts lightest edges while Union-Find blocks cycles, whereas naive addition can create loops; for example, designing a minimal road network among towns. -VI. *Network Flow Algorithms*: +VI. Network Flow Algorithms: -* In capacity networks, iterating augmenting paths with the *Ford-Fulkerson method* increases flow until none remain, whereas treating edges independently leaves throughput unused; for example, pushing more water through a pipe system. -* To control search effort, adopting *Edmonds–Karp* chooses shortest augmenting paths by BFS, whereas arbitrary choices may require many iterations; for example, allocating bandwidth across shared links. +- In networks with non-negative integer capacities, the Ford–Fulkerson method increases flow along residual augmenting paths until none remain. Reverse residual edges allow earlier choices to be revised. Arbitrary augmenting-path choices do not guarantee termination for irrational capacities; whereas treating edges independently leaves throughput unused; for example, pushing more water through a pipe system. +- To control search effort, adopting Edmonds–Karp chooses shortest augmenting paths by BFS, whereas arbitrary choices may require many iterations; for example, allocating bandwidth across shared links. -VII. *Other Important Concepts*: +VII. Other Important Concepts: -* For acyclic dependencies, performing *topological sorting* produces a valid task order, whereas ignoring it risks starting work before prerequisites; for example, scheduling courses with prerequisites. -* On directed graphs, computing *strongly connected components* via Kosaraju or Tarjan groups mutually reachable nodes, whereas skipping this step hides tightly coupled modules; for example, isolating service clusters in a microservice graph. -* For conflict separation, checking *bipartite graphs* via BFS/DFS enables 2-coloring and reveals odd cycles, whereas ad-hoc coloring may waste time and colors; for example, assigning exam slots so no shared-course students collide. +- For acyclic dependencies, performing topological sorting produces a valid task order, whereas ignoring it risks starting work before prerequisites; for example, scheduling courses with prerequisites. +- On directed graphs, computing strongly connected components via Kosaraju or Tarjan groups mutually reachable nodes, whereas skipping this step hides tightly coupled modules; for example, isolating service clusters in a microservice graph. +- For conflict separation, checking bipartite graphs via BFS/DFS enables 2-coloring and reveals odd cycles, whereas ad-hoc coloring may waste time and colors; for example, assigning exam slots so no shared-course students collide. -#### Working with Hash Tables +### Working with Hash Tables Hash tables store key-value pairs for efficient lookup, insertion, and deletion. -Hash tables are the go-to “make it fast” tool because they turn searching into (usually) constant time. In puzzles, they often appear when you need to remember what you’ve seen: duplicates, frequencies, complements, visited states. The “do” is to leverage them for counting and membership. The “don’t” is using mutable keys or assuming worst-case can’t happen. +Hash tables are the go-to “make it fast” tool because they turn searching into (usually) constant time. In puzzles, they often appear when you need to remember what you’ve seen: duplicates, frequencies, complements, visited states. Leverage them for counting and membership. The “don’t” is using mutable keys or assuming worst-case can’t happen. -* When *choosing a good hash function*, ensure it distributes keys uniformly across the table to reduce the likelihood of collisions. For custom objects, it's important to override hash functions carefully, ensuring they are consistent and align with the object’s equality logic. -* To deal with *collisions*, two main methods are commonly used: *Chaining*, where multiple elements that hash to the same index are stored in a linked list or another structure at that index. *Open addressing*, which resolves collisions by finding another empty slot through probing methods, such as linear probing, quadratic probing, or double hashing. -* The *load factor and resizing* are important aspects of hash table management. The load factor, which is the ratio of the number of elements to the number of buckets, should be monitored. When the load factor exceeds a certain threshold (often around 0.7), it's time to resize and rehash the table to maintain performance. -* Always use *immutable keys* in a hash table to avoid complications from key mutation. If a key changes after insertion, it can corrupt the structure of the hash table and lead to retrieval failures. -* Hash tables have numerous *applications*, including frequency counting, caching, indexing, and implementing associative arrays like dictionaries or maps in various programming languages. -* While hash tables offer *$O(1)$ average-case time complexity*, it’s important to understand that the worst-case time complexity can degrade to $O(n)$ if the hash function leads to many collisions, particularly when the load factor is too high. +- When choosing a good hash function, ensure it distributes keys uniformly across the table to reduce the likelihood of collisions. For custom objects, it's important to override hash functions carefully, ensuring they are consistent and align with the object’s equality logic. +- To deal with collisions, two main methods are commonly used: Chaining, where multiple elements that hash to the same index are stored in a linked list or another structure at that index. Open addressing, which resolves collisions by finding another empty slot through probing methods, such as linear probing, quadratic probing, or double hashing. +- The load factor and resizing are important aspects of hash table management. The load factor, which is the ratio of the number of elements to the number of buckets, should be monitored. When the load factor exceeds a certain threshold (often around 0.7), it's time to resize and rehash the table to maintain performance. +- Always use immutable keys in a hash table to avoid complications from key mutation. If a key changes after insertion, it can corrupt the structure of the hash table and lead to retrieval failures. +- Hash tables have numerous applications, including frequency counting, caching, indexing, and implementing associative arrays like dictionaries or maps in various programming languages. +- While hash tables offer *$O(1)$ average-case time complexity*, it’s important to understand that the worst-case time complexity can degrade to $O(n)$ if the hash function leads to many collisions, particularly when the load factor is too high. -### Algorithms +## Algorithms Mastering algorithms is helpful for solving programming problems more efficiently by understanding patterns and techniques that reduce time and space usage. -Algorithms are the “moves” you apply once you’ve chosen your data structures. A simple way to build skill is to recognize the story the puzzle is telling: *pair finding* (two pointers), *explore choices* (backtracking), *best so far* (greedy), *reuse work* (DP), *split and combine* (divide and conquer). The “do” is to name the pattern out loud, once you name it, you can reach for proven templates. +Algorithms are the “moves” you apply once you’ve chosen your data structures. A simple way to build skill is to recognize the story the puzzle is telling: pair finding (two pointers), explore choices (backtracking), best so far (greedy), reuse work (DP), split and combine (divide and conquer). Name the pattern out loud, once you name it, you can reach for proven templates. -#### Two-Pointer Technique +### Two-Pointer Technique Use two pointers when: -* The data is **sorted** (or can be sorted), and you need **pairs/triplets** that satisfy a condition (e.g., sum to a target). -* The condition is **monotonic** as pointers move (moving one side predictably increases/decreases the value you’re comparing). +- The data is sorted (or can be sorted), and you need pairs/triplets that satisfy a condition (e.g., sum to a target). +- The condition is monotonic as pointers move (moving one side predictably increases/decreases the value you’re comparing). Typical uses: 2-Sum (sorted), 3-Sum (outer loop + two pointers), merging intervals, palindrome checks, removing duplicates, sliding-window constraints. Two pointers are powerful because they replace nested loops with a single controlled scan. The “why” is geometry: when the array is sorted, moving left or right has a predictable effect, so you can steer toward the answer instead of brute forcing combinations. -**How it works** +How it works: -* Before searching for pairs in a sorted array, set the *pointers* `left` to index 0 and `right` to index n−1, which is useful because without this setup you may miss edge combinations or scan redundantly (e.g., starting at the ends of prices [1,2,9,11] for budget 13 quickly reveals 2+11). -* During the scan, continue while the *condition* `left < right` so that you stop before indices cross, whereas omitting this check can reuse the same element or cause an out-of-bounds access (e.g., halting when `left == right` prevents pairing one ticket price with itself). -* When the current sum equals the *target*, record the pair and move both pointers inward, but if the sum is too small move `left` rightward and if too large move `right` leftward, and skipping these directional moves would prolong the search or miss valid pairs (e.g., item prices [1,3,4,7,10] with budget 11 yield 1+10 then 3+7). -* After recording a valid pair, advance *indices* past duplicate values on both sides to avoid repeats, whereas failing to skip duplicates will emit the same pair multiple times (e.g., with prices [1,1,2,3,4,4] for 5 you keep {1,4} and {2,3} once each). -* In terms of performance, a single pass runs in linear *time* O(n) with O(1) extra space after an O(n log n) sort, whereas omitting sorting or pointer discipline often devolves into quadratic checks with nested loops (e.g., ~10,000 versus ~50,000,000 comparisons on 10,000 items). +- Before searching for pairs in a sorted array, set the pointers `left` to index 0 and `right` to index n−1, which is useful because without this setup you may miss edge combinations or scan redundantly (e.g., starting at the ends of prices [1,2,9,11] for budget 13 quickly reveals 2+11). +- During the scan, continue while the condition `left < right` so that you stop before indices cross, whereas omitting this check can reuse the same element or cause an out-of-bounds access (e.g., halting when `left == right` prevents pairing one ticket price with itself). +- When the current sum equals the target, record the pair and move both pointers inward, but if the sum is too small move `left` rightward and if too large move `right` leftward, and skipping these directional moves would prolong the search or miss valid pairs (e.g., item prices [1,3,4,7,10] with budget 11 yield 1+10 then 4+7). +- After recording a valid pair, advance indices past duplicate values on both sides to avoid repeats, whereas failing to skip duplicates will emit the same pair multiple times (e.g., with prices [1,1,2,3,4,4] for 5 you keep {1,4} and {2,3} once each). +- In terms of performance, a single pass runs in linear time O(n) with O(1) extra space after an O(n log n) sort, whereas omitting sorting or pointer discipline often devolves into quadratic checks with nested loops (e.g., ~10,000 versus ~50,000,000 comparisons on 10,000 items). -**Walkthrough example , all pairs summing to 10** +For pair enumeration, decide whether the output means distinct value pairs or every pair of indices. Skipping equal values gives distinct value pairs. Listing every index pair can require quadratic output when many values repeat. + +Walkthrough example: distinct value pairs summing to 10 Array (sorted): `[1, 2, 3, 4, 5, 6, 7, 8, 9]` @@ -248,28 +250,28 @@ Stop when `left >= right`. Two quick counter-examples for moves: -* **Too small:** `1 + 8 = 9 < 10` → move `left` rightward. -* **Too large:** `2 + 9 = 11 > 10` → move `right` leftward. +- Too small: `1 + 8 = 9 < 10` → move `left` rightward. +- Too large: `2 + 9 = 11 > 10` → move `right` leftward. -**Variants you’ll use often** +Variants you’ll use often -* When finding zero-sum triplets, apply the *3-sum* pattern by sorting and for each index scanning `i+1…n−1` with two pointers while skipping duplicates of `i` and inner pointers, which is useful because omitting sorting or duplicate skipping causes missed or repeated triplets (e.g., from `[-4,-1,-1,0,1,2]` you report `[-1,-1,2]` and `[-1,0,1]` once). -* If you only need to know whether any pair reaches the target, treat the task as an *existence* check and stop at the first hit, whereas continuing the scan enumerates all pairs and, without duplicate skipping, repeats results (e.g., with target `10` in `[1,3,7,7,9]` you may stop at `3+7` or keep scanning to list it once despite the second `7`). -* For range constraints over contiguous elements, use a *sliding window* by expanding `right` to grow and contracting `left` to restore the constraint, whereas moving pointers only in one direction can overshoot and miss beneficial ranges (e.g., counting subarrays with sum ≤ `8` in `[2,3,1,2,4,3]` or tracking at most `2` distinct items in a rolling shopping cart). +- When finding zero-sum triplets, apply the 3-sum pattern by sorting and for each index scanning `i+1…n−1` with two pointers while skipping duplicates of `i` and inner pointers, which is useful because omitting sorting or duplicate skipping causes missed or repeated triplets (e.g., from `[-4,-1,-1,0,1,2]` you report `[-1,-1,2]` and `[-1,0,1]` once). +- If you only need to know whether any pair reaches the target, treat the task as an existence check and stop at the first hit, whereas continuing the scan enumerates all pairs and, without duplicate skipping, repeats results (e.g., with target `10` in `[1,3,7,7,9]` you may stop at `3+7` or keep scanning to list it once despite the second `7`). +- For range constraints with suitable monotonicity, use a sliding window by expanding `right` and contracting `left` to restore the constraint. Sum-based shrinking requires non-negative elements; negative elements can invalidate the reasoning, whereas moving pointers only in one direction can overshoot and miss beneficial ranges (e.g., counting subarrays with sum ≤ `8` in `[2,3,1,2,4,3]` or tracking at most `2` distinct items in a rolling shopping cart). -#### Recursion +### Recursion -I. *Recursion* works by breaking a problem into smaller instances of itself, with each call reducing the size of the problem. +I. Recursion works by breaking a problem into smaller instances of itself, with each call reducing the size of the problem. II. Recursion process includes: -* Defining a clear *base case* to ensure that recursion terminates when the simplest version of the problem is solved. -* Ensuring that the *recursive case* moves toward the base case by reducing the problem’s size or complexity with each call. -* Verifying the *correctness* of your recursion logic to ensure it mirrors the structure of the problem. +- Defining a clear base case to ensure that recursion terminates when the simplest version of the problem is solved. +- Ensuring that the recursive case moves toward the base case by reducing the problem’s size or complexity with each call. +- Verifying the correctness of your recursion logic to ensure it mirrors the structure of the problem. -III. Consider *stack overflow* when recursion depth is too great. You can either switch to an iterative approach or use tail recursion optimization (if supported by the language). *Memoization* is another technique to improve efficiency by caching results of recursive calls to avoid redundant computations. +III. Consider stack overflow when recursion depth is too great. You can either switch to an iterative approach or use tail recursion optimization (if supported by the language). Memoization is another technique to improve efficiency by caching results of recursive calls to avoid redundant computations. -Recursion is less about “calling yourself” and more about writing a solution that matches the shape of the problem. The “do” is to define what a smaller instance looks like and make sure each call gets you closer to it. The “don’t” is to hide progress, if the input doesn’t clearly shrink toward a base case, the recursion will either loop forever or crash. +Recursion is less about “calling yourself” and more about writing a solution that matches the shape of the problem. Define what a smaller instance looks like and make sure each call gets you closer to it. Do not hide progress, if the input doesn’t clearly shrink toward a base case, the recursion will either loop forever or crash. ``` factorial(4) @@ -285,26 +287,26 @@ factorial(4) |---> Base case: 1 ``` -#### Backtracking +### Backtracking -I. *Backtracking* is well-suited for problems that require exploring all potential configurations, such as puzzles like Sudoku, the N-Queens problem, or combinatorial tasks. +I. Backtracking is well-suited for problems that require exploring all potential configurations, such as puzzles like Sudoku, the N-Queens problem, or combinatorial tasks. -II. *Implementation tips* for backtracking include: +II. Implementation tips for backtracking include: -* Clearly defining how to represent the *current state* of the problem. -* Using *constraints checking* before progressing to ensure that the current path meets the problem's criteria, which helps prune invalid paths early. -* Recursively *exploring* valid options by modifying the state and recursively diving deeper into the solution space. -* After exploring each path, *backtracking* involves undoing the last change and trying the next alternative, ensuring all possible solutions are considered. -* Consider *optimizations* such as using heuristics to explore more promising paths first. +- Clearly defining how to represent the current state of the problem. +- Using constraints checking before progressing to ensure that the current path meets the problem's criteria, which helps prune invalid paths early. +- Recursively exploring valid options by modifying the state and recursively diving deeper into the solution space. +- After exploring each path, backtracking involves undoing the last change and trying the next alternative, ensuring all possible solutions are considered. +- Consider optimizations such as using heuristics to explore more promising paths first. -III. A classic example is solving the *N-Queens problem*, where queens are placed row by row, and conflicts are resolved by backtracking when necessary. +III. A classic example is solving the N-Queens problem, where queens are placed row by row and invalid placements are rejected immediately. The following walkthrough uses one-based row and column numbers. -Backtracking is structured trial-and-error. The “why” it works is that you don’t blindly try everything, you quickly reject choices that violate constraints, which can cut the search space down dramatically. The “do” is to make constraint checks cheap and early. The “don’t” is to forget to undo changes, state leaks are the most common backtracking bug. +Backtracking is structured trial-and-error. The “why” it works is that you don’t blindly try everything, you quickly reject choices that violate constraints, which can cut the search space down dramatically. Make constraint checks cheap and early. Do not forget to undo changes, state leaks are the most common backtracking bug. -*Step 1: Start with an empty board.* +Step 1: Start with an empty board. -* Place a queen in Row 1, Column 1. -* Try placing queens row by row. +- Place a queen in Row 1, Column 1. +- Try placing queens row by row. ``` +---+---+---+---+ @@ -318,11 +320,11 @@ Backtracking is structured trial-and-error. The “why” it works is that you d +---+---+---+---+ ``` -*Step 2: Move to Row 2.* +Step 2: Move to Row 2. -* Try Column 1 → Conflict! (Same column as Queen in Row 1). -* Try Column 2 → Conflict! (Diagonal from Queen in Row 1). -* Try Column 3 → Place Queen. +- Try Column 1 → Conflict! (Same column as Queen in Row 1). +- Try Column 2 → Conflict! (Diagonal from Queen in Row 1). +- Try Column 3 → Place Queen. ``` +---+---+---+---+ @@ -336,42 +338,38 @@ Backtracking is structured trial-and-error. The “why” it works is that you d +---+---+---+---+ ``` -*Step 3: Move to Row 3.* +Step 3: Check Row 3. -* Try Column 1 → Conflict! (Same column as Queen in Row 1). -* Try Column 2 → Place Queen. +- Column 1 conflicts with the queen in Row 1. +- Column 2 conflicts diagonally with the queen in Row 2. +- Column 3 conflicts with the queen in Row 2 and diagonally with Row 1. +- Column 4 conflicts diagonally with the queen in Row 2. -``` +No queen is placed in Row 3 because every column is invalid. + +Step 4: Backtrack before entering Row 4. + +There is no valid partial solution to extend, so Row 4 is never reached on this branch. Rejecting the branch here is the purpose of checking constraints after each choice. + +Step 5: Undo the last placement. + +Remove the queen from Row 2, Column 3. Keep the queen in Row 1 while trying the next candidate in Row 2. The board is now: + +```text +---+---+---+---+ | Q | | | | +---+---+---+---+ -| | | Q | | +| | | | | +---+---+---+---+ -| | Q | | | +| | | | | +---+---+---+---+ | | | | | +---+---+---+---+ ``` -*Step 4: Move to Row 4.* +Step 6: Try other possibilities in Row 2. -* Try Column 1 → Conflict! (Same column as Queen in Row 1). -* Try Column 2 → Conflict! (Same column as Queen in Row 3). -* Try Column 3 → Conflict! (Diagonal from Queen in Row 2). -* Try Column 4 → Conflict! (Diagonal from Queen in Row 1). -* No valid positions! (All columns conflict). -* Backtrack: Remove the queen from Row 3, Column 2. - -*Step 5: Try other possibilities in Row 3.* - -* Try Column 3 → Conflict! (Diagonal from Queen in Row 1). -* Try Column 4 → Conflict! (Diagonal from Queen in Row 2). -* No valid positions in Row 3. -* Backtrack: Remove the queen from Row 2, Column 3. - -*Step 6: Try other possibilities in Row 2.* - -* Try Column 4 → Place Queen. +- Try Column 4 → Place Queen. ``` +---+---+---+---+ @@ -385,10 +383,10 @@ Backtracking is structured trial-and-error. The “why” it works is that you d +---+---+---+---+ ``` -*Step 7: Move to Row 3.* +Step 7: Move to Row 3. -* Try Column 1 → Conflict! (Same column as Queen in Row 1). -* Try Column 2 → Place Queen. +- Try Column 1 → Conflict! (Same column as Queen in Row 1). +- Try Column 2 → Place Queen. ``` +---+---+---+---+ @@ -402,52 +400,54 @@ Backtracking is structured trial-and-error. The “why” it works is that you d +---+---+---+---+ ``` -Continue exploring alternatives and backtracking as necessary until all solutions are found. +Row 4 has no valid column for this partial placement either, so backtrack again. Eventually the queen in Row 1 must move. The two solutions, written as one-based column choices for Rows 1 through 4, are `[2, 4, 1, 3]` and `[3, 1, 4, 2]`. -#### Dynamic Programming +### Dynamic Programming -I. *Dynamic programming* is used to solve optimization problems by breaking them down into overlapping subproblems with an optimal substructure, meaning solutions to smaller subproblems can be reused to solve larger problems. +I. Dynamic programming stores reusable subproblem results. Optimization problems rely on optimal substructure; counting and feasibility problems instead combine counts or Boolean results. In each case, a state must capture all information needed to solve its subproblem. -II. There are two primary *approaches* to dynamic programming: +II. There are two primary approaches to dynamic programming: -* The *top-down approach (memoization)* involves using recursion while caching the results of subproblems to avoid redundant computations. -* The *bottom-up approach (tabulation)* builds a solution iteratively by solving the smallest subproblems first and using their solutions to solve larger subproblems. +- The top-down approach (memoization) involves using recursion while caching the results of subproblems to avoid redundant computations. +- The bottom-up approach (tabulation) builds a solution iteratively by solving the smallest subproblems first and using their solutions to solve larger subproblems. III. Some steps that might be taken in dynamic programming: -* *Defining the subproblems* to break the main problem into manageable parts. -* *Identifying the state variables* that uniquely define each subproblem. -* Establishing a *recurrence relation* to compute the solution of each subproblem based on smaller subproblems. -* Proper *initialization* of base cases in a table or memoization cache. -* Determining the correct *iteration order* to fill the DP table for bottom-up approaches. +- Defining the subproblems to break the main problem into manageable parts. +- Identifying the state variables that uniquely define each subproblem. +- Establishing a recurrence relation to compute the solution of each subproblem based on smaller subproblems. +- Proper initialization of base cases in a table or memoization cache. +- Determining the correct iteration order to fill the DP table for bottom-up approaches. -IV. *Space optimization* is often possible by realizing that only a few recent states are needed, reducing space complexity. +IV. Space optimization is often possible by realizing that only a few recent states are needed, reducing space complexity. V. Classic examples include problems like the Fibonacci sequence, the Knapsack problem, or calculating the minimum edit distance between two strings. -Dynamic programming is what you reach for when brute force repeats itself. The “why” is efficiency: if the same subproblem appears again and again, you should only solve it once. The “do” is to define a state you can memoize and a recurrence you trust. The “don’t” is to build a giant table without a clear meaning for what each cell represents. +Dynamic programming is what you reach for when brute force repeats itself. The “why” is efficiency: if the same subproblem appears again and again, you should only solve it once. Define a state you can memoize and a recurrence you trust. Do not build a giant table without a clear meaning for what each cell represents. ``` -DP Table (Rows = Items, Columns = Knapsack Capacities): +0/1 Knapsack: items (weight, value) = (1,1), (3,4), (4,5), (4,7) +Each item is available once; capacity is 7. +Rows = first i available items, columns = capacity. Capacity -> 0 1 2 3 4 5 6 7 Item 0 [ 0 0 0 0 0 0 0 0 ] (Base case: No items) Item 1 [ 0 1 1 1 1 1 1 1 ] (Include Item 1) Item 2 [ 0 1 1 4 5 5 5 5 ] (Include Item 2) Item 3 [ 0 1 1 4 5 6 6 9 ] (Include Item 3) - Item 4 [ 0 1 1 4 5 7 8 11 ] (Include Item 4) + Item 4 [ 0 1 1 4 7 8 8 11 ] (First 4 items available) ``` -*How the Table Is Filled* +How the Table Is Filled For each cell $DP[i][w]$ if the weight of the item $i$ is less than or equal to the current capacity $w$, choose the maximum of: -* Value without including the item ($DP[i-1][w]$). -* Value including the item ($DP[i-1][w-\text{weight}[i]] + \text{value}[i]$). +- Value without including the item ($DP[i-1][w]$). +- Value including the item ($DP[i-1][w-\text{weight}[i]] + \text{value}[i]$). -*Visualization of Choices* +Visualization of Choices -I. For *Item 1* (Weight = 1, Value = 1): +I. For Item 1 (Weight = 1, Value = 1): ``` Capacity = 0 → Can't include Item 1 → DP[1][0] = 0 @@ -457,7 +457,7 @@ Capacity = 2 → Include Item 1 → DP[1][2] = 1 Capacity = 7 → Include Item 1 → DP[1][7] = 1 ``` -II. For *Item 2* (Weight = 3, Value = 4): +II. For Item 2 (Weight = 3, Value = 4): ``` Capacity = 0 → Can't include Item 2 → DP[2][0] = 0 @@ -469,46 +469,47 @@ Capacity = 7 → Include Item 2 → DP[2][7] = 5 III. Continue for all items, progressively updating the table. -*Extract the Solution* +Extract the Solution To find the maximum value look at the last cell: $DP[4][7] = 11$. To find the items included trace back from $DP[4][7]$, checking where values changed: -* $DP[4][7] → Include Item 4$ -* $DP[3][3] → Include Item 3$ +- $DP[4][7]=11$ exceeds $DP[3][7]=9$, so include Item 4 and reduce capacity from 7 to 3. +- $DP[3][3]=DP[2][3]=4$, so skip Item 3. +- $DP[2][3]=4$ exceeds $DP[1][3]=1$, so include Item 2; the remaining capacity is zero. -*Final Knapsack Contents:* +Final Knapsack Contents: -* Item 3 (Weight 4, Value 5) -* Item 4 (Weight 3, Value 7) -* Total Weight = 7, Total Value = 11 +- Item 2 (Weight 3, Value 4) +- Item 4 (Weight 4, Value 7) +- Total Weight = 7, Total Value = 11 ### Greedy Algorithms -I. *Greedy algorithms* are used when making a locally optimal choice at each step leads to a globally optimal solution. +I. Greedy algorithms are used when making a locally optimal choice at each step leads to a globally optimal solution. -II. The two *characteristics* of greedy algorithms are: +II. The two characteristics of greedy algorithms are: -* *Optimal substructure*, meaning the overall solution incorporates optimal solutions to subproblems. -* The *greedy choice property*, where making the best local decision at each step results in the globally best solution. +- Optimal substructure, meaning the overall solution incorporates optimal solutions to subproblems. +- The greedy choice property, where making the best local decision at each step results in the globally best solution. -III. Common *implementation tips* for greedy algorithms include: +III. Common implementation tips for greedy algorithms include: -* *Sorting* the input data according to a specific criterion before applying the greedy strategy. -* Always ensure the greedy choice leads to an optimal solution by providing a *proof of correctness* or counterexamples. +- Sorting the input data according to a specific criterion before applying the greedy strategy. +- Prove that the greedy choice is safe, for example through an exchange argument. A counterexample disproves a rule; failing to find one does not prove correctness. IV. Examples of greedy algorithms include the activity selection problem, Huffman coding, and algorithms for finding minimum spanning trees (Prim's and Kruskal's). -Greedy algorithms are tempting because they feel simple: pick the best-looking option and move on. Sometimes that works beautifully, and sometimes it fails spectacularly. The “do” is to either know the problem is greedy-friendly (via proof or known pattern) or actively search for a counterexample. The “don’t” is assuming that “best right now” must lead to “best overall.” +Greedy algorithms are tempting because they feel simple: pick the best-looking option and move on. Sometimes that works beautifully, and sometimes it fails spectacularly. Either know the problem is greedy-friendly (via proof or known pattern) or actively search for a counterexample. The “don’t” is assuming that “best right now” must lead to “best overall.” -*Example Huffman coding Input*: +Example Huffman coding Input: Characters: $[A, B, C, D, E, F]$ Frequencies: $[5, 9, 12, 13, 16, 45]$ -*Build a Min-Heap* +Build a Min-Heap Create a priority queue (min-heap) with the characters and their frequencies. @@ -517,7 +518,7 @@ Initial Min-Heap: 5(A) 9(B) 12(C) 13(D) 16(E) 45(F) ``` -*Build the Huffman Tree* +Build the Huffman Tree Combine the two smallest frequency nodes into a new node. Repeat until there is one tree. @@ -563,124 +564,124 @@ Updated Heap: $[45(F), 55(CDABE)]$ V. Combine 45(F) and 55(CDABE): -``` -Tree: - (100) - / \ - 45(F) 55(CDABE) - / \ - 25(CD) 30(ABE) - / \ / \ - 12(C) 13(D) 14(AB) 16(E) - / \ - 5(A) 9(B) +```text + (100) + / \ + F:45 (55) + / \ + (25) (30) + / \ / \ + C:12 D:13 (14) E:16 + / \ + A:5 B:9 ``` -*Assign Binary Codes* +Assign Binary Codes Traverse the tree to assign codes: -* Left edge = `0` -* Right edge = `1` +- Left edge = `0` +- Right edge = `1` ``` Codes: -A = 000 -B = 001 +A = 1100 +B = 1101 C = 100 D = 101 -E = 01 -F = 11 +E = 111 +F = 0 ``` -*Final Huffman Tree Diagram:* +Final Huffman Tree Diagram: +```text + (100) + 0 / \ 1 + F (55) + 0 / \ 1 + (25) (30) + 0 / \1 0/ \1 + C D (14) E + 0/ \1 + A B ``` -Tree: - (100) - / \ - (45) (55) - / / \ - (F) (25) (30) - / \ / \ - (C) (D) (14) (E) - / \ - (A) (B) -``` -#### Divide and Conquer +The weighted length is `5×4 + 9×4 + 12×3 + 13×3 + 16×3 + 45×1 = 224` bits for 100 symbols, or 2.24 bits per symbol. It also equals the sum of merge weights: `14 + 25 + 30 + 55 + 100`. + +### Divide and Conquer -I. The *divide and conquer* strategy solves problems by dividing them into smaller subproblems, solving those independently, and then combining their solutions. +I. The divide and conquer strategy solves problems by dividing them into smaller subproblems, solving those independently, and then combining their solutions. -II. *Implementation tips* for divide and conquer: +II. Implementation tips for divide and conquer: -* Use *recursion* to divide the problem, with each recursive call handling a subproblem. -* Pay attention to the *combine step*, as the efficiency of combining subproblem solutions can affect the overall performance. -* Define *base cases* for small subproblems that can be solved directly without further division. +- Use recursion to divide the problem, with each recursive call handling a subproblem. +- Pay attention to the combine step, as the efficiency of combining subproblem solutions can affect the overall performance. +- Define base cases for small subproblems that can be solved directly without further division. III. Examples of divide and conquer algorithms include Merge Sort, Quick Sort, and Binary Search. -Divide and conquer is your “zoom lens.” Instead of trying to solve the whole problem at once, you solve smaller independent versions and merge results. The “do” is to keep subproblems truly independent and make the combine step efficient. The “don’t” is accidentally recomputing the same work across branches, if that happens, you may be drifting into dynamic programming territory. +Divide and conquer is your “zoom lens.” Instead of trying to solve the whole problem at once, you solve smaller independent versions and merge results. Keep subproblems truly independent and make the combine step efficient. The “don’t” is accidentally recomputing the same work across branches, if that happens, you may be drifting into dynamic programming territory. -### Sorting Algorithms +## Sorting Algorithms Sorting algorithms are fundamental to computer science and programming. They are used to rearrange elements in a list or array so that they follow a specific order (ascending or descending). Efficient sorting is a first step for optimizing other algorithms (like search and merge algorithms) that require input data to be in sorted lists. Understanding the different sorting algorithms, their time and space complexities, stability, and suitable use cases is essential for problem-solving and technical interviews. -#### Overview of Common Sorting Algorithms +### Overview of Common Sorting Algorithms -Below is a detailed comparison of commonly used sorting algorithms: +The table uses auxiliary space, including recursive stack storage. Here $n$ is the number of items, $k$ is a key range or bucket count, and radix sort uses $d$ digit positions with radix $b$. Average bounds depend on input assumptions; stability can depend on implementation. -| *Algorithm* | *Average Time Complexity* | *Worst-Case Time Complexity* | *Space Complexity* | *Stability* | *Best Use Case* | *Notes* | +| Algorithm | Average Time Complexity | Worst-Case Time Complexity | Space Complexity | Stability | Best Use Case | Notes | |---------------------|-----------------------------|-------------------------------|----------------------|---------------|----------------------------------------------|---------------------------------------------| -| *Bubble Sort* | $O(n^2)$ | $O(n^2)$ | $O(1)$ | Stable | Educational purposes, small datasets | Simple but inefficient for large datasets | -| *Insertion Sort* | $O(n^2)$ | $O(n^2)$ | $O(1)$ | Stable | Nearly sorted or small datasets | Efficient for small or nearly sorted datasets | -| *Selection Sort* | $O(n^2)$ | $O(n^2)$ | $O(1)$ | Unstable | Small datasets, when memory is limited | Inefficient for large datasets | -| *Merge Sort* | $O(n \log n)$ | $O(n \log n)$ | $O(n)$ | Stable | Large datasets, linked lists | Requires additional memory for merging | -| *Quick Sort* | $O(n \log n)$ | $O(n^2)$ | $O(\log n)$ | Unstable | Large datasets, general-purpose sorting | Pivot selection strategy affects performance | -| *Heap Sort* | $O(n \log n)$ | $O(n \log n)$ | $O(1)$ | Unstable | Large datasets, in-place sorting | Efficient with minimal memory usage | -| *Radix Sort* | $O(nk)$ | $O(nk)$ | $O(n + k)$ | Stable | Large datasets with integer keys | Non-comparative sorting algorithm | -| *Counting Sort* | $O(n + k)$ | $O(n + k)$ | $O(k)$ | Stable | Small range of integer keys | Efficient when range $k$ is small | -| *Tim Sort* | $O(n \log n)$ | $O(n \log n)$ | $O(n)$ | Stable | Real-world data, hybrid sorting | Default sorting algorithm in Python | -| *Bucket Sort* | $O(n + k)$ | $O(n^2)$ | $O(n)$ | Stable | Uniformly distributed data | Divides elements into buckets | -| *Shell Sort* | $O(n \log n)$ to $O(n^2)$ | $O(n^2)$ | $O(1)$ | Unstable | Medium-sized datasets | Improves upon Insertion Sort | - -#### General Tips for Sorting Algorithms - -- It’s important to *understand the data* when selecting a sorting algorithm. Factors such as the size of the dataset, its distribution, and the data type being sorted can significantly influence the choice of algorithm. -- If the *stability requirement* is important, meaning the relative order of equal elements must be preserved, you should opt for stable sorting algorithms like Merge Sort or Tim Sort. -- When memory is a concern, *in-place sorting* algorithms, such as Quick Sort or Heap Sort, are preferable because they require minimal additional memory. -- Consider the *time complexity trade-offs* when choosing a sorting algorithm. While Quick Sort is generally fast, its performance can degrade to $O(n^2)$ in the worst case if not implemented carefully. -- *Hybrid approaches* like Tim Sort, which combines Merge Sort and Insertion Sort, are designed to optimize performance by leveraging the strengths of multiple sorting techniques. - -#### Practical Applications - -- *Merge Sort* is particularly useful for sorting linked lists since it does not require random access to the elements. -- *Quick Sort* is often chosen for general-purpose sorting due to its efficiency in the average case and its ability to sort in-place. -- *Counting Sort* and *Radix Sort* are effective when sorting integers within a known and limited range, offering linear-time complexity under the right conditions. -- *Heap Sort* is a good option when memory usage is restricted and a guaranteed $O(n \log n)$ time complexity is necessary. -- *Insertion Sort* is ideal for very small datasets or as a subroutine in more complex algorithms due to its simplicity and efficiency on nearly sorted data. - -### Bit Manipulation +| Bubble Sort | $O(n^2)$ | $O(n^2)$ | $O(1)$ | Stable | Educational purposes, small datasets | Simple but inefficient for large datasets | +| Insertion Sort | $O(n^2)$ | $O(n^2)$ | $O(1)$ | Stable | Nearly sorted or small datasets | Efficient for small or nearly sorted datasets | +| Selection Sort | $O(n^2)$ | $O(n^2)$ | $O(1)$ | Unstable | Small datasets, when memory is limited | Inefficient for large datasets | +| Merge Sort | $O(n \log n)$ | $O(n \log n)$ | $O(n)$ | Stable | Large datasets, linked lists | Requires additional memory for merging | +| Quick Sort | $O(n \log n)$ | $O(n^2)$ | $O(n)$ worst; $O(\log n)$ with smaller-side recursion | Unstable | Large datasets, general-purpose sorting | Pivot selection strategy affects performance | +| Heap Sort | $O(n \log n)$ | $O(n \log n)$ | $O(1)$ | Unstable | Large datasets, in-place sorting | Efficient with minimal memory usage | +| LSD Radix Sort | $O(d(n+b))$ | $O(d(n+b))$ | $O(n+b)$ | Stable with stable digit passes | Large datasets with integer keys | Non-comparative sorting algorithm | +| Counting Sort | $O(n + k)$ | $O(n + k)$ | $O(n + k)$ | Stable | Small range of integer keys | Efficient when range $k$ is small | +| Tim Sort | $O(n \log n)$ | $O(n \log n)$ | $O(n)$ | Stable | Real-world data, hybrid sorting | Default sorting algorithm in Python | +| Bucket Sort | $O(n+k)$ under suitable distribution | $O(n^2)$ with insertion-sorted buckets | $O(n+k)$ | Implementation dependent | Uniformly distributed data | Divides elements into buckets | +| Shell Sort | Depends on gap sequence | Depends on gap sequence; $O(n^2)$ for halving gaps | $O(1)$ | Unstable | Medium-sized datasets | Improves upon Insertion Sort | + +### General Tips for Sorting Algorithms + +- It’s important to understand the data when selecting a sorting algorithm. Factors such as the size of the dataset, its distribution, and the data type being sorted can significantly influence the choice of algorithm. +- If the stability requirement is important, meaning the relative order of equal elements must be preserved, you should opt for stable sorting algorithms like Merge Sort or Tim Sort. +- When memory is a concern, in-place sorting algorithms, such as Quick Sort or Heap Sort, are preferable because they require minimal additional memory. +- Consider the time complexity trade-offs when choosing a sorting algorithm. While Quick Sort is generally fast, its performance can degrade to $O(n^2)$ in the worst case if not implemented carefully. +- Hybrid approaches like Tim Sort, which combines Merge Sort and Insertion Sort, are designed to optimize performance by leveraging the strengths of multiple sorting techniques. + +### Practical Applications + +- Merge Sort is particularly useful for sorting linked lists since it does not require random access to the elements. +- Quick Sort is often chosen for general-purpose sorting due to its efficiency in the average case and its ability to sort in-place. +- Counting Sort and Radix Sort are effective when sorting integers within a known and limited range, offering linear-time complexity under the right conditions. +- Heap Sort is a good option when memory usage is restricted and a guaranteed $O(n \log n)$ time complexity is necessary. +- Insertion Sort is ideal for very small datasets or as a subroutine in more complex algorithms due to its simplicity and efficiency on nearly sorted data. + +## Bit Manipulation Bit manipulation involves algorithms that operate directly on bits, the basic units of data in computing. By leveraging bit-level operations, you can achieve performance optimizations, reduce memory usage, and solve certain problems more elegantly. Bit manipulation is particularly useful in systems programming, cryptography, graphics, and competitive programming. -#### Fundamental Concepts +### Fundamental Concepts -- A solid grasp of *binary representation* is necessary for working with bitwise operations. In binary, each digit (or bit) represents a power of 2, starting with the least significant bit (LSB) on the right, which corresponds to $2^0$, and increasing as you move left. -- *Signed and unsigned integers* differ in how they represent numbers. *Unsigned integers* can only represent non-negative values, while *signed integers* use the most significant bit (MSB) as a sign bit, with 0 representing positive numbers and 1 representing negative numbers, typically using two's complement representation. -- *Bitwise operators* are tools for manipulating individual bits. The *AND (`&`)* operator produces 1 only when both corresponding bits are 1, making it useful for masking bits. The *OR (`|`)* operator sets a bit to 1 if at least one of the corresponding bits is 1, often used for setting bits. -- The *XOR (`^`)* operator produces 1 when the bits are different, useful for toggling bits or swapping values. The *NOT (`~`)* operator flips all bits, performing a bitwise negation. -- The *left shift (`<<`)* operation shifts bits to the left, filling with zeros from the right, effectively multiplying the number by powers of two. Conversely, *right shift (`>>`)* operations shift bits to the right, with two variations: *logical shifts*, which fill with zeros from the left (used for unsigned integers), and *arithmetic shifts*, which preserve the sign bit, used for signed integers. +- A solid grasp of binary representation is necessary for working with bitwise operations. In binary, each digit (or bit) represents a power of 2, starting with the least significant bit (LSB) on the right, which corresponds to $2^0$, and increasing as you move left. +- Signed and unsigned integers differ in how they represent numbers. Unsigned integers can only represent non-negative values, while signed integers use the most significant bit (MSB) as a sign bit, with 0 representing positive numbers and 1 representing negative numbers, typically using two's complement representation. +- Bitwise operators are tools for manipulating individual bits. The *AND (`&`) operator produces 1 only when both corresponding bits are 1, making it useful for masking bits. The OR (`|`)* operator sets a bit to 1 if at least one of the corresponding bits is 1, often used for setting bits. +- The *XOR (`^`) operator produces 1 when the bits are different, useful for toggling bits or swapping values. The NOT (`~`)* operator flips all bits, performing a bitwise negation. +- The *left shift (`<<`) operation shifts bits to the left, filling with zeros from the right, effectively multiplying the number by powers of two. Conversely, right shift (`>>`) operations shift bits to the right, with two variations: logical shifts*, which fill with zeros from the left (used for unsigned integers), and arithmetic shifts, which preserve the sign bit, used for signed integers. -#### Bit Manipulation Techniques +### Bit Manipulation Techniques -- To *set a bit* at position $n$ to 1, the operation `number |= (1 << n)` can be used. This works by left-shifting 1 by $n$ positions to create a mask with only the $n$-th bit set, and then applying bitwise OR to modify the original number. -- To *clear a bit* at position $n$, use the operation `number &= ~(1 << n)`. Here, 1 is left-shifted by $n$ and negated to form a mask where only the $n$-th bit is 0, and applying bitwise AND clears that bit. -- To *toggle a bit* at position $n$, the operation `number ^= (1 << n)` is used. By left-shifting 1 by $n$ and applying XOR, the target bit at $n$ is flipped. -- To *check if a bit* at position $n$ is set to 1, the operation `(number & (1 << n)) != 0` is employed. It works by left-shifting 1 by $n$ and applying bitwise AND; a non-zero result indicates the bit is set. -- To *clear the least significant bit (LSB)*, the operation `number &= (number - 1)` is effective. Subtracting 1 flips all bits from the LSB onward, and applying AND clears the lowest set bit. -- To *isolate the least significant bit*, the operation `isolated_bit = number & (-number)` is used. In two’s complement, `-number` is the bitwise complement plus one, so ANDing it with `number` isolates the LSB. -- To *count the set bits* (Hamming weight) in a number, Kernighan's Algorithm is applied using the following code: +- To set a bit at position $n$ to 1, the operation `number |= (1 << n)` can be used. This works by left-shifting 1 by $n$ positions to create a mask with only the $n$-th bit set, and then applying bitwise OR to modify the original number. +- To clear a bit at position $n$, use the operation `number &= ~(1 << n)`. Here, 1 is left-shifted by $n$ and negated to form a mask where only the $n$-th bit is 0, and applying bitwise AND clears that bit. +- To toggle a bit at position $n$, the operation `number ^= (1 << n)` is used. By left-shifting 1 by $n$ and applying XOR, the target bit at $n$ is flipped. +- To check if a bit at position $n$ is set to 1, the operation `(number & (1 << n)) != 0` is employed. It works by left-shifting 1 by $n$ and applying bitwise AND; a non-zero result indicates the bit is set. +- To clear the lowest set bit, the operation `number &= (number - 1)` is effective. Subtracting 1 flips all bits from the LSB onward, and applying AND clears the lowest set bit. +- To isolate the lowest set bit, the operation `isolated_bit = number & (-number)` is used. In two’s complement, `-number` is the bitwise complement plus one, so ANDing it with `number` isolates the LSB. +- To count the set bits (Hamming weight) in a number, Kernighan's Algorithm is applied using the following code: ```c int count = 0; @@ -690,10 +691,10 @@ while (number) { } ``` -This works by repeatedly clearing the LSB and incrementing the count. +This repeatedly clears the lowest set bit and counts how many such bits were present. The displayed C loop assumes an unsigned integer. -- To *check if a number is a power of two*, the condition `(number != 0) && ((number & (number - 1)) == 0)` is used. This works because powers of two have only one set bit, and subtracting 1 flips all bits after that set bit, resulting in zero when ANDed. -- To *swap two variables* without using a temporary variable, the following sequence is used: +- To check if a number is a power of two, the condition `(number > 0) && ((number & (number - 1)) == 0)` is used. This works because powers of two have only one set bit, and subtracting 1 flips all bits after that set bit, resulting in zero when ANDed. +- To swap two variables without using a temporary variable, the following sequence is used: ```c a ^= b; @@ -701,136 +702,138 @@ b ^= a; a ^= b; ``` -This sequence of XOR operations swaps the values of `a` and `b` without needing extra space. +This sequence swaps the values only when the two names refer to distinct storage locations. If they alias, the first operation zeroes the value. A normal swap is clearer and does not require this special-case reasoning. -- To *reverse the bits* in a number, bitwise operations and shifting can be used to swap bits from opposite ends, moving towards the center of the number. This is commonly achieved using a loop or bit manipulation techniques to systematically reverse bit positions. +- To reverse the bits in a number, bitwise operations and shifting can be used to swap bits from opposite ends, moving towards the center of the number. This is commonly achieved using a loop or bit manipulation techniques to systematically reverse bit positions. -#### Bit Masks +### Bit Masks -- A *bitmask* is a binary pattern used to select or manipulate specific bits within a byte or word, enabling efficient operations on particular bits of a number. -- *Common uses of bitmasks* include applications such as feature flags, where each bit represents a boolean option or setting, and permission sets in file systems, where bits represent different permissions like read, write, or execute. -- *State representation* is another common use, where multiple boolean states are stored in a single integer, allowing efficient space usage and manipulation of these states. -- To *create a bitmask* that selects specific bits, you can combine shifts and bitwise OR. For example, to create a mask for bits 0 and 3, the operation `mask = (1 << 0) | (1 << 3)` will set only those bits to 1. +- A bitmask is a binary pattern used to select or manipulate specific bits within a byte or word, enabling efficient operations on particular bits of a number. +- Common uses of bitmasks include applications such as feature flags, where each bit represents a boolean option or setting, and permission sets in file systems, where bits represent different permissions like read, write, or execute. +- State representation is another common use, where multiple boolean states are stored in a single integer, allowing efficient space usage and manipulation of these states. +- To create a bitmask that selects specific bits, you can combine shifts and bitwise OR. For example, to create a mask for bits 0 and 3, the operation `mask = (1 << 0) | (1 << 3)` will set only those bits to 1. -#### Bit Shifting Tricks +### Bit Shifting Tricks -- *Left shifting* (`<<`) is a useful technique for *multiplying by powers of two*. For instance, `number << 3` multiplies `number` by $2^3 = 8$, which is a fast alternative to regular multiplication. -- Similarly, *right shifting* (`>>`) can be used for *dividing by powers of two*. For example, `number >> 2` divides `number` by $2^2 = 4$, making it an efficient way to handle division for unsigned integers or logical shifts. -- To *extract specific bits* from a number, you can use a combination of shifting and masking. For instance, to extract bits from position $p$ to $p + n - 1$, the operation `(number >> p) & ((1 << n) - 1)` can be applied. This shifts the target bits to the right and uses a mask to isolate only those bits. +- Left shifting (`<<`) is a useful technique for multiplying by powers of two. For instance, `number << 3` multiplies `number` by $2^3 = 8$, provided the value is representable and the language's shift rules permit it. Compilers already optimize multiplication by constants; shifting is not automatically faster. +- Similarly, right shifting (`>>`) can be used for dividing by powers of two. For example, `number >> 2` divides `number` by $2^2 = 4$, making it an efficient way to handle division for unsigned integers or logical shifts. +- To extract specific bits from a number, you can use a combination of shifting and masking. For instance, to extract bits from position $p$ to $p + n - 1$, the operation `(number >> p) & ((1 << n) - 1)` can be applied. This shifts the target bits to the right and uses a mask to isolate only those bits. -#### Cautions and Best Practices +### Cautions and Best Practices -- In programming, it is important to note that *bitwise operators* have lower precedence than arithmetic and relational operators, so parentheses should be used to ensure the correct order of operations. -- When working with shifts, be aware that *arithmetic shifts* (using `>>`) preserve the sign bit in signed integers, which is useful when working with negative numbers. -- Some languages, like Java, also support *logical shifts*, represented as `>>>`, which shift zeros into the high-order bits regardless of the sign of the number. -- *Portability* can be a concern in bit manipulation code because differences in integer sizes and endianness across various systems may affect how the code behaves. -- Be careful to avoid *overflow and underflow* when performing shifts, as shifting bits beyond the size of the data type can lead to unintended results. -- While bit manipulation can be powerful, excessive use can hurt *readability*. To make the code easier to maintain, it’s a good practice to include comments and, where appropriate, use macros or inline functions. +- In programming, it is important to note that bitwise operators have lower precedence than arithmetic and relational operators, so parentheses should be used to ensure the correct order of operations. +- When working with shifts, be aware that arithmetic shifts (using `>>`) preserve the sign bit in signed integers, which is useful when working with negative numbers. +- Some languages, like Java, also support logical shifts, represented as `>>>`, which shift zeros into the high-order bits regardless of the sign of the number. +- Integer width and signedness affect masks and shifts. Endianness affects how bytes are stored or serialized, not the mathematical result of a bitwise operation on an integer value. +- Be careful to avoid overflow and underflow when performing shifts, as shifting bits beyond the size of the data type can lead to unintended results. +- While bit manipulation can be powerful, excessive use can hurt readability. To make the code easier to maintain, it’s a good practice to include comments and, where appropriate, use macros or inline functions. -### List of problems +## List of Problems -#### Minimum deletions to make valid parentheses +### Minimum deletions to make valid parentheses Given a string of parentheses, determine the minimum number of parentheses that need to be removed to make the string valid. This can be solved using a stack data structure to track the open parentheses as we iterate through the string. If we encounter an open parenthesis, we add it to the stack. If we encounter a closing parenthesis, we check if there is a matching open parenthesis on the top of the stack. If there is, we pop the open parenthesis from the stack. If there is not, we add the closing parenthesis to a list of characters to remove. After we have processed the entire string, we remove the remaining open parentheses from the stack. -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/cpp/deletions_to_make_valid_parentheses) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/python/deletions_to_make_valid_parentheses) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/cpp/deletions_to_make_valid_parentheses) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/python/deletions_to_make_valid_parentheses) + +### Is palindrome after at most one char delete? -#### Is palindrome after at most one char delete? +Determine whether a string becomes a palindrome after deleting at most one character. Move two pointers inward while the characters match. At the first mismatch, test the two remaining possibilities: skip the left character or skip the right character, then require the remaining range to be a palindrome without further deletions. Deleting both would exceed the limit. The two checks still give $O(n)$ time and $O(1)$ auxiliary space when implemented by indices. -Determinine whether a string can be transformed into a palindrome by deleting at most one character. This can be solved using a two pointer approach, where we start from both ends of the string and move towards the middle. If we encounter a pair of characters that are not equal, we have three options: delete the character at the left pointer, delete the character at the right pointer, or delete both characters. We check if any of these options results in a palindrome and return the result. +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/cpp/is_palindrome_after_char_deletion) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/python/is_palindrome_after_char_deletion) -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/cpp/is_palindrome_after_char_deletion) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/python/is_palindrome_after_char_deletion) +### K closest points to origin -#### K closest points to origin +Find the K points closest to the origin. Maintain a max-heap of at most K points, ordered by squared distance: its root is the farthest retained point. Replace that root when a closer point arrives. This takes $O(n\log K)$ time and $O(K)$ space for $1\le K\le n$; handle $K=0$ separately. Alternatively, build a min-heap of all points and extract K times in $O(n+K\log n)$ time with $O(n)$ storage. Squared distance preserves ordering without computing square roots. -Find the K points in a list of points that are closest to the origin (0, 0). This can be solved using a min heap data structure, where we store the K closest points so far. As we iterate through the list of points, we compute the distance from each point to the origin and add it to the heap if it is among the K closest points so far. +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/cpp/k_closest_points) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/python/k_closest_points) -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/cpp/k_closest_points) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/python/k_closest_points) +### Subarray sum equals K -#### Subarray sum equals K +To find a contiguous subarray summing to K when negative numbers are allowed, maintain prefix sums and a map of earlier sums. At prefix sum `current`, an earlier prefix equal to `current - K` identifies a matching range. Seed the map with the empty prefix sum zero. Store indices to recover one range, or frequencies to count every matching range, in expected $O(n)$ time and $O(n)$ space. A sum-based sliding window is appropriate only when the input and objective support monotonic shrinking, such as non-negative values. -Find a contiguous subarray of an array that has a sum of K. This can be solved using a sliding window approach, where we maintain a running sum of the elements in the window and move the window through the array until we find a subarray with the desired sum. +The linked exercise is a related range-sum task with supplied endpoints rather than target-sum search. Its Python description uses an exclusive end index; specify endpoint conventions before translating a formula or comparing examples. -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/cpp/subarray_sum) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/python/subarray_sum) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/cpp/subarray_sum) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/python/subarray_sum) -#### Add numbers given as strings +### Add numbers given as strings Add two numbers represented as strings. This can be solved by treating the strings as arrays of digits and using a carry variable to keep track of the carryover from one place value to the next. -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/cpp/add_string_numbers) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/python/add_string_numbers) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/cpp/add_string_numbers) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/python/add_string_numbers) -#### Dot product of two sparse vectors +### Dot product of two sparse vectors Compute the dot product of two sparse vectors, where a vector is represented as a list of (index, value) pairs. This can be solved by iterating through both vectors and adding up the products of the values at the same index. - - -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/cpp/sparse_vectors_product) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/python/sparse_vectors_product) -#### Range sum of BST + +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/cpp/sparse_vectors_product) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/python/sparse_vectors_product) + +### Range sum of BST Compute the sum of the values of all the nodes in a binary search tree within a given range. This can be solved using a recursive in-order traversal of the tree, where we add the value of each node to the sum if it is within the range. -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/cpp/range_sum_of_bst) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/python/range_sum_of_bst) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/cpp/range_sum_of_bst) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/python/range_sum_of_bst) -#### Product of array except self +### Product of array except self Compute an array where each element is the product of all the other elements in the input array. This can be solved using two pass approach, where we first compute the product of all the elements before each index and then compute the product of all the elements after each index. -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/cpp/array_product) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/python/array_product) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/cpp/array_product) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/python/array_product) -#### Convert BST to sorted doubly linked list +### Convert BST to sorted doubly linked list Convert a binary search tree to a sorted doubly linked list. This can be solved using a recursive in-order traversal of the tree, where we build the linked list by adding each node to the end of the list as we visit it. -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/cpp/bst_to_list) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/python/bst_to_list) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/cpp/bst_to_list) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/python/bst_to_list) -#### Lowest common ancestor of a binary tree +### Lowest common ancestor of a binary tree Find the lowest common ancestor (LCA) of two nodes in a binary tree. The LCA is the node in the tree that is the ancestor of both nodes and is the deepest node in the tree. To solve this problem, you can use a variety of techniques such as traversing the tree in a depth-first or breadth-first manner, or using a recursive approach to traverse the tree and find the LCA. You can also use a divide and conquer approach, where you split the tree into left and right subtrees and find the LCA in each subtree. Another approach is to use a hash table or a map to store the ancestors of each node and then use this information to find the LCA. -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/cpp/lowest_common_ancestor) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/python/lowest_common_ancestor) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/cpp/lowest_common_ancestor) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/python/lowest_common_ancestor) -#### LRU Cache +### LRU Cache -Designing a cache data structure that stores a limited number of items and removes the least recently used items when the capacity is reached. This can be solved using a doubly linked list and a hash table. The doubly linked list is used to store the items in the cache in the order in which they were accessed, with the most recently accessed item at the front of the list and the least recently accessed item at the back. The hash table is used to store the keys and values of the items in the cache and to quickly retrieve items from the cache based on their keys. +Designing a cache data structure that stores a limited number of items and removes the least recently used items when the capacity is reached. This can be solved using a doubly linked list and a hash table. The doubly linked list is used to store the items in the cache in the order in which they were accessed, with the most recently accessed item at the front of the list and the least recently accessed item at the back. The hash table maps keys directly to linked-list nodes, allowing lookup, movement to the front, and eviction from the back in expected $O(1)$ time. Storing only values would still require a list search to update recency. -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/cpp/lru_cache) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/python/lru_cache) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/cpp/lru_cache) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/python/lru_cache) -#### Randomize An Array +### Randomize An Array -Shuffle the elements of an array randomly. This can be solved using a random number generator and a Fisher-Yates shuffle algorithm. The Fisher-Yates shuffle algorithm works by starting at the end of the array and randomly swapping each element with an element preceding it. This results in a randomly shuffled array. +Shuffle the elements of an array randomly. This can be solved using a random number generator and a Fisher-Yates shuffle algorithm. The Fisher-Yates shuffle algorithm works by starting at the end of the array and swapping the item at index `i` with an item chosen uniformly from indices `0..i`, including itself. This results in a randomly shuffled array. -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/cpp/randomize_array) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/python/randomize_array) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/cpp/randomize_array) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/python/randomize_array) -#### Binary Tree Right Side View +### Binary Tree Right Side View -Given a binary tree, return an array containing the values of the nodes on the right side of the tree, when viewed from the right. One way to solve this problem might be to perform a breadth-first search on the tree, keeping track of the maximum depth of each node as it is visited. When visiting a node at a particular depth for the first time, add its value to the result array. +Given a binary tree, return an array containing the values of the nodes on the right side of the tree, when viewed from the right. Use breadth-first traversal and record the last node of each level when visiting children left to right. Alternatively, traverse right-first and record the first node encountered at each depth. The visit order determines which side is visible. -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/cpp/binary_tree_right_side_view) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/python/binary_tree_right_side_view) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/cpp/binary_tree_right_side_view) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/python/binary_tree_right_side_view) -#### Design Browser History +### Design Browser History -Implement a browser history system that supports the following operations: visit a URL, go back to the previous URL, and go forward to the next URL. One potential approach to this problem could involve using a doubly-linked list to store the URLs visited, with a pointer to the current URL. Going back or forward would simply involve moving the pointer to the previous or next URL in the list. +Implement a browser history system that supports the following operations: visit a URL, go back to the previous URL, and go forward to the next URL. One potential approach to this problem could involve using a doubly-linked list to store the URLs visited, with a pointer to the current URL. Going back or forward moves the pointer to the previous or next entry, stopping at the ends. Visiting a new URL after moving back discards the forward history before appending the new entry. -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/cpp/design_browser_history) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/python/design_browser_history) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/cpp/design_browser_history) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/python/design_browser_history) -#### Score After Flipping Matrix +### Score After Flipping Matrix -Given a matrix of 0s and 1s, flip the rows and columns of the matrix such that the maximum possible score is achieved. The score is calculated as the number of 1s in the matrix. One way to solve this problem might be to use dynamic programming to compute the maximum score for each submatrix of the original matrix, taking into account whether the rows and columns of the submatrix should be flipped. +Each row of the binary matrix represents a binary number, with the most significant bit on the left. The score is the sum of those row values, not the total number of ones. First flip any row whose leading bit is zero: that bit is worth more than all later bits in the row combined. Then flip each remaining column if it contains more zeros than ones. Once leading bits are fixed, each column can be optimized independently. This greedy approach takes $O(RC)$ time and $O(1)$ auxiliary space if flips are performed in place. -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/cpp/score_after_flipping_matrix) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/python/score_after_flipping_matrix) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/cpp/score_after_flipping_matrix) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/brain_teasers/python/score_after_flipping_matrix) diff --git a/notes/data_structures.md b/notes/data_structures.md index 940b703..f186c80 100644 --- a/notes/data_structures.md +++ b/notes/data_structures.md @@ -1,21 +1,23 @@ -## Data Structures: Collections and Containers +# Data Structures: Collections and Containers -In computer science, a *collection* (often interchangeably referred to as a *container*) is a sophisticated data structure designed to hold multiple entities, these could be simple elements like numbers or text strings, or more complex objects like user-defined structures. Collections help you store, organize, and manipulate different types of data in your programs. +In computer science, a collection (often interchangeably referred to as a container) is a data structure that holds multiple values. These may be simple elements like numbers or text strings, or more complex objects like user-defined structures. Collections help you store, organize, and manipulate different types of data in your programs. -Collections are one of those topics that feel “obvious” until you start building real software. The moment your program needs to store more than a couple values, you’re forced to answer questions like: *How fast do I need to find things? Will I be inserting a lot? Does order matter? Do I need uniqueness?* Picking a collection is basically picking the rules of your data’s world, and that choice quietly controls your performance, your code clarity, and how painful future changes will be. +Choose a collection by listing the operations the program needs: lookup, insertion, deletion, iteration, ordering, and uniqueness. The right choice simplifies those operations and makes their performance costs predictable. -1. At **The Abstract Level**, we focus on the conceptual understanding of collections, defining what they are and their characteristics. This includes how they store and retrieve items, whether it's based on a specific order (as in arrays or lists) or using unique keys (as in dictionaries or maps). -2. **The Machine Level** refers to the practical implementation of collections, where we concentrate on efficiently realizing the abstract models. Factors like memory usage, operation speeds (insertion, deletion, search), and flexibility help determine the best data structures and algorithms for a specific task. +1. At The Abstract Level, we focus on the conceptual understanding of collections, defining what they are and their characteristics. This includes how they store and retrieve items, whether it's based on a specific order (as in arrays or lists) or using unique keys (as in dictionaries or maps). +2. The Machine Level refers to the practical implementation of collections, where we concentrate on efficiently realizing the abstract models. Factors like memory usage, operation speeds (insertion, deletion, search), and flexibility help determine the best data structures and algorithms for a specific task. -This “two-level view” is how you stop memorizing data structures and start *using* them. Abstractly, you care about behavior and guarantees (“fast lookup”, “keeps order”, “no duplicates”). At the machine level, you care about what those guarantees cost (extra memory, copying, pointer chasing, cache misses). The best developers keep both views in their head: the abstract model tells you what’s possible, and the machine model tells you what’s smart. +The abstract interface defines behavior, such as ordering and uniqueness. The implementation determines costs such as memory allocation, copying, pointer traversal, and cache misses. Consider both when comparing containers. Every developer should really get to know collections well. It helps you pick the right type of collection for the job, making your code faster, more efficient, and easier to manage in the long run. -diagram +![Overview of collection types](https://github.com/user-attachments/assets/3a2fd42a-68d2-4cdd-abed-f087d1d96ede) -### Linked Lists +The interfaces below describe conceptual operations; exact method names and return values vary by implementation. Complexity tables assume constant-time element comparisons, moves, and hashing unless otherwise stated. Here, $n$ is the number of stored elements. -A *Linked List* is a basic way to organize data where you have a chain of nodes. Each node contains two distinct parts: +## Linked Lists + +A Linked List is a basic way to organize data where you have a chain of nodes. Each node contains two distinct parts: 1. A value (or data), which could range from simple data types to more complex structures. 2. A reference (commonly referred to as a link or a pointer) to the subsequent node in the sequence. @@ -28,24 +30,24 @@ Linked lists can be highly valuable when the data doesn't require contiguous sto A quick “do/don’t” intuition that helps: -* **Do** consider a linked list when your operations naturally happen at the ends or via known node references (like “remove this node I already have”). -* **Don’t** reach for a linked list if you constantly need `list[i]` style access, because every index lookup becomes a walk. +- Do consider a linked list when your operations naturally happen at the ends or via known node references (like “remove this node I already have”). +- Don’t reach for a linked list if you constantly need `list[i]` style access, because every index lookup becomes a walk. ``` [HEAD]-->[1| ]-->[2| ]-->[3| ]-->[4| ]-->[NULL] ``` -#### Typical Interface +### Typical Interface Here are some standard operations associated with linked lists: -* `is_empty()`: This checks if the list is devoid of any elements. -* `front()`: This retrieves the first element in the list. -* `append(element)`: This adds a new element to the end of the list. -* `remove(index)`: This removes an element at a specified index. -* `replace(index, data)`: This substitutes the element at a specified index with a new value. +- `is_empty()`: This checks if the list is devoid of any elements. +- `front()`: This retrieves the first element in the list. +- `append(element)`: This adds a new element to the end of the list. +- `remove(index)`: This removes an element at a specified index. +- `replace(index, data)`: This substitutes the element at a specified index with a new value. -#### Time Complexity +### Time Complexity | Operation | Average case | Worst case | | --------- | ------------ | ---------- | @@ -54,34 +56,34 @@ Here are some standard operations associated with linked lists: | Delete | `O(1)` | `O(1)` | | Search | `O(n)` | `O(n)` | -Note that while insertion and deletion are constant time operations, access and search operations require traversing the list, resulting in linear time complexity. +The insertion and deletion rows assume the required links are already known. In a singly linked list, deleting an arbitrary node generally requires its predecessor; having only the node reference is insufficient. Appending is $O(1)$ with a tail pointer and $O(n)$ without one. Index-based removal and replacement require traversal. -One detail that’s easy to miss: linked-list **insert/delete** being `O(1)` is only true when you’re inserting/deleting at a spot you can reach in `O(1)`, like the head, or a node reference you already hold. If you’re deleting “the element at index 500,” you still have to *walk* to index 500 first, which is `O(n)`. So the real takeaway is: linked lists are great at **local edits**, not great at **random access**. +One detail that’s easy to miss: linked-list insert/delete being `O(1)` is only true when you’re inserting/deleting at a spot you can reach in `O(1)`, like the head, or a node reference you already hold. If you’re deleting “the element at index 500,” you still have to walk to index 500 first, which is `O(n)`. So the real takeaway is: linked lists are great at local edits, not great at random access. -#### Common Applications +### Common Applications Linked lists find utility in a wide variety of areas including, but not limited to: -* Storing items that undergo frequent addition or removal operations. This is due to the cost-effective manipulation (insertion/deletion) capabilities of linked lists. -* Implementing other high-level data structures such as stacks and queues. -* Storing buckets in hash tables or elements in associative arrays. -* Maintaining key-value pairs in symbol tables or dictionaries. -* Representing nodes in advanced structures like graphs or tree data structures. -* Managing allocation or deallocation of memory blocks in memory management systems. -* Storing sequences of operations to support undo or redo functionality in software applications. +- Storing items that undergo frequent addition or removal operations. This is due to the cost-effective manipulation (insertion/deletion) capabilities of linked lists. +- Implementing other high-level data structures such as stacks and queues. +- Storing buckets in hash tables or elements in associative arrays. +- Maintaining key-value pairs in symbol tables or dictionaries. +- Representing nodes in advanced structures like graphs or tree data structures. +- Managing allocation or deallocation of memory blocks in memory management systems. +- Storing sequences of operations to support undo or redo functionality in software applications. -#### Implementation Details +### Implementation Details Linked lists can be implemented in various ways, largely dependent on the programming language used. A conventional implementation strategy involves defining a custom class or structure to represent a node. This class encapsulates a value and a reference to the next node. Additionally, methods for adding, removing, and accessing elements are provided to manipulate the list. In some languages, linked lists are provided as part of the standard library, abstracting away the underlying implementation details and providing a rich set of functionality. -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/collections_and_containers/cpp/linked_list) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/collections_and_containers/python/linked_list) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/collections_and_containers/cpp/linked_list) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/collections_and_containers/python/linked_list) -### Vectors +## Vectors -In computer science, a *vector* can be thought of as a dynamic array with the ability to adjust its size as required. While similar to traditional arrays, vectors distinguish themselves with their superior flexibility and efficiency in numerous scenarios. They store elements in contiguous blocks of memory, ensuring swift access to elements through index-based referencing. +In computer science, a vector can be thought of as a dynamic array with the ability to adjust its size as required. While similar to traditional arrays, vectors distinguish themselves with their superior flexibility and efficiency in numerous scenarios. They store elements in contiguous blocks of memory, ensuring swift access to elements through index-based referencing. Vectors are often the “default best friend” in many programs because they play nicely with the hardware. Contiguous memory means the CPU cache can help you, iteration is fast, and indexing is instant. You get the simplicity of an array without the pain of manually resizing it. @@ -93,53 +95,54 @@ Vectors are often favored over low-level arrays due to their augmented feature s ------------------------- ``` -#### Typical Interface +### Typical Interface Some of the commonly used operations for vectors are: -* `sort(start, end)`: This function sorts the elements in the specified range within the vector. -* `reverse(start, end)`: This operation reverses the order of elements between specified positions. -* `size()`: This function retrieves the current number of elements in the vector. -* `capacity()`: This retrieves the maximum number of elements the vector can hold before needing to resize. -* `resize(new_size)`: This operation alters the vector size, adding or removing elements as necessary. -* `reserve(new_capacity)`: This increases the vector capacity, if required, without changing the vector size. +- `sort(start, end)`: This function sorts the elements in the specified range within the vector. +- `reverse(start, end)`: This operation reverses the order of elements between specified positions. +- `size()`: This function retrieves the current number of elements in the vector. +- `capacity()`: This retrieves the maximum number of elements the vector can hold before needing to resize. +- `resize(new_size)`: This operation alters the vector size, adding or removing elements as necessary. +- `reserve(new_capacity)`: This increases the vector capacity, if required, without changing the vector size. -#### Time Complexity +### Time Complexity -| Operation | Average case | Worst case | -| --------- | ------------ | ---------- | -| Access | `O(1)` | `O(1)` | -| Insert | `O(1)` | `O(n)` | -| Delete | `O(1)` | `O(n)` | -| Search | `O(n)` | `O(n)` | +| Operation | Cost | Worst case per operation | +| --- | --- | --- | +| Access by index | `O(1)` | `O(1)` | +| Append | Amortized `O(1)` | `O(n)` | +| Remove last | `O(1)` without shrinking | `O(1)` without shrinking | +| Insert or delete at an arbitrary position | `O(n)` | `O(n)` | +| Search unsorted values | `O(n)` | `O(n)` | -While access to elements is a constant time operation, insertions and deletions can vary depending on where these operations are performed. If at the end, they are constant time operations; however, elsewhere, these operations can result in shifting of elements, causing linear time complexity. +Appending is amortized $O(1)$ when capacity grows geometrically: although an individual reallocation copies $O(n)$ elements, a sequence of $n$ appends takes $O(n)$ total work. Amortized analysis concerns a sequence of operations and does not assume random inputs. Removing the last element is constant time unless the implementation shrinks its buffer. Inserting or deleting in the middle shifts the remaining elements. -The practical story here is: vectors are fast until they aren’t, specifically, until you force them to move lots of elements. Insert at the end is usually cheap, but inserting at the front means everything shifts. And resizing sometimes triggers a bigger hidden cost: copying the entire array into a bigger one. That’s why capacity exists, and why `reserve()` can be a performance cheat code when you already know roughly how big you’ll get. +Appending to a vector is usually cheap, while insertion at the front shifts every existing element. Reallocation may move the whole buffer. If the approximate final size is known, reserving capacity can avoid repeated reallocations without changing the current size. -#### Common Applications +### Common Applications Vectors can be leveraged for a wide array of tasks, such as: -* Storing collections of values. Vectors are adept at storing and manipulating ordered lists of homogenous data types. -* Executing matrix operations. Vectors provide a good structure for representing matrices and performing computations on them. -* Implementing various search and sorting algorithms. Due to their continuous memory allocation and indexing, vectors are ideal for executing such algorithms. -* Creating intricate data structures. Vectors can serve as foundational building blocks in the implementation of more complex data structures. +- Storing collections of values. Vectors are adept at storing and manipulating ordered lists of homogenous data types. +- Executing matrix operations. Vectors provide a good structure for representing matrices and performing computations on them. +- Implementing various search and sorting algorithms. Due to their continuous memory allocation and indexing, vectors are ideal for executing such algorithms. +- Creating intricate data structures. Vectors can serve as foundational building blocks in the implementation of more complex data structures. -#### Implementation Details +### Implementation Details Vectors are typically implemented using an array of elements accompanied by additional metadata such as the current size and total capacity. When a resize operation is necessary, a new, larger array is created, and the existing elements are copied over to the new array. As resizing can be time and resource-consuming, vectors often allocate extra memory ahead of time to minimize the frequency of resizing. In some programming languages or libraries, containers may be described as vector-like even if they are internally implemented using other data structures, such as linked lists; however, regardless of such claims, a true vector (or dynamic array) is fundamentally defined by its contiguous memory storage, which ensures efficient random access and predictable time complexity for indexing operations. While alternative implementations might present a similar interface to the programmer, they cannot fully replicate the behavior or performance characteristics of a standard vector, especially for indexing, because using non-contiguous structures like linked lists inherently changes how these operations work. -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/collections_and_containers/cpp/vector) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/collections_and_containers/python/vector) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/collections_and_containers/cpp/vector) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/collections_and_containers/python/vector) -### Stacks +## Stacks -A *stack* is a fundamental data structure that models a First-In-Last-Out (FILO) or Last-In-First-Out (LIFO) strategy. In essence, the most recently added element (the top of the stack) is the first one to be removed. This data structure is widely utilized across various programming languages to execute a range of operations, such as arithmetic computations, character printing, and more. +A stack is a fundamental data structure that models a First-In-Last-Out (FILO) or Last-In-First-Out (LIFO) strategy. In essence, the most recently added element (the top of the stack) is the first one to be removed. This data structure is widely utilized across various programming languages to execute a range of operations, such as arithmetic computations, character printing, and more. -Stacks feel almost too simple, until you realize how often your brain already uses them. “Undo” in an editor? Stack. Nested function calls? Stack. Matching parentheses? Stack. Any time the *most recent thing* needs to be handled first, a stack fits naturally and keeps your logic clean. +Stacks feel almost too simple, until you realize how often your brain already uses them. “Undo” in an editor? Stack. Nested function calls? Stack. Matching parentheses? Stack. Any time the most recent thing needs to be handled first, a stack fits naturally and keeps your logic clean. ``` [5] <- top @@ -149,16 +152,16 @@ Stacks feel almost too simple, until you realize how often your brain already us [1] <- bottom ``` -#### Typical Interface +### Typical Interface Typical operations associated with a stack include: -* `push(element)`: This method places a new element on top of the stack, effectively becoming the most recent addition. -* `pop()`: This method returns the top element of the stack and simultaneously removes it. -* `top()`: This method simply returns the top element without removing it from the stack, allowing you to inspect the most recent addition. -* `is_empty()`: This method checks if the stack is empty, a useful operation before attempting to `pop()` or `top()` to prevent errors. +- `push(element)`: This method places a new element on top of the stack, effectively becoming the most recent addition. +- `pop()`: This method returns the top element of the stack and simultaneously removes it. +- `top()`: This method simply returns the top element without removing it from the stack, allowing you to inspect the most recent addition. +- `is_empty()`: This method checks if the stack is empty, a useful operation before attempting to `pop()` or `top()` to prevent errors. -#### Time Complexity +### Time Complexity | Operation | Average case | Worst case | | ------------- | ------------ | ---------- | @@ -167,28 +170,28 @@ Typical operations associated with a stack include: | Delete (Pop) | `O(1)` | `O(1)` | | Search | `O(n)` | `O(n)` | -As can be seen, while push and pop operations are constant time, access (retrieving an element without popping) and search operations require linear time since, in the worst case, all elements must be inspected. +Peeking at the top is $O(1)$. The access row refers to reaching an arbitrary deeper element, which a stack interface does not normally expose. Push and pop are worst-case $O(1)$ for a linked stack; for a dynamic-array stack, push is amortized $O(1)$ and can take $O(n)$ when the buffer grows. A quick intuition: stacks aren’t meant for “random access.” If you find yourself wanting “the 7th element from the top,” you’re fighting the data structure. The win is that they make a certain workflow effortless: add to top, remove from top, repeat. -#### Common Applications +### Common Applications Stacks are incredibly versatile and are used in a multitude of applications, including: -* **Stack-oriented programming languages** such as Forth and PostScript rely heavily on the stack data structure to perform operations. -* In **graph algorithms**, stacks are widely used in depth-first search (DFS) algorithms, where backtracking is necessary to explore unvisited nodes. -* When **finding Eulerian cycles in graphs**, stacks help keep track of the current path, as the cycle must visit every edge exactly once. -* **Identifying strongly connected components** in graphs often involves the use of stacks, as seen in algorithms like Tarjan’s algorithm, which utilizes stacks to efficiently track components. +- Stack-oriented programming languages such as Forth and PostScript rely heavily on the stack data structure to perform operations. +- In graph algorithms, stacks are widely used in depth-first search (DFS) algorithms, where backtracking is necessary to explore unvisited nodes. +- When finding Eulerian cycles in graphs, stacks help keep track of the current path, as the cycle must visit every edge exactly once. +- Identifying strongly connected components in graphs often involves the use of stacks, as seen in algorithms like Tarjan’s algorithm, which utilizes stacks to efficiently track components. -#### Implementation Details +### Implementation Details Stacks can be implemented using various underlying data structures, but arrays and linked lists are among the most common due to their inherent characteristics which nicely complement the operations and needs of a stack. The choice between arrays and linked lists often depends on the specific requirements of the application, such as memory usage and the cost of operations. With an array, resizing can be expensive but access is faster, while with a linked list, insertion and deletion are faster but more memory is used due to the extra storage of pointers. -### Queues +## Queues -A *queue* is a fundamental data structure that stores elements in a sequence, with modifications made by adding elements to the end (enqueuing) and removing them from the front (dequeuing). The queue follows a First In First Out (FIFO) strategy, implying that the oldest elements are processed first. +A queue is a fundamental data structure that stores elements in a sequence, with modifications made by adding elements to the end (enqueuing) and removing them from the front (dequeuing). The queue follows a First In First Out (FIFO) strategy, implying that the oldest elements are processed first. Queues are what you reach for when “fairness” or “arrival order” matters. They show up everywhere you need to process tasks in the order they came in: requests, jobs, events, messages. If stacks are great for backtracking and nested logic, queues are great for pipelines and scheduling. @@ -200,27 +203,27 @@ Queues are what you reach for when “fairness” or “arrival order” matters [5] <- rear ``` -#### Typical Interface +### Typical Interface The standard operations associated with a queue are: -* `enqueue(element)`: This method adds a new element to the end of the queue. -* `dequeue()`: This method returns the front element of the queue and simultaneously removes it. -* `front()`: This method simply returns the front element without removing it from the queue, giving a peek at the next item to be dequeued. -* `is_empty()`: This method checks if the queue is empty, an important step before calling `dequeue()` or `front()` to prevent errors. +- `enqueue(element)`: This method adds a new element to the end of the queue. +- `dequeue()`: This method returns the front element of the queue and simultaneously removes it. +- `front()`: This method simply returns the front element without removing it from the queue, giving a peek at the next item to be dequeued. +- `is_empty()`: This method checks if the queue is empty, an important step before calling `dequeue()` or `front()` to prevent errors. -#### Comparing Stacks and Queues +### Comparing Stacks and Queues While stacks and queues are similar data structures, they differ in several aspects: -* The stack operates on a Last-In-First-Out (LIFO) principle, while the queue follows a **First-In-First-Out (FIFO)** principle. -* In terms of **accessibility**, stacks allow access only to the top element, while queues provide access to both the front and rear elements. -* **Implementation** of both stacks and queues can be done using arrays or linked lists, but queues can also be implemented using circular arrays, heaps, or double-ended queues (deques). -* In **usage**, stacks are ideal for reversing the order of elements, while queues are useful for maintaining the original order of elements. +- The stack operates on a Last-In-First-Out (LIFO) principle, while the queue follows a First-In-First-Out (FIFO) principle. +- In terms of accessibility, stacks allow access only to the top element, while queues provide access to both the front and rear elements. +- Implementation of both stacks and queues can be done using arrays or linked lists, and FIFO queues commonly use circular arrays or double-ended queues (deques). A heap implements a priority queue, whose removal order depends on priority rather than arrival time. +- In usage, stacks are ideal for reversing the order of elements, while queues are useful for maintaining the original order of elements. A nice way to make this stick: stacks are “latest-first,” queues are “earliest-first.” When you pick one, you’re picking the story your data will follow. -#### Time Complexity +### Time Complexity | Operation | Average case | Worst case | | ---------------- | ------------ | ---------- | @@ -229,34 +232,34 @@ A nice way to make this stick: stacks are “latest-first,” queues are “earl | Delete (Dequeue) | `O(1)` | `O(1)` | | Search | `O(n)` | `O(n)` | -Enqueue and dequeue operations are constant time, while accessing or searching for a specific element require linear time as all elements may need to be inspected. +Peeking at the front is $O(1)$; the access row refers to an arbitrary interior element. Enqueue and dequeue are $O(1)$ with a linked queue that maintains both head and tail pointers, or with a fixed-capacity circular buffer. A growing circular buffer gives amortized $O(1)$ enqueue. Removing index zero from an ordinary dynamic array shifts elements and takes $O(n)$. -#### Common Applications +### Common Applications Queues are incredibly versatile and find use in a variety of real-world scenarios, including: -* In **resource scheduling**, queues are used to manage requests for resources like servers, printers, or CPU tasks, ensuring they are handled in the order they are received. -* **Communication systems** rely on queues to handle data packets and messages in a FIFO manner, making them the ideal data structure for this purpose. -* **Simulations** of real-world systems, such as event processing or customer service, often process tasks in a FIFO order, which aligns naturally with queue usage. -* In **Breadth-First Search (BFS)** for graph algorithms, a queue is used to keep track of the vertices that need to be processed. +- In resource scheduling, queues are used to manage requests for resources like servers, printers, or CPU tasks, ensuring they are handled in the order they are received. +- Communication systems rely on queues to handle data packets and messages in a FIFO manner, making them the ideal data structure for this purpose. +- Simulations of real-world systems, such as event processing or customer service, often process tasks in a FIFO order, which aligns naturally with queue usage. +- In Breadth-First Search (BFS) for graph algorithms, a queue is used to keep track of the vertices that need to be processed. -#### Implementation Details +### Implementation Details Queues can be implemented using various underlying data structures, including arrays, linked lists, circular buffers, and dynamic arrays. The choice of implementation depends on the specific requirements of the application, such as memory usage, the cost of operations, and whether the size of the queue changes frequently. -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/collections_and_containers/cpp/queue) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/collections_and_containers/python/simple_queue) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/collections_and_containers/cpp/queue) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/collections_and_containers/python/simple_queue) -### Heaps +## Heaps -A *heap* is a specialized tree-based data structure that satisfies the heap property. There are two main types of heaps: *max-heaps* and *min-heaps*. +A heap is a specialized tree-based data structure that satisfies the heap property. There are two main types of heaps: max-heaps and min-heaps. -* In a *max-heap*, the keys of parent nodes are always greater than or equal to those of the children, with the maximum key present at the root node. -* In a *min-heap*, the keys of parent nodes are always less than or equal to those of the children, with the minimum key present at the root node. +- In a max-heap, the keys of parent nodes are always greater than or equal to those of the children, with the maximum key present at the root node. +- In a min-heap, the keys of parent nodes are always less than or equal to those of the children, with the minimum key present at the root node. -Heaps are always complete binary trees, which means they are filled at all levels, except the last, which is filled from left to right. This structure ensures that the heap remains balanced. +The binary heaps discussed here are complete binary trees, which means they are filled at all levels, except the last, which is filled from left to right. This structure ensures that the heap remains balanced. -Heaps are what you use when you don’t just want “the next item,” you want “the most important next item.” That’s the whole personality of a heap: it’s built to make the minimum or maximum easy to grab *repeatedly*, even while new things keep arriving. +Heaps are what you use when you don’t just want “the next item,” you want “the most important next item.” That’s the whole personality of a heap: it’s built to make the minimum or maximum easy to grab repeatedly, even while new things keep arriving. ``` [100] @@ -268,17 +271,17 @@ Heaps are what you use when you don’t just want “the next item,” you want └── [1] ``` -#### Typical Interface +### Typical Interface The basic operations of a heap are as follows: -* `insert(value)`: Adds a new element to the heap while ensuring that the heap property is preserved. -* `search(value)`: Searches for a specific value in the heap. Note that heaps do not allow efficient arbitrary searches. -* `remove(value)`: Deletes a specific value from the heap, then restructures the heap to maintain its properties. +- `insert(value)`: Adds a new element to the heap while ensuring that the heap property is preserved. +- `search(value)`: Searches for a specific value in the heap. Note that heaps do not allow efficient arbitrary searches. +- `remove(value)`: Deletes a specific value from the heap, then restructures the heap to maintain its properties. -In practice, heaps shine most when you remove **min/max**, not arbitrary values. That’s why many heap APIs focus on `push` / `pop` (or `insert` / `extract_min`). Arbitrary `remove(value)` can exist, but it’s usually not the main event and often needs extra bookkeeping to be efficient. +In practice, heaps shine most when you remove min/max, not arbitrary values. That’s why many heap APIs focus on `push` / `pop` (or `insert` / `extract_min`). Arbitrary `remove(value)` can exist, but it’s usually not the main event and often needs extra bookkeeping to be efficient. -#### Time Complexity +### Time Complexity | Operation | Average case | | -------------- | --------------- | @@ -287,60 +290,62 @@ In practice, heaps shine most when you remove **min/max**, not arbitrary values. | Insert | `O(log n)` | | Merge | `O(m log(m+n))` | +The table describes a binary heap. Root lookup, insertion, and root removal have these worst-case heap-operation bounds, excluding buffer reallocation. Merging by inserting all $m$ elements into an $n$-element heap takes $O(m\log(m+n))$; concatenating the arrays and applying bottom-up heap construction takes $O(m+n)$. Arbitrary-value search takes $O(n)$, while removing a known index takes $O(\log n)$ after restoring heap order. + These time complexities make heaps especially useful in situations where we need to repeatedly remove the minimum (or maximum) element. -#### Common Applications +### Common Applications Heaps are versatile and find widespread use across various computational problems: -* In **graph algorithms**, heaps are used in algorithms like Dijkstra's for finding shortest paths and Prim's for constructing minimum spanning trees. -* **Priority queues** are efficiently implemented using heaps, as they provide fast insertion and extraction of the highest or lowest priority elements. -* In **sorting**, the heap sort algorithm utilizes heap properties to sort an array in-place, achieving a worst-case time complexity of **$O(n \log n)$**. +- In graph algorithms, heaps are used in algorithms like Dijkstra's for finding shortest paths and Prim's for constructing minimum spanning trees. +- Priority queues are efficiently implemented using heaps, as they provide fast insertion and extraction of the highest or lowest priority elements. +- In sorting, the heap sort algorithm utilizes heap properties to sort an array in-place, achieving a worst-case time complexity of $O(n \log n)$. -#### Implementation Details +### Implementation Details Heaps can be represented efficiently using a one-dimensional array, which allows for easy calculation of the parent, left, and right child positions. The standard mapping from the heap to this array representation is as follows: -* In a *heap*, the root of the tree is stored at index `1` of the array representation. -* The *left child* of any node at index `i` is positioned at index `2*i` in the array. -* The *right child* of a node located at index `i` can be found at index `2*i + 1`. -* To find the *parent* of any node at index `i`, calculate the index as `floor(i/2)`. +- In a heap, the root of the tree is stored at index `1` of the array representation. +- The left child of any node at index `i` is positioned at index `2*i` in the array. +- The right child of a node located at index `i` can be found at index `2*i + 1`. +- To find the parent of any node at index `i`, calculate the index as `floor(i/2)`. This layout keeps the parent-child relationships consistent and makes navigation within the heap straightforward and efficient. -One small real-world note: many implementations store the root at index `0` instead of `1`. The formulas shift slightly, but the core idea stays the same: the array layout is what makes heaps fast and compact. +One small real-world note: many implementations store the root at index `0` instead of `1`. The children of index `i` are then `2*i + 1` and `2*i + 2`, and the parent of a non-root node is `floor((i-1)/2)`. The core idea stays the same: the array layout is what makes heaps fast and compact. -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/collections_and_containers/cpp/heap) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/collections_and_containers/python/heap) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/collections_and_containers/cpp/heap) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/collections_and_containers/python/heap) -### Binary Search Trees (BST) +## Binary Search Trees (BST) -A Binary Search Tree (BST) is a type of binary tree where each node has up to two children and maintains a specific ordering property among its values. For every node, the values of all nodes in its left subtree are less than its own value, and the values of all nodes in its right subtree are greater than its own value. In a well-balanced BST, this property enables efficient lookup, addition, and deletion operations, ideally distributing the nodes evenly across the left and right subtrees. Note that BSTs typically disallow duplicate values. +A Binary Search Tree (BST) is a type of binary tree where each node has up to two children and maintains a specific ordering property among its values. For every node, the values of all nodes in its left subtree are less than its own value, and the values of all nodes in its right subtree are greater than its own value. In a well-balanced BST, this property enables efficient lookup, addition, and deletion operations, ideally distributing the nodes evenly across the left and right subtrees. This presentation assumes distinct keys. A BST can support duplicates by storing a count or a collection of values per key, provided insertion, lookup, and deletion use the same policy. -BSTs are the “ordered dictionary” idea in tree form. You get structure (everything left is smaller, everything right is bigger), which means you can search by repeatedly eliminating half the remaining options, *if* the tree stays balanced. That “if” is why the next sections (AVL and Red-Black trees) exist. +BSTs are the “ordered dictionary” idea in tree form. You get structure (everything left is smaller, everything right is bigger), which means you can search by repeatedly eliminating half the remaining options, if the tree stays balanced. That “if” is why the next sections (AVL and Red-Black trees) exist. ``` # - [1] + [4] / \ - [2] [3] + [2] [5] / \ \ -[4] [5] [6] +[1] [3] [6] ``` -#### Typical Interface +### Typical Interface The fundamental operations provided by a BST are: -* `insert(value)`: Inserts a new value into the tree, maintaining the BST property. -* `search(value)`: Searches for a specific value in the tree. If the value exists, the operation returns the node. If not, it typically returns `null`. -* `remove(value)`: Deletes a value from the tree, preserving the BST property. If the node to be removed has two children, it's replaced with its in-order predecessor or successor. -* `minimum()`: Retrieves the smallest value in the tree, which is the leftmost node. -* `maximum()`: Retrieves the largest value in the tree, which is the rightmost node. +- `insert(value)`: Inserts a new value into the tree, maintaining the BST property. +- `search(value)`: Searches for a specific value in the tree. If the value exists, the operation returns the node. If not, it typically returns `null`. +- `remove(value)`: Deletes a value from the tree, preserving the BST property. If the node to be removed has two children, it's replaced with its in-order predecessor or successor. +- `minimum()`: Retrieves the smallest value in the tree, which is the leftmost node. +- `maximum()`: Retrieves the largest value in the tree, which is the rightmost node. -#### Time Complexity +### Time Complexity -The efficiency of BST operations depends on the height of the tree. In a balanced tree, this results in an average-case time complexity of `O(log n)`. However, in the worst-case scenario (a degenerate or unbalanced tree), the time complexity degrades to `O(n)`. +The efficiency of BST operations depends on the height of the tree. Each search, insertion, or deletion takes $O(h)$ for height $h$. A balanced tree guarantees $h=O(\log n)$; the average-case column for an ordinary BST assumes a suitable distribution, such as random insertion order. However, in the worst-case scenario (a degenerate or unbalanced tree), the time complexity degrades to `O(n)`. | Operation | Average case | Worst case | | --------- | ------------ | ---------- | @@ -351,16 +356,16 @@ The efficiency of BST operations depends on the height of the tree. In a balance This is the big caution label on plain BSTs: if you insert sorted data into a basic BST, it can collapse into a linked list shape, and you lose the whole point. Balanced variants exist because real data isn’t always friendly. -#### Implementation Details +### Implementation Details A BST is typically implemented using a node-based model where each node encapsulates a value and pointers to the left and right child nodes. A special root pointer points to the first node, if any. Insertion, deletion, and search operations are often implemented recursively to traverse the tree starting from the root, making comparisons at each node to guide the direction of traversal (left or right). This design pattern exploits the recursive nature of the tree structure and results in clean, comprehensible code. -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/collections_and_containers/cpp/binary_search_tree) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/collections_and_containers/python/binary_search_tree) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/collections_and_containers/cpp/binary_search_tree) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/collections_and_containers/python/binary_search_tree) -### AVL Trees +## AVL Trees Named after its inventors, G.M. Adelson-Velsky and E.M. Landis, AVL trees are a subtype of binary search trees (BSTs) that self-balance. Unlike regular BSTs, AVL trees maintain a stringent balance by ensuring the height difference between the left and right subtrees of any node is at most one. This tight control on balance guarantees optimal performance for lookups, insertions, and deletions, making AVL trees highly efficient. @@ -368,30 +373,30 @@ AVL trees are what you use when you want BST behavior without the “it might tu ``` # - [20] + [20] / \ [10] [30] / \ / \ [5] [15] [25] [35] ``` -#### Properties +### Properties The defining characteristics of AVL trees include: -* Every node's left and right subtrees differ in height by at most one. -* Both left and right subtrees of any node are also AVL trees, recursively applying the AVL property across the tree. +- Every node's left and right subtrees differ in height by at most one. +- Both left and right subtrees of any node are also AVL trees, recursively applying the AVL property across the tree. -#### Typical Interface +### Typical Interface AVL trees support standard BST operations with some modifications to maintain balance: -* `insert(value)`: Add a value to the tree while preserving the AVL properties. -* `search(value)`: Find and return a node with the given value in the tree. -* `remove(value)`: Delete a node with the given value while keeping the tree balanced. -* `rotate_left()`, `rotate_right()`: Perform rotation operations to balance the tree after insertions or deletions. +- `insert(value)`: Add a value to the tree while preserving the AVL properties. +- `search(value)`: Find and return a node with the given value in the tree. +- `remove(value)`: Delete a node with the given value while keeping the tree balanced. +- `rotate_left()`, `rotate_right()`: Perform rotation operations to balance the tree after insertions or deletions. -#### Time Complexity +### Time Complexity | Operation | Average | Worst | | --------- | ---------- | ---------- | @@ -400,31 +405,31 @@ AVL trees support standard BST operations with some modifications to maintain ba | Delete | `O(log n)` | `O(log n)` | | Search | `O(log n)` | `O(log n)` | -#### Common Applications +### Common Applications The AVL tree's self-balancing property lends itself to various applications, such as: -* Implementing associative arrays or sets in programming languages. -* Database storage and indexing, where rapid data retrieval is essential. -* Priority queues and dictionaries, where maintaining order matters. -* Word suggestion systems and spell checkers, which benefit from fast lookups. -* Networking algorithms and data compression, requiring efficient data structure. +- Implementing associative arrays or sets in programming languages. +- Database storage and indexing, where rapid data retrieval is essential. +- Priority queues and dictionaries, where maintaining order matters. +- Word suggestion systems and spell checkers, which benefit from fast lookups. +- Networking algorithms and data compression, requiring efficient data structure. -#### Implementation Details +### Implementation Details -The secret to maintaining AVL trees' balance is the application of rotation operations. Four types of rotations can occur based on the tree's balance state: +An AVL node needs rebalancing when the height difference reaches two, not merely when one subtree is taller. Rotations preserve the in-order sequence of keys. Choose a single rotation for an outside-heavy child, or a double rotation for an inside-heavy child; after deletion, a child with equal subtree heights uses a single rotation. Four types of rotations can occur based on the tree's balance state: -1. **Right rotation** occurs when the left subtree of a node is taller than its right subtree. In this rotation, the node's left child becomes the new root, and the old root becomes the right child of this new root. The final step involves updating the heights of the affected nodes. -2. **Left rotation** is used when the right subtree of a node is taller than its left subtree. During this rotation, the node's right child becomes the new root, while the old root becomes the left child of the new root. Afterward, the nodes' heights are updated accordingly. -3. **Right-left rotation** is necessary when the right subtree of a node is taller, but the right child's left subtree is taller than its right subtree. This rotation involves first applying a right rotation on the right child, followed by a left rotation on the original node, with height updates for all affected nodes. -4. **Left-right rotation** is applied when the left subtree of a node is taller, and the left child's right subtree is taller than its left subtree. The process starts with a left rotation on the left child, followed by a right rotation on the original node, and concludes with height updates. +1. Right rotation occurs when the left subtree of a node is taller than its right subtree. In this rotation, the node's left child becomes the new root, and the old root becomes the right child of this new root. The final step involves updating the heights of the affected nodes. +2. Left rotation is used when the right subtree of a node is taller than its left subtree. During this rotation, the node's right child becomes the new root, while the old root becomes the left child of the new root. Afterward, the nodes' heights are updated accordingly. +3. Right-left rotation is necessary when the right subtree of a node is taller, but the right child's left subtree is taller than its right subtree. This rotation involves first applying a right rotation on the right child, followed by a left rotation on the original node, with height updates for all affected nodes. +4. Left-right rotation is applied when the left subtree of a node is taller, and the left child's right subtree is taller than its left subtree. The process starts with a left rotation on the left child, followed by a right rotation on the original node, and concludes with height updates. These rotations keep AVL trees balanced, ensuring consistent performance. -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/collections_and_containers/cpp/avl_tree) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/collections_and_containers/python/avl_tree) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/collections_and_containers/cpp/avl_tree) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/collections_and_containers/python/avl_tree) -### Red-Black Trees +## Red-Black Trees Red-Black trees are a type of self-balancing binary search trees, with each node bearing an additional attribute: color. Each node in a Red-Black tree is colored either red or black, and these color attributes follow specific rules to maintain the tree's balance. This balance ensures that fundamental tree operations such as insertions, deletions, and searches remain efficient. @@ -434,31 +439,31 @@ If AVL trees are “strict about balance,” Red-Black trees are “relaxed but # (30B) / \ - (20R) (40B) + (20R) (40R) / \ / \ (10B) (25B) (35B) (50B) ``` -#### Properties +### Properties Red-Black trees abide by the following essential properties: -* Every node in the tree is colored either red or black. -* The root node of the tree is always black. -* All leaves (NIL or null nodes) are always black. -* If a node is red, then both its parent and its children nodes must be black. -* For each node, any simple path from this node to any of its descendant NIL nodes contains the same number of black nodes. +- Every node in the tree is colored either red or black. +- The root node of the tree is always black. +- All leaves (NIL or null nodes) are always black. +- If a node is red, then both its parent and its children nodes must be black. +- For each node, any simple path from this node to any of its descendant NIL nodes contains the same number of black nodes. -#### Typical Interface +### Typical Interface Red-Black trees support typical BST operations, albeit with additional steps to manage node colors and balance: -* `insert(value)`: Add a value to the tree, managing colors and rotations to maintain balance. -* `search(value)`: Find and return a node with the given value in the tree. -* `remove(value)`: Delete a node with the given value, adjusting colors and performing rotations as necessary to keep the tree balanced. -* `rotate_left()`, `rotate_right()`: Perform rotation operations to maintain balance after insertions or deletions. +- `insert(value)`: Add a value to the tree, managing colors and rotations to maintain balance. +- `search(value)`: Find and return a node with the given value in the tree. +- `remove(value)`: Delete a node with the given value, adjusting colors and performing rotations as necessary to keep the tree balanced. +- `rotate_left()`, `rotate_right()`: Perform rotation operations to maintain balance after insertions or deletions. -#### Time Complexity +### Time Complexity | Operation | Average | Worst | | --------- | ---------- | ---------- | @@ -467,31 +472,31 @@ Red-Black trees support typical BST operations, albeit with additional steps to | Delete | `O(log n)` | `O(log n)` | | Search | `O(log n)` | `O(log n)` | -#### Common Applications +### Common Applications Red-Black trees find application in several areas, including: -* Kernel and system programming, where they are used in the internals of many operating systems. -* Implementing maps, sets, and multi-maps in programming languages like C++ (as part of the Standard Template Library) and Java (in the TreeMap and TreeSet classes). -* Underlying data structures for various advanced tree-based data structures and algorithms, such as splay trees and treaps. +- Kernel and system programming, where they are used in the internals of many operating systems. +- Implementing maps, sets, and multi-maps in programming languages like C++ (as part of the Standard Template Library) and Java (in the TreeMap and TreeSet classes). +- Building augmented ordered structures, such as order-statistic trees and interval trees. Splay trees and treaps are alternative BST designs, not structures built on red-black trees. -#### Implementation Details +### Implementation Details Red-Black trees are a sophisticated binary search tree variant. They require careful management of node colors during insertions and deletions, often involving multiple rebalancing steps: -1. Standard BST insertion is performed initially. +1. Perform standard BST insertion and color the new node red, with black NIL children. 2. If the new node is the tree root, it's colored black to satisfy the Red-Black tree properties. 3. If the new node's parent is black, no further action is needed. -4. If the new node's parent is red, and its uncle (the sibling of its parent) is also red, both the parent and uncle are recolored black, and the grandparent is recolored red (if not the root). This step is known as "recoloring." +4. If the new node's parent is red, and its uncle (the sibling of its parent) is also red, recolor the parent and uncle black and the grandparent red. Continue checking from the grandparent because the violation may have moved upward; restore the root to black at the end. 5. If the new node's parent is red, but its uncle is black or NIL, the tree might need a rotation or a series of rotations. These rotations can be left, right, or a combination of both based on the relative placements of the new node, its parent, and its grandparent. These rotations aim to restructure the tree while preserving the BST property and restoring the Red-Black tree properties. These careful rotations and recoloring steps enable Red-Black trees to maintain balance and ensure efficient performance across various operations. -### Hash Tables +## Hash Tables Hash tables, often referred to as hash maps or dictionaries, are a type of data structure that employs a hash function to pair keys with their corresponding values. This mechanism enables efficient insertion, deletion, and lookup of items within a collection, making hash tables a fundamental tool in many computing scenarios. They are often used in contexts where fast lookup and insertion operations are important, such as database management and networking applications. -Hash tables are the opposite of “keep things in order.” They’re built for one purpose: *find things fast by key*. If you’ve ever wanted the “phone book” experience, given a name, instantly get the number, that’s the vibe. The cost is that you generally give up sorted order, and your performance depends heavily on how well your hash function spreads keys out. +Hash tables organize entries for lookup by key. A name-to-phone-number map is a typical example. Hashing does not itself provide sorted order, and performance depends on how evenly keys are distributed and how collisions are handled. ``` +-----------------+ @@ -505,18 +510,20 @@ Hash tables are the opposite of “keep things in order.” They’re built for +-----------------+ ``` -#### Typical Interface +### Typical Interface Hash tables support a range of operations, which include: -* `is_empty()`: This method checks if the hash table is empty and returns `True` if so, or `False` otherwise. -* `insert(key, value)`: This method adds a new key-value pair to the hash table. -* `retrieve(key)`: This function fetches the value associated with a specified key. -* `update(key, value)`: This operation modifies the value associated with a given key. -* `delete(key)`: This function removes the key-value pair corresponding to the provided key from the hash table. -* `traverse()`: This method allows iteration over all the key-value pairs in the hash table. +- `is_empty()`: This method checks if the hash table is empty and returns `True` if so, or `False` otherwise. +- `insert(key, value)`: This method adds a new key-value pair to the hash table. +- `retrieve(key)`: This function fetches the value associated with a specified key. +- `update(key, value)`: This operation modifies the value associated with a given key. +- `delete(key)`: This function removes the key-value pair corresponding to the provided key from the hash table. +- `traverse()`: This method allows iteration over all the key-value pairs in the hash table. -#### Time Complexity +### Time Complexity + +Expected constant-time lookup assumes bounded load, suitable hash distribution, and constant-time key hashing and comparison. Insertion is also amortized over occasional resizing; hashing a previously unhashed string of length $L$ can itself cost $O(L)$. The performance of hash table operations depends on the hash function's quality, the strategy for resolving collisions, and the load factor (the ratio of the number of key-value pairs to the number of slots or buckets in the table). @@ -527,28 +534,28 @@ The performance of hash table operations depends on the hash function's quality, | Delete | `O(1)` | `O(n)` | | Search | `O(1)` | `O(n)` | -The important “why should I care?” here is: that `O(1)` is the reason hash tables are everywhere. But it’s also not magic, when collisions pile up or the table is overloaded, the worst case can degrade. Good implementations manage load factor and resizing behind the scenes to keep the average case fast. +Expected constant-time lookup explains the usefulness of hash tables. Poor hash distribution or an overloaded table can lengthen searches substantially, so implementations manage their load factor and resize when needed. -#### Common Applications +### Common Applications Hash tables are ubiquitous across various domains, including: -* Databases and caching systems leverage hash tables to achieve quick insertion, deletion, and lookup operations. -* Networking applications use hash tables for rapid data retrieval and updates. -* They are employed within string searching algorithms, such as the Rabin-Karp algorithm, enabling efficient substring search. -* Data visualization often involves creating histograms, which use hash tables for binning data. +- Databases and caching systems leverage hash tables to achieve quick insertion, deletion, and lookup operations. +- Networking applications use hash tables for rapid data retrieval and updates. +- Multi-pattern string search can use a hash table of pattern fingerprints. Basic Rabin–Karp uses rolling hashes and does not itself require a hash table. +- Data visualization often involves creating histograms, which use hash tables for binning data. -#### Implementation Details +### Implementation Details The creation of a hash table involves setting up a data structure to store key-value pairs, designing a hash function to map keys to slots in the data structure, and devising a collision resolution strategy to handle instances where multiple keys hash to the same slot. Common data structures used include arrays (for open addressing) and linked lists or balanced trees (for separate chaining). -##### Hash Function +#### Hash Function -A hash function takes a key (like a string) and turns it into a number, which then decides where to store the data in a table. The goal is to spread out the keys evenly so that no single spot gets overloaded. One simple method for strings is to add up the ASCII values of the characters and then use the remainder from dividing by the table size to pick a slot. Using prime numbers for the table size or in the calculations can help achieve a more balanced spread. +A hash function takes a key (like a string) and turns it into a number, which then decides where to store the data in a table. The goal is to spread out the keys evenly so that no single spot gets overloaded. Adding character codes and reducing modulo the table size is a simple teaching example, but it gives all anagrams the same hash and distributes many inputs poorly. A practical string hash must account for character order. Prime moduli can help particular constructions, but cannot repair a fundamentally poor hash function. -One correction for accuracy: the hash function should depend on the **key**, not on the number of keys. A typical sketch is: +One correction for accuracy: the hash function should depend on the key, not on the number of keys. A typical sketch is: ``` # x: key @@ -558,13 +565,13 @@ h(x) = hash(x) % m That’s the whole game: turn the key into a hash code, then map it into the table range. -##### Collision Resolution +#### Collision Resolution Collisions occur when two or more keys hash to the same slot. There are two common strategies to handle these collisions: chaining and open addressing. Collisions are not a rare edge case, they’re expected. A good hash table design doesn’t try to “avoid collisions completely,” it tries to make collisions cheap to deal with. That’s why understanding the collision strategy matters: it tells you what “worst case” looks like in practice. -##### Chaining +#### Chaining With chaining, each slot or bucket in the hash table acts as a linked list. Each new key-value pair is appended to its designated list. While this approach allows for efficient insertions and deletions, the worst-case time complexity for lookups escalates to `O(n)`, where `n` is the number of keys stored. @@ -583,7 +590,7 @@ index Bucket In the above example, "banana" and "mango" hash to the same index (1), so they are stored in the same linked list. -##### Open Addressing +#### Open Addressing In open addressing, the hash table is an array, and each key-value pair is stored directly in an array slot. When a collision arises, the hash table searches for the next available slot according to a predefined probe sequence. Common probing techniques include linear probing, where the table checks each slot one by one, and quadratic probing, which checks slots based on a quadratic function of the distance from the original hash. @@ -604,10 +611,12 @@ index Bucket In this case, "banana" and "mango" were hashed to index 1, but since it was already occupied by "cherry", they were moved to the next available slots (2 and 3, respectively). -A practical “do/don’t” here: +Deleting an open-addressed entry must preserve its probe chain, for example with a tombstone or backward-shift deletion. Turning an occupied slot directly into a never-used empty slot can make later keys unreachable. + +Practical considerations: -* **Do** understand that open addressing performance depends a lot on how full the table is (high load factors cause longer probe chains). -* **Don’t** assume iteration order means anything useful, hash tables are about key access, not ordering. +- Do understand that open addressing performance depends a lot on how full the table is (high load factors cause longer probe chains). +- Rely on iteration order only if the implementation documents it. Some hash tables preserve insertion order, but hashing alone does not provide sorted order. - [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/collections_and_containers/cpp/hash_table) - [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/collections_and_containers/python/hash_table) diff --git a/notes/dynamic_programming.md b/notes/dynamic_programming.md index 19d8a10..af4f7fa 100644 --- a/notes/dynamic_programming.md +++ b/notes/dynamic_programming.md @@ -1,85 +1,89 @@ -## Dynamic Programming +# Dynamic Programming -Dynamic Programming (DP) is a way to solve complex problems by breaking them into smaller, easier problems. Instead of solving the same small problems again and again, DP **stores their solutions** in a structure like an array, table, or map. This avoids wasting time on repeated calculations and makes the process much faster and more efficient. +Dynamic Programming (DP) is a way to solve complex problems by breaking them into smaller, easier problems. Instead of solving the same small problems again and again, DP stores their solutions in a structure like an array, table, or map. This avoids wasting time on repeated calculations and makes the process much faster and more efficient. -Before you even touch formulas, it helps to know *why DP feels like a superpower*: a lot of “hard” problems aren’t hard because each step is complicated, they’re hard because you keep stumbling into the **same** steps repeatedly. DP is basically the art of noticing that repetition and saying, “Cool, I’ll pay the cost once, then reuse it.” +The main opportunity for DP is repeated work. When different choices lead to the same subproblem, solve that state once and reuse its result. The benefit depends on having fewer distinct states than recursive calls in the naive solution. -DP works best for problems that have two features. The first is **optimal substructure**, which means you can build the solution to a big problem from the solutions to smaller problems. The second is **overlapping subproblems**, where the same smaller problems show up multiple times during the process. By focusing on these features, DP ensures that each part of the problem is solved only once. +DP needs a state definition that contains everything needed to solve a subproblem and a recurrence that combines previously solved states. Optimization problems require optimal substructure; counting and feasibility problems combine counts or Boolean answers instead. Overlapping subproblems make storing these results useful. With a valid dependency order, each state can be computed once. -A practical way to “care” about these two features is this: they’re your green lights. If a problem has both, DP usually turns what feels impossible into something routine. If it doesn’t, DP can become a slow, memory-hungry distraction. So these aren’t academic definitions, they’re decision tools. +Check both the recurrence and the amount of overlap before adding a cache. Reuse can save substantial work, but a large state space can still make a DP solution expensive. This method was introduced by Richard Bellman in the 1950s and has become a valuable tool in areas like computer science, economics, and operations research. It has been used to solve problems that would otherwise take too long by turning slow, exponential-time algorithms into much faster polynomial-time solutions. DP is used in practice for tackling real-world optimization challenges. -And that’s the real payoff: DP isn’t about writing big tables for fun, it’s about making your program behave like a thoughtful planner instead of a frantic guesser. When you see DP in the wild, it’s usually hiding inside something that needs to be fast *and* correct: routing, scheduling, matching, compression, resource allocation, and more. +DP appears in routing, scheduling, matching, compression, and resource allocation. In each setting, the central task is to identify which partial results can be reused without losing information needed for correctness. -### Principles +## Principles -To effectively apply dynamic programming, a problem must satisfy two properties: +Two properties explain when dynamic programming is useful, especially for optimization: A useful mindset here is to treat DP like a story you’re building: first you define the “chapters” (subproblems), then you decide how chapters connect (transitions), and finally you make sure you never rewrite a chapter you already finished (storage). -#### 1. Optimal Substructure +### 1. Optimal Substructure -A problem has **optimal substructure** when the best solution to the overall problem can be built from the best solutions to its smaller parts. In simple terms, solving the smaller pieces perfectly ensures the entire problem is solved perfectly too. +A problem has optimal substructure when the best solution to the overall problem can be built from the best solutions to its smaller parts. In simple terms, solving the smaller pieces perfectly ensures the entire problem is solved perfectly too. Formally, if the optimal solution $S_n$ for a problem of size $n$ can be created by combining the optimal solutions $S_k$ for smaller sizes $k < n$, then the problem is said to exhibit optimal substructure. Why this matters: DP depends on trust. You’re trusting that if you make the best choice for a subproblem, you’re not accidentally ruining the final answer. If that trust holds, you can safely “lock in” sub-results. If it doesn’t, DP will confidently assemble a solution that looks logical but isn’t globally optimal. -**Mathematical Representation**: +Mathematical Representation: Consider a problem where we want to find the optimal value $V(n)$ for a given parameter $n$. If there exists a function $f$ such that: -$$ V(n) = \min_{k} { f(V(k), V(n - k)) } $$ +$ V(n) = \min_{1 \le k < n} f(V(k), V(n-k)), \qquad n \ge 2 $ -then the problem exhibits optimal substructure. +then this is one possible form of an optimal-substructure recurrence, provided the state is sufficient and the combination accounts for every valid choice and its cost. It is not a universal DP formula: many problems need several state parameters or different transitions. Base cases must be defined separately. -**Example**: Shortest Path in Graphs +Example: Shortest Path in Graphs -In the context of graph algorithms, suppose we want to find the shortest path from vertex $A$ to vertex $C$. If $B$ is an intermediate vertex on the shortest path from $A$ to $C$, then the shortest path from $A$ to $C$ is the concatenation of the shortest path from $A$ to $B$ and the shortest path from $B$ to $C$. +In the context of graph algorithms, suppose we want to find the shortest path from vertex $A$ to vertex $C$. If $B$ is an intermediate vertex on the shortest path from $A$ to $C$, then the two segments of that path are shortest paths between their respective endpoints. Otherwise, replacing a segment with a shorter one would improve the whole route. This reasoning assumes finite shortest-path distances; relevant negative cycles can make a shortest walk undefined. -This example also highlights a common “do”: make sure you can clearly explain how a best solution is *composed*. If you can’t describe how the big answer is built from smaller best answers, your “DP solution” is probably just a recursive solution with a table taped to it. +This example also highlights a common “do”: make sure you can clearly explain how a best solution is composed. If you can’t describe how the big answer is built from smaller best answers, your “DP solution” is probably just a recursive solution with a table taped to it. -#### 2. Overlapping Subproblems +### 2. Overlapping Subproblems -A problem has **overlapping subproblems** when it can be divided into smaller problems that are solved multiple times. This happens when the same subproblem appears in different parts of the solution process. In a straightforward recursive approach, this leads to solving the same subproblem repeatedly, which wastes time and resources. Dynamic Programming addresses this by solving each subproblem once and storing the result for reuse, improving efficiency significantly. +A problem has overlapping subproblems when it can be divided into smaller problems that are solved multiple times. This happens when the same subproblem appears in different parts of the solution process. In a straightforward recursive approach, this leads to solving the same subproblem repeatedly, which wastes time and resources. Dynamic Programming addresses this by solving each subproblem once and storing the result for reuse, improving efficiency significantly. -This is the efficiency engine. Optimal substructure tells you DP *can* work; overlapping subproblems tells you DP is *worth it*. If subproblems don’t repeat, caching doesn’t buy you much, and you’re better off with a different approach. +A valid recurrence establishes correctness; repeated subproblems explain why caching helps. Without repetition, storing every intermediate result may add memory overhead without reducing the amount of computation. -**Mathematical Representation**: +Mathematical Representation: Let $S(n)$ be the set of subproblems for problem size $n$. If there exists $s \in S(n)$ such that $s$ appears in $S(k)$ for multiple $k$, the problem has overlapping subproblems. -**Example**: Fibonacci Numbers +Example: Fibonacci Numbers The recursive computation of Fibonacci numbers $F(n) = F(n - 1) + F(n - 2)$ involves recalculating the same Fibonacci numbers multiple times. For instance, to compute $F(5)$, we need to compute $F(4)$ and $F(3)$, both of which require computing $F(2)$ and $F(1)$ multiple times. -Fibonacci is the classic demo not because it’s deep, but because it’s obvious. It teaches the key emotional lesson: recursion can feel elegant while quietly doing ridiculous repeated work, DP keeps the elegance and removes the waste. +Fibonacci makes the overlap easy to see: the same values are recomputed in separate recursive branches. Memoization preserves the recurrence while eliminating those repeated computations. -### Techniques +## Techniques There are two primary methods for implementing dynamic programming algorithms: You can think of these as two different writing styles for the same story. Memoization writes chapters only when needed (and bookmarks them). Tabulation writes every chapter in order, from page 1 onward. Both can end at the same ending; the difference is how they get there. -#### 1. Memoization (Top-Down Approach) +### 1. Memoization (Top-Down Approach) -**Memoization** is a technique used to optimize recursive problem-solving by storing the results of solved subproblems in a data structure, such as a hash table or an array. When the algorithm encounters a subproblem, it first checks if the result is already stored. If it is, the algorithm simply retrieves the cached result instead of recomputing it. This approach reduces redundant calculations, making the solution process faster and more efficient. +Memoization is a technique used to optimize recursive problem-solving by storing the results of solved subproblems in a data structure, such as a hash table or an array. When the algorithm encounters a subproblem, it first checks if the result is already stored. If it is, the algorithm simply retrieves the cached result instead of recomputing it. This approach reduces redundant calculations, making the solution process faster and more efficient. A “do” for memoization: keep the recursion because it matches the problem’s natural definition, but make every subproblem cheap after the first time. A “don’t”: let your memo structure accidentally depend on hidden state (like mutable defaults or changing globals) in a way that makes results inconsistent. -**Algorithm Steps**: +Algorithm Steps: -1. Identify the parameters that uniquely define a subproblem and create a **recursive function** using these parameters. -2. Establish **base cases** to terminate recursion. -3. Initialize a data structure to **store** computed subproblem results. -4. Before computing a subproblem, **check if its result is already stored**. -5. If not already computed, **compute the subproblem's result** and store it. +1. Identify the parameters that uniquely define a subproblem and create a recursive function using these parameters. +2. Establish base cases to terminate recursion. +3. Initialize a data structure to store computed subproblem results. +4. Before computing a subproblem, check if its result is already stored. +5. If not already computed, compute the subproblem's result and store it. -**Example**: Computing Fibonacci Numbers with Memoization +Example: Computing Fibonacci Numbers with Memoization ```python -def fibonacci(n, memo={}): +def fibonacci(n, memo=None): + if n < 0: + raise ValueError("n must be non-negative") + if memo is None: + memo = {} if n in memo: return memo[n] if n <= 1: @@ -89,29 +93,33 @@ def fibonacci(n, memo={}): return memo[n] ``` -**Analysis**: +Analysis: -* **Time Complexity**: $O(n)$, since each number up to $n$ is computed once. -* **Space Complexity**: $O(n)$, due to the memoization structure. +- Time Complexity: $O(n)$, since each number up to $n$ is computed once. +- Space complexity: $O(n)$ stored values and an $O(n)$ recursion stack. Python's recursion limit can prevent this version from handling large $n$ even when sufficient heap memory is available. + +These Fibonacci examples assume integer $n$ and count arithmetic operations as constant time. Exact Fibonacci numbers grow to $\Theta(n)$ bits, so arbitrary-precision arithmetic adds cost; $O(n)$ additions is not the same as $O(n)$ bit operations. One practical note when learning: top-down DP is often easier to write correctly first, because you start from the real question (“solve $n$”) and naturally reach for smaller questions. If you can write the recurrence cleanly, memoization is usually your fastest path to a working solution. -#### 2. Tabulation (Bottom-Up Approach) +### 2. Tabulation (Bottom-Up Approach) -**Tabulation** is a technique that solves a problem by working from the smallest subproblems up to the larger ones in a step-by-step, iterative way. It stores the solutions to these subproblems in a table, often an array, and uses these stored values to solve bigger subproblems until the final solution is reached. Unlike memoization, which works recursively and checks if a solution is already computed, tabulation systematically builds the solution from the ground up. This method is efficient and avoids the overhead of recursive calls. +Tabulation is a technique that solves a problem by working from the smallest subproblems up to the larger ones in a step-by-step, iterative way. It stores the solutions to these subproblems in a table, often an array, and uses these stored values to solve bigger subproblems until the final solution is reached. Unlike memoization, which works recursively and checks if a solution is already computed, tabulation systematically builds the solution from the ground up. This method is efficient and avoids the overhead of recursive calls. -The “why” behind tabulation is control: you decide the exact order states are computed, and you avoid recursion limits and call overhead. The “do” is to fill the table in dependency order. The “don’t” is to fill it in a convenient order that violates dependencies, because then you’ll be reading values that aren’t valid yet. +The “why” behind tabulation is control: you decide the exact order states are computed, and you avoid recursion limits and call overhead. Fill the table in dependency order. Do not fill it in a convenient order that violates dependencies, because then you’ll be reading values that aren’t valid yet. -**Algorithm Steps**: +Algorithm Steps: -1. **Initialize the table** by setting up a structure to store the results of subproblems, and ensure that base cases are properly initialized to handle the simplest instances of the problem. -2. **Iterative computation** is carried out using loops to fill the table, making sure that each subproblem is solved in the correct order before being used to solve larger subproblems. -3. **Construct the solution** by referencing the filled table, using the stored values to derive the final solution to the original problem efficiently. +1. Initialize the table by setting up a structure to store the results of subproblems, and ensure that base cases are properly initialized to handle the simplest instances of the problem. +2. Iterative computation is carried out using loops to fill the table, making sure that each subproblem is solved in the correct order before being used to solve larger subproblems. +3. Construct the solution by referencing the filled table, using the stored values to derive the final solution to the original problem efficiently. -**Example**: Computing Fibonacci Numbers with Tabulation +Example: Computing Fibonacci Numbers with Tabulation ```python def fibonacci(n): + if n < 0: + raise ValueError("n must be non-negative") if n <= 1: return n fib_table = [0] * (n + 1) @@ -121,117 +129,126 @@ def fibonacci(n): return fib_table[n] ``` -**Analysis**: +Analysis: -* The **time complexity** is $O(n)$, since the algorithm iterates from 2 to $n$, computing each Fibonacci number sequentially. -* The **space complexity** is $O(n)$, due to the storage required for the table that holds the Fibonacci numbers up to $n$. +- The time complexity is $O(n)$, since the algorithm iterates from 2 to $n$, computing each Fibonacci number sequentially. +- The space complexity is $O(n)$, due to the storage required for the table that holds the Fibonacci numbers up to $n$. -A nice way to make tabulation feel natural is to ask: “What’s the smallest thing I must know before I can know the next thing?” That question basically *is* the loop order. +A nice way to make tabulation feel natural is to ask: “What’s the smallest thing I must know before I can know the next thing?” That question basically is the loop order. -#### Comparison of Memoization and Tabulation +### Comparison of Memoization and Tabulation | Aspect | Memoization (Top-Down) | Tabulation (Bottom-Up) | | -------------------- | --------------------------------------------------------- | ----------------------------------------------------- | -| **Approach** | Recursive | Iterative | -| **Storage** | Stores solutions as needed | Pre-fills table with solutions to all subproblems | -| **Overhead** | Function call overhead due to recursion | Minimal overhead due to iteration | -| **Flexibility** | May be easier to implement for complex recursive problems | May require careful ordering of computations | -| **Space Efficiency** | Potentially higher due to recursion stack and memoization | Can be more space-efficient with careful table design | +| Approach | Recursive | Iterative | +| Storage | Stores solutions as needed | Pre-fills table with solutions to all subproblems | +| Overhead | Function call overhead due to recursion | Minimal overhead due to iteration | +| Flexibility | May be easier to implement for complex recursive problems | May require careful ordering of computations | +| Space Efficiency | Potentially higher due to recursion stack and memoization | Can be more space-efficient with careful table design | If you’re choosing between them, a simple rule of thumb is: start with memoization to get the recurrence right, then switch to tabulation when you want tighter performance or cleaner memory control. -### Implementation Techniques +## Implementation Techniques Dynamic programming problems are often formulated using recurrence relations, which express the solution to a problem in terms of its subproblems. -This step is where DP stops being “a concept” and becomes “a tool.” Your recurrence is the blueprint: it tells you what a state means and exactly how it is built. Without a recurrence (explicit or implicit), you don’t have DP, you have hope and a table. +A recurrence specifies how the answer for a state depends on other states. Define that relationship, its base cases, and a valid evaluation order before deciding how to store the results. -#### Formulating Recurrence Relations +### Formulating Recurrence Relations -**Example**: Longest Common Subsequence (LCS) +Example: Longest Common Subsequence (LCS) -Given two sequences $X = x_1, x_2, ..., x_m$ and $Y = y_1, y_2, ..., y_n$, the length of their LCS can be defined recursively: +Given two sequences $X = x_1, x_2, ..., x_m$ and $Y = y_1, y_2, ..., y_n$, let $LCS(i,j)$ be the LCS length of their first $i$ and $j$ elements. The recurrence uses one-based character indices and a zero row and column for empty prefixes: $$ LCS(i, j) = \begin{cases} -0 & \text{if } i = 0 \text{ or } j = 0 \ -LCS(i - 1, j - 1) + 1 & \text{if } x_i = y_j \ +0 & \text{if } i = 0 \text{ or } j = 0 \\ +LCS(i - 1, j - 1) + 1 & \text{if } x_i = y_j \\ \max(LCS(i - 1, j), LCS(i, j - 1)) & \text{if } x_i \neq y_j \end{cases} $$ -**Implementation**: +Implementation: We can implement the LCS problem using either memoization or tabulation. With tabulation, we build a two-dimensional table $LCS[0..m][0..n]$ iteratively. -* The **time complexity** is $O(mn)$, as the algorithm processes a grid or matrix of size $m \times n$, iterating through each cell. -* The **space complexity** is $O(mn)$, due to the table storing intermediate results, but this can be reduced to $O(n)$ by optimizing the storage to only keep necessary data for the current and previous rows. +- The time complexity is $O(mn)$, as the algorithm processes a grid or matrix of size $m \times n$, iterating through each cell. +- The space complexity is $O(mn)$, due to the table storing intermediate results, but this can be reduced to $O(n)$ by optimizing the storage to only keep necessary data for the current and previous rows. The reason LCS is such a beloved DP example is that the “choices” are easy to explain: if characters match, you take the diagonal; if they don’t, you take the best of left/top. That clear decision structure is exactly what good DP feels like: simple local rules that reliably build a global answer. -#### State Representation +### State Representation Properly defining the state is crucial for dynamic programming. -* **State variables** are the parameters that uniquely define each subproblem, helping to break down the problem into smaller, manageable components. -* **State transition** refers to the rules or formulas that describe how to move from one state to another, typically using the results of smaller subproblems to solve larger ones. +- State variables are the parameters that uniquely define each subproblem, helping to break down the problem into smaller, manageable components. +- State transition refers to the rules or formulas that describe how to move from one state to another, typically using the results of smaller subproblems to solve larger ones. Think of state design like labeling drawers in a workshop. If the labels are precise, you can find what you need instantly and build bigger things confidently. If the labels are vague, you’ll keep opening drawers, guessing, and making mistakes, even if the math is technically “there.” -**Example**: 0/1 Knapsack Problem +Example: 0/1 Knapsack Problem -* The **problem statement** focuses on selecting $n$ items, each with a weight $w_i$ and value $v_i$, while ensuring the total weight stays within the knapsack capacity $W$, in order to maximize the total value. -* In **state representation**, $dp[i][w]$ represents the maximum value that can be achieved using the first $i$ items with a total weight capacity of $w$. -* **State Transition**: +- Assume positive integer weights, non-negative integer capacity, and at most one copy of each item. Initialize $dp[0][w]=0$ and $dp[i][0]=0$. +- The problem statement focuses on selecting $n$ items, each with a weight $w_i$ and value $v_i$, while ensuring the total weight stays within the knapsack capacity $W$, in order to maximize the total value. +- In state representation, $dp[i][w]$ represents the maximum value that can be achieved using the first $i$ items with a total weight capacity of $w$. +- State Transition: $$ dp[i][w] = \begin{cases} -dp[i - 1][w] & \text{if } w_i > w \ +dp[i - 1][w] & \text{if } w_i > w \\ \max(dp[i - 1][w], dp[i - 1][w - w_i] + v_i) & \text{if } w_i \leq w \end{cases} $$ -**Implementation**: +Implementation: We fill the table $dp[0..n][0..W]$ iteratively based on the state transition. -* The **time complexity** is $O(nW)$, where $n$ is the number of items and $W$ is the capacity of the knapsack, as the algorithm iterates through both items and weights. -* The **space complexity** is $O(nW)$, but this can be optimized to $O(W)$ because each row in the table depends only on the values from the previous row, allowing for space reduction. +- The time complexity is $O(nW)$, where $n$ is the number of items and $W$ is the capacity of the knapsack, as the algorithm iterates through both items and weights. +- The space complexity is $O(nW)$, but this can be optimized to $O(W)$ because each row in the table depends only on the values from the previous row, allowing for space reduction. + +The $O(nW)$ bound is pseudopolynomial: it is polynomial in the numeric capacity $W$, but not in the number of bits needed to encode $W$. DP does not always turn exponential problems into polynomial-time algorithms. -A key “do” here is to interpret your state in plain language. If you can’t say what `dp[i][w]` *means* in one sentence, debugging will be painful. A key “don’t” is mixing meanings (for example, letting `w` sometimes mean “remaining capacity” and sometimes mean “used weight”), that’s how DP tables become nonsense. +A key “do” here is to interpret your state in plain language. If you can’t say what `dp[i][w]` means in one sentence, debugging will be painful. A key “don’t” is mixing meanings (for example, letting `w` sometimes mean “remaining capacity” and sometimes mean “used weight”), that’s how DP tables become nonsense. -### Advanced Concepts +## Advanced Concepts -#### Memory Optimization +### Memory Optimization In some cases, we can optimize space complexity by noticing dependencies between states. -This is where DP graduates from “works” to “works well.” Once the recurrence is correct, you can ask: “Do I truly need *all* previous states, or only a slice of them?” Many DP solutions only depend on the previous row, previous column, or a small window, so storing everything is optional. +This is where DP graduates from “works” to “works well.” Once the recurrence is correct, you can ask: “Do I truly need all previous states, or only a slice of them?” Many DP solutions only depend on the previous row, previous column, or a small window, so storing everything is optional. -**Example**: Since $dp[i][w]$ depends only on $dp[i - 1][w]$ and $dp[i - 1][w - w_i]$, we can use a one-dimensional array and update it in reverse. +Example: Since $dp[i][w]$ depends only on $dp[i - 1][w]$ and $dp[i - 1][w - w_i]$, we can use a one-dimensional array and update it in reverse. ```python -dp = [0] * (W + 1) -for i in range(1, n + 1): - for w in range(W, w_i - 1, -1): - dp[w] = max(dp[w], dp[w - w_i] + v_i) +def knapsack(weights, values, capacity): + best = [0] * (capacity + 1) + for weight, value in zip(weights, values): + for remaining in range(capacity, weight - 1, -1): + best[remaining] = max( + best[remaining], best[remaining - weight] + value + ) + return best[capacity] ``` +Here `weights` and `values` must have the same length, weights must be positive integers, and capacity must be a non-negative integer. Keeping only values is enough to return the optimum; reconstructing the chosen items requires additional predecessor information or recomputation. + One important “do/don’t” hidden in this snippet is the reverse loop. Updating from high to low prevents using the same item multiple times in a single iteration. If you go forward, you silently switch the problem you’re solving. -#### Dealing with Non-Overlapping Subproblems +### Dealing with Non-Overlapping Subproblems -If a problem has **optimal substructure** but does not have **overlapping subproblems**, it is often better to use **Divide and Conquer** instead of Dynamic Programming. Divide and Conquer works by breaking the problem into independent subproblems, solving each one separately, and then combining their solutions. Since there are no repeated subproblems to reuse, storing intermediate results (as in Dynamic Programming) is unnecessary, making Divide and Conquer a more suitable and efficient choice in such cases. +If a problem has optimal substructure but does not have overlapping subproblems, it is often better to use Divide and Conquer instead of Dynamic Programming. Divide and Conquer works by breaking the problem into independent subproblems, solving each one separately, and then combining their solutions. Since there are no repeated subproblems to reuse, storing intermediate results (as in Dynamic Programming) is unnecessary, making Divide and Conquer a more suitable and efficient choice in such cases. This distinction matters because DP isn’t “better recursion”, it’s “recursion plus reuse.” If there’s nothing to reuse, DP is extra work for no gain. Picking the right paradigm is part of writing efficient algorithms, not just correct ones. -**Example**: Merge Sort algorithm divides the list into halves, sorts each half, and then merges the sorted halves. +Example: Merge Sort algorithm divides the list into halves, sorts each half, and then merges the sorted halves. -### Common Terms +## Common Terms These terms show up constantly around DP because DP is really about breaking a big idea into structured pieces. The more comfortable you are with these building blocks, the faster DP problems start to feel like patterns instead of puzzles. -#### Recursion +### Recursion -A process is called **recursion** when a function solves a problem by calling itself, either directly or indirectly, with a smaller instance of the same problem. This continues until the function reaches a **base case**, which is a condition that stops further recursive calls and provides a straightforward solution. Recursion is useful for problems that can naturally be divided into similar smaller subproblems. +A process is called recursion when a function solves a problem by calling itself, either directly or indirectly, with a smaller instance of the same problem. This continues until the function reaches a base case, which is a condition that stops further recursive calls and provides a straightforward solution. Recursion is useful for problems that can naturally be divided into similar smaller subproblems. -**Mathematical Perspective**: +Mathematical Perspective: A recursive function $f(n)$ satisfies: @@ -239,7 +256,7 @@ $$ f(n) = g(f(k), f(n - k)) $$ for some function $g$ and smaller subproblem size $k < n$. -**Example**: Computing $n!$: +Example: Computing $n!$: $$ n! = n \times (n - 1)! $$ @@ -247,44 +264,44 @@ with base case $0! = 1$. In DP, recursion is often your first draft: it expresses the logic cleanly. Then DP adds the missing ingredient: remembering what you already solved. -#### Subset +### Subset For a set $S$, a subset $T$ is a set where every element of $T$ is also an element of $S$. Denoted as $T \subseteq S$. -**Mathematical Properties**: +Mathematical Properties: -* The **total subsets** of a set with $n$ elements is $2^n$, as each element can either be included or excluded from a subset. -* The **power set** of a set $S$, denoted as $\mathcal{P}(S)$, is the set of all possible subsets of $S$, including the empty set and $S$ itself. +- The total subsets of a set with $n$ elements is $2^n$, as each element can either be included or excluded from a subset. +- The power set of a set $S$, denoted as $\mathcal{P}(S)$, is the set of all possible subsets of $S$, including the empty set and $S$ itself. -**Relevance to DP**: +Relevance to DP: Subsets often represent different states or configurations in combinatorial problems, such as the subset-sum problem. The “why you should care” is right in the number $2^n$: subsets explode fast. DP is often the difference between “impossible after n=40” and “fine for n=10,000,” depending on the structure. -#### Subarray +### Subarray A contiguous segment of an array $A$. A subarray is defined by a starting index $i$ and an ending index $j$, with $0 \leq i \leq j < n$, where $n$ is the length of the array. -**Mathematical Representation**: +Mathematical Representation: $$ \text{Subarray } A[i..j] = [A_i, A_{i+1}, ..., A_j] $$ -**Example**: +Example: Given $A = [3, 5, 7, 9]$, the subarray from index $1$ to $2$ is $[5, 7]$. -**Relevance to DP**: +Relevance to DP: Subarray problems include finding the maximum subarray sum (Kadane's algorithm), where dynamic programming efficiently computes optimal subarrays. Subarrays matter in DP because contiguity gives you a natural order, perfect for transitions. When a problem is about “best segment,” “best window,” or “best range,” DP patterns show up immediately. -#### Substring +### Substring A contiguous sequence of characters within a string $S$. Analogous to subarrays in arrays. -**Mathematical Representation**: +Mathematical Representation: A substring $S[i..j]$ is: @@ -292,21 +309,21 @@ $$ S_iS_{i+1}...S_j $$ with $0 \leq i \leq j < \text{length}(S)$. -**Example**: +Example: For $S = "dynamic"$, the substring from index $2$ to $4$ is $"nam"$. -**Relevance to DP**: +Relevance to DP: Substring problems include finding the longest palindromic substring or the longest common substring between two strings. Strings are a DP playground because “prefixes” and “ends at i/j” are easy states to define. If you’ve ever seen a 2D DP table with characters along the top and side, you’ve met this world. -#### Subsequence +### Subsequence A sequence derived from another sequence by deleting zero or more elements without changing the order of the remaining elements. -**Mathematical Representation**: +Mathematical Representation: Given sequence $S$, subsequence $T$ is: @@ -314,210 +331,212 @@ $$ T = [S_{i_1}, S_{i_2}, ..., S_{i_k}] $$ where $0 \leq i_1 < i_2 < ... < i_k < n$. -**Example**: +Example: For $S = [a, b, c, d]$, $[a, c, d]$ is a subsequence. -**Relevance to DP**: +Relevance to DP: The Longest Common Subsequence (LCS) problem is a classic dynamic programming problem. -**LCS Dynamic Programming Formulation**: +LCS Dynamic Programming Formulation: Let $X = x_1 x_2 ... x_m$ and $Y = y_1 y_2 ... y_n$. Define $L[i][j]$ as the length of the LCS of $X[1..i]$ and $Y[1..j]$. -**Recurrence Relation**: +Recurrence Relation: $$ L[i][j] = \begin{cases} -0 & \text{if } i = 0 \text{ or } j = 0 \ -L[i - 1][j - 1] + 1 & \text{if } x_i = y_j \ +0 & \text{if } i = 0 \text{ or } j = 0 \\ +L[i - 1][j - 1] + 1 & \text{if } x_i = y_j \\ \max(L[i - 1][j], L[i][j - 1]) & \text{if } x_i \neq y_j \end{cases} $$ -**Implementation**: +Implementation: We build a two-dimensional table $L[0..m][0..n]$ using the above recurrence. -**Time Complexity**: $O(mn)$ +Time Complexity: $O(mn)$ Subsequences are where DP really earns its reputation, because “skip or take” decisions can branch exponentially. DP keeps that branching logically, but prevents it from becoming computational chaos. -### Practical Considerations +## Practical Considerations -#### Identifying DP Problems +### Identifying DP Problems This section is your pattern-matching toolkit. DP gets dramatically easier once you stop trying to “invent DP” every time and instead learn to recognize the signals that a table of reused sub-results will pay off. -I. If the problem asks for the number of *ways* to do something: +I. If the problem asks for the number of ways to do something: -* Counting paths in a grid. -* Without DP, you would need to enumerate every route. +- Counting paths in a grid. +- DP avoids enumerating every route. An obstacle-free rectangular grid also has a direct binomial-coefficient formula. -II. If the task is to find the *minimum* or *maximum* value under constraints: +II. If the task is to find the minimum or maximum value under constraints: -* Knapsack problem. -* Without DP, you would need to check every subset of items. +- Knapsack problem. +- Subset enumeration is a baseline; branch-and-bound and other methods may also apply. -III. If the same *inputs* appear again during recursion: +III. If the same inputs appear again during recursion: -* Fibonacci numbers. -* Without DP, Fibonacci numbers would be recomputed many times. +- Fibonacci numbers. +- Without DP, Fibonacci numbers would be recomputed many times. -IV. If the solution depends on both the *current step* and *remaining resources* (time, weight, money, length): +IV. If the solution depends on both the current step and remaining resources (time, weight, money, length): -* Scheduling tasks within a time limit. -* Without DP, brute force would be required. +- Scheduling tasks within a time limit. +- A state can capture remaining time, although the exact scheduling constraints determine whether DP, greedy selection, or another method is appropriate. -V. If the problem works with *prefixes, substrings, or subsequences*: +V. If the problem works with prefixes, substrings, or subsequences: -* Longest common subsequence. -* Without DP, exponential checking would be needed. +- Longest common subsequence. +- A naive recursion explores exponentially many choices; caching prefix states removes repeated work. VI. If choices at each step must be explored and combined carefully: -* Coin change with mixed denominations. -* Without DP, you cannot guarantee the fewest coins. +- Coin change with mixed denominations. +- DP guarantees the fewest coins for arbitrary positive integer denominations when its recurrence is correct. Exhaustive search can also guarantee optimality, while greedy works only for suitable coin systems. -VII. If the state space can be stored in a *table or array*: +VII. If the state space can be stored in a table or array: -* Problems with discrete states. -* Without this, problems with infinitely many possibilities (like arbitrary real numbers) cannot be handled. +- Problems with discrete states. +- A finite state space makes direct tabulation possible. Continuous-state problems need additional structure, a different representation, or an approximation; an array is not a requirement for every DP formulation. A good “do” when scanning a problem is to ask: “Can I describe a state with a small number of integers (like i, j, w)?” If yes, DP is often on the table. A good “don’t” is jumping into DP just because the problem is hard, hard problems also show up in greedy, graph, and divide-and-conquer territory. -#### State Design and Transition +### State Design and Transition -* A well-chosen *state* defines what each subproblem represents, while a poorly chosen one leaves the formulation incomplete; for example, `dp[i][w]` in the knapsack problem captures value using `i` items and capacity `w`. -* A correct *transition* connects states consistently, while skipping this leads to undefined progress; in knapsack, the choice to include or exclude an item gives the formula for moving between states. +- A well-chosen state defines what each subproblem represents, while a poorly chosen one leaves the formulation incomplete; for example, `dp[i][w]` in the knapsack problem captures value using `i` items and capacity `w`. +- A correct transition connects states consistently, while skipping this leads to undefined progress; in knapsack, the choice to include or exclude an item gives the formula for moving between states. If DP ever feels “mysterious,” it’s usually because the state meaning isn’t crisp. Once the state is clear, the transition almost writes itself: you ask what decision moves you from smaller states to bigger ones, and you encode that decision as a recurrence. -#### Complexity Optimization +### Complexity Optimization -* Reducing *memory usage* by discarding unnecessary states makes solutions efficient, while failing to do so can waste resources; for example, knapsack space can shrink from `O(nW)` to `O(W)` with a one-dimensional array. -* Using *pruning* to skip impossible paths speeds up computation, while omitting it allows redundant work; in recursive search with memoization, branches exceeding a current best value can be safely ignored. +- Reducing memory usage by discarding unnecessary states makes solutions efficient, while failing to do so can waste resources; for example, knapsack space can shrink from `O(nW)` to `O(W)` with a one-dimensional array. +- Using pruning to skip impossible paths speeds up computation, while omitting it allows redundant work; in recursive search with memoization, a branch can be ignored only when a proven bound shows it cannot improve the answer. Do not cache a partially explored or pruned result as an exact state value if it depends on the current global bound. The key “why” here is that DP gives you structure, but structure can still be expensive. Optimization is about keeping the structure while trimming what you don’t truly need, whether that’s table size, transitions, or states that can never be reached. -#### Common Pitfalls +### Common Pitfalls I. Failure to Define Proper Base Cases -* *Example*: In grid path counting, omitting `dp[0][0] = 1` prevents any valid paths from being constructed. -* *Consequence*: Without correct starting values, the DP table propagates errors and produces incorrect results. +- Example: In grid path counting, omitting `dp[0][0] = 1` prevents any valid paths from being constructed. +- Consequence: Without correct starting values, the DP table propagates errors and produces incorrect results. A simple “do”: before filling anything, write down the smallest cases and ensure they make sense. A simple “don’t”: assume the base case is “obvious” and skip it, DP will punish that immediately. II. Updating States in the Wrong Dependency Order -* *Example*: In knapsack with a 1D array, iterating weights from low to high causes items to be reused multiple times. -* *Consequence*: Using the wrong order inflates computed values and leads to invalid or impossible solutions. +- Example: In knapsack with a 1D array, iterating weights from low to high causes items to be reused multiple times. +- Consequence: Using the wrong order inflates computed values and leads to invalid or impossible solutions. The ordering rule is not a style preference, it’s part of correctness. If your transition reads from states you’ve already updated in the same iteration, you may be solving a different problem than you think. III. Ignoring Special or Edge Case Inputs -* *Example*: In knapsack, a zero-capacity input should return zero value rather than throwing an error. -* *Consequence*: Overlooking edge inputs causes crashes or incorrect answers in boundary conditions. +- Example: In knapsack, a zero-capacity input should return zero value rather than throwing an error. +- Consequence: Overlooking edge inputs causes crashes or incorrect answers in boundary conditions. A good “do” is to test the edges early: zeros, ones, empty inputs, minimal sizes. DP solutions can look perfect on “normal” cases while quietly breaking on the boundaries where states and base cases are most exposed. -### List of Problems +## List of Problems + +For the sum and coin problems below, assume a non-negative integer target and positive integer choices. Zero or negative choices can create nonterminating recurrences when reuse is allowed. For string construction, use nonempty word-bank entries; words may be reused, and their order in the concatenation matters. Enumerating every construction remains output-sensitive even with memoization. -#### Fibonacci Sequence +### Fibonacci Sequence The Fibonacci sequence is a series of numbers where each number is the sum of the two preceding ones, starting with 0 and 1. This sequence is a classic example used to demonstrate recursive algorithms and dynamic programming techniques. -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/cpp/fibonacci/) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/python/fibonacci/) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/cpp/fibonacci/) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/python/fibonacci/) -#### Grid Traveler +### Grid Traveler The Grid Traveler problem involves finding the total number of ways to traverse an `m x n` grid from the top-left corner to the bottom-right corner, with the constraint that you can only move right or down. -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/cpp/grid_traveler) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/python/grid_traveler) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/cpp/grid_traveler) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/python/grid_traveler) -#### Climbing Stairs +### Climbing Stairs The Climbing Stairs problem requires determining the number of distinct ways to reach the top of a staircase with 'n' steps, given that you can climb either 1, 2, or 3 steps at a time. -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/cpp/climb_stairs) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/python/climbing_stairs) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/cpp/climb_stairs) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/python/climbing_stairs) -#### Can Sum +### Can Sum The Can Sum problem involves determining if it is possible to achieve a target sum using any number of elements from a given list of numbers. Each number in the list can be used multiple times. -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/cpp/can_sum) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/python/can_sum) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/cpp/can_sum) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/python/can_sum) -#### How Sum +### How Sum The How Sum problem extends the Can Sum problem by identifying which elements from the list can be combined to sum up to the target value. -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/cpp/how_sum) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/python/how_sum) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/cpp/how_sum) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/python/how_sum) -#### Best Sum +### Best Sum The Best Sum problem further extends the How Sum problem by finding the smallest combination of numbers that add up to exactly the target sum. -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/cpp/best_sum) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/python/best_sum) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/cpp/best_sum) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/python/best_sum) -#### Can Construct +### Can Construct The Can Construct problem involves determining if a target string can be constructed from a given list of substrings. -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/cpp/can_construct) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/python/can_construct) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/cpp/can_construct) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/python/can_construct) -#### Count Construct +### Count Construct The Count Construct problem expands on the Can Construct problem by determining the number of ways a target string can be constructed using a list of substrings. -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/cpp/count_construct) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/python/count_construct) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/cpp/count_construct) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/python/count_construct) -#### All Constructs +### All Constructs The All Constructs problem is a variation of the Count Construct problem, which identifies all the possible combinations of substrings from a list that can be used to form the target string. -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/cpp/all_construct) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/python/all_construct) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/cpp/all_construct) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/python/all_construct) -#### Coins +### Coins The Coins problem aims to find the minimum number of coins needed to make a given value, provided an infinite supply of each coin denomination. -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/cpp/coin_change) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/python/coins) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/cpp/coins) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/python/coins) -#### Longest Common Subsequence +### Longest Common Subsequence The Longest Common Subsequence problem involves finding the longest subsequence that two sequences have in common, where a subsequence is derived by deleting some or no elements without changing the order of the remaining elements. -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/cpp/longest_common_subsequence) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/python/longest_common_subsequence) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/cpp/longest_common_subsequence) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/python/longest_common_subsequence) -#### Longest Increasing Subarray +### Longest Increasing Subarray The Longest Increasing Subarray problem involves identifying the longest contiguous subarray where the elements are strictly increasing. -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/cpp/longest_increasing_subarray) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/python/longest_increasing_subarray) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/cpp/longest_increasing_subarray) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/python/longest_increasing_subarray) -#### Knuth-Morris-Pratt +### Knuth-Morris-Pratt -The Knuth-Morris-Pratt (KMP) algorithm is a pattern searching algorithm that looks for occurrences of a "word" within a main "text string" using preprocessing over the pattern to achieve linear time complexity. +The Knuth-Morris-Pratt (KMP) algorithm is a pattern searching algorithm that looks for occurrences of a "word" within a main "text string" using preprocessing over the pattern to achieve linear time complexity. It is primarily a string-matching algorithm; it appears here because its prefix table reuses information from shorter prefixes. -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/cpp/kmp) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/python/kmp) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/cpp/kmp) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/python/kmp) -#### Minimum Insertions to Form a Palindrome +### Minimum Insertions to Form a Palindrome This problem involves finding the minimum number of insertions needed to transform a given string into a palindrome. The goal is to make the string read the same forwards and backwards with the fewest insertions possible. -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/cpp/minimum_insertions_for_palindrome) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/python/minimum_insertions_for_palindrome) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/cpp/minimum_insertions_for_palindrome) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/dynamic_programming/python/minimum_insertions_for_palindrome) diff --git a/notes/graphs.md b/notes/graphs.md index 3fb1023..59c42a5 100644 --- a/notes/graphs.md +++ b/notes/graphs.md @@ -1,82 +1,84 @@ -## Graphs +# Graphs In many areas of life, we come across systems where elements are deeply interconnected, whether through physical routes, digital networks, or abstract relationships. Graphs offer a flexible way to represent and make sense of these connections. -The reason graphs matter isn’t that they *describe* connections, it’s that they let you **reason** about them. Once you turn a messy “web of relationships” into a graph, you can start asking precise questions: *What’s the fastest route? What’s the weakest link? Who or what is most connected? What happens if this node fails?* That’s why we say that graphs are a useful tool. +The reason graphs matter isn’t that they describe connections, it’s that they let you reason about them. Once you turn a messy “web of relationships” into a graph, you can start asking precise questions: What’s the fastest route? What’s the weakest link? Who or what is most connected? What happens if this node fails? That’s why we say that graphs are a useful tool. Some real-world examples include: -* Underground tunnel networks (subways and transportation systems below the city surface) -* Railway maps (train routes connecting different towns and cities) -* Cities linked by flights (air travel routes between global destinations) -* Networks of pipes (piping systems transporting water, gas, or oil) -* Electrical grids (networks distributing electricity across regions) -* Carbon atoms in a molecule (chemical compounds and the bonds between their constituent atoms) -* Internet (webpages interlinked through hyperlinks or networks of computers) -* Task scheduling (dependencies among tasks that determine their sequence or priority) -* Spread of a disease (understanding how diseases propagate through populations) -* Social networks (people connected through friendships, family ties, or professional relationships) -* Countries and their political alliances (diplomatic ties, trade partnerships, or defense pacts between nations) +- Underground tunnel networks (subways and transportation systems below the city surface) +- Railway maps (train routes connecting different towns and cities) +- Cities linked by flights (air travel routes between global destinations) +- Networks of pipes (piping systems transporting water, gas, or oil) +- Electrical grids (networks distributing electricity across regions) +- Carbon atoms in a molecule (chemical compounds and the bonds between their constituent atoms) +- Internet (webpages interlinked through hyperlinks or networks of computers) +- Task scheduling (dependencies among tasks that determine their sequence or priority) +- Spread of a disease (understanding how diseases propagate through populations) +- Social networks (people connected through friendships, family ties, or professional relationships) +- Countries and their political alliances (diplomatic ties, trade partnerships, or defense pacts between nations) Viewing these systems as graphs lets us dig deeper into how they work, optimize them better, and even predict their behavior more accurately. In fields like computer science and math, graphs are a powerful tool that helps us model and understand these systems effectively. A helpful way to think about this is: graphs are a “universal adapter.” They don’t care whether your nodes are cities, atoms, or tasks. If you can describe “things” and “relationships,” you can model it as a graph, and suddenly a huge library of algorithms becomes available to you. -### Graph Terminology +Throughout the algorithm sections, $V$ and $E$ denote the numbers of vertices and edges in complexity bounds, while $V(G)$ and $E(G)$ denote the sets themselves. Space bounds exclude the input graph unless stated otherwise. Examples use explicit edge lists as the source of truth for weights and directions. + +## Graph Terminology Graph theory has its own language, full of terms that make it easier to talk about and define different types of graphs and their features. Here's an explanation of some important terms: -This vocabulary isn’t just academic; it’s how you avoid confusion later. Graph problems often look similar on the surface, but one word, *directed*, *weighted*, *connected*, can completely change which algorithms work and which ones fail. Learning the terms is like learning traffic signs before driving. - -* A *graph* $G$ is a mathematical structure composed of vertices (also known as nodes or points), forming the set $V(G)$, and edges (also called links or lines), forming the set $E(G)$. Each edge connects two distinct vertices, denoted by pairs ${x, y} \in E(G)$. -* When two vertices, say $x$ and $y$, share an edge, they are termed *adjacent*. The count of adjacent vertices to any vertex $v$ defines its *degree*. Notably, summing the degrees of all vertices in a graph always yields an even number. -* A *path* of length $n$ is a sequence of vertices $v_1 \sim v_2 \sim dots \sim v_{n+1}$, each consecutively connected by edges, with no vertex repeated. -* A *cycle* is a special path where the first and last vertices are identical, forming a closed loop, and no other vertex is repeated within the cycle. -* *Distance* in a graph refers to the shortest path length between two vertices. Essentially, it measures how closely two vertices are connected through the minimal number of edges. -* A *simple graph* is characterized by the absence of self-loops, edges that connect vertices to themselves, and contains no more than one edge between any two distinct vertices. -* A *directed graph* (or digraph) includes edges with specific directions, known as arcs, represented as ordered pairs of vertices. Conversely, an undirected graph treats connections symmetrically, meaning the connection between vertices $A$ and $B$ is identical to the one from $B$ to $A$. -* A *weighted graph* assigns numerical values (weights) to edges, typically non-negative integers. Binary weights (0 or 1) indicate connection presence or absence; numeric weights quantify costs or strengths; and normalized weights scale outgoing connections from a vertex to sum to one, commonly used in probability models. -* Regarding *connectivity*, an undirected graph is *connected* if there's a path linking any pair of vertices. For directed graphs, *weak connectivity* means at least one directional path exists between vertex pairs, while *strong connectivity* demands a path in both directions between every vertex pair. -* Two vertices connected by an edge are called *neighbors*, and the edge itself is said to be *incident* to these vertices. Edges sharing a common vertex are described as *adjacent* edges. -* An *isolated vertex* refers to a vertex with no connecting edges, hence having a degree of zero and no neighbors. -* *Subgraphs* are smaller structures derived from selecting subsets of vertices and edges from a larger graph. Formally, a subgraph $H$ of a graph $G$ has vertex and edge sets entirely contained within those of $G$. -* A *spanning tree* of a graph $G$ is a subgraph that connects all vertices using the minimal number of edges required, specifically $|V| - 1$ edges, and contains no cycles. -* *Bipartite graphs* partition vertices into two distinct groups such that no edges connect vertices within the same group. Formally, a graph is bipartite if it can be split into sets $V_1$ and $V_2$ where each edge links vertices across the two sets only. -* A *complete graph*, denoted by $K_n$, is a simple graph with an edge between every possible pair of vertices. Thus, a complete graph with $n$ vertices contains exactly $\frac{n(n-1)}{2}$ edges. -* *Planar graphs* are graphs that can be drawn on a flat plane without edge intersections except at vertices. Such graphs can be embedded clearly onto a two-dimensional surface without visual confusion. -* An *Eulerian path* travels through each edge exactly once. If this path forms a loop, starting and ending at the same vertex, it becomes an *Eulerian circuit*. For a graph to have an Eulerian circuit, all vertices must have even degrees and the graph must be connected; an Eulerian path exists if exactly zero or two vertices have odd degrees. -* A *Hamiltonian path* traverses every vertex exactly once. If such a path returns to its starting vertex, it becomes a *Hamiltonian circuit*. The decision about the existence of Hamiltonian paths or circuits is notably difficult and classified as an NP-complete problem. -* *Graph isomorphism* describes two graphs, $G$ and $H$, that are structurally identical despite potentially different vertex or edge labels. Formally, an isomorphism is a one-to-one correspondence preserving adjacency between their vertex sets. -* The *degree sequence* of a graph lists all vertex degrees, typically ordered from highest to lowest. This sequence serves as a graph invariant, meaning it is consistent across isomorphic graphs. -* *Graph coloring* involves assigning different labels or "colors" to adjacent vertices. The minimum number of colors required to color a graph without adjacent vertices sharing the same color is called the chromatic number. -* A *tree* is a connected graph with no cycles. It possesses a unique path between every pair of vertices, and adding any extra edge will inevitably create a cycle. Trees are important in many computational applications, especially algorithms and data structures. -* A *forest* consists of multiple disconnected trees. It's essentially an acyclic graph that isn't necessarily connected, serving as a generalization of trees. -* Regarding *connectivity*, a vertex whose removal increases the number of disconnected parts is an *articulation vertex* (or cut vertex). Similarly, an edge whose removal disconnects the graph is called a *bridge* or cut edge. -* A *matching* in a graph is a set of edges with no shared vertices. A *maximum matching* is the largest possible such set. -* An *independent set* is a vertex set with no edges connecting any two vertices. The largest size of such a set is known as the independence number. -* A *clique* is a subset of vertices in which each vertex connects directly to all others. The largest clique size is called the clique number. -* A *vertex cover* is a collection of vertices such that every graph edge touches at least one vertex from the collection. -* A *clique* is the opposite concept, vertices in a clique are all mutually adjacent. The largest possible clique size defines the clique number. -* *Planarity testing* checks if a graph can be drawn without intersecting edges. Kuratowski’s theorem states that non-planarity occurs precisely when a graph contains subgraphs similar to $K_5$ (complete graph of five vertices) or $K_{3,3}$ (complete bipartite graph). -* Common algorithms on graphs include methods for finding shortest paths (Bellman-Ford, Dijkstra's algorithm), discovering minimum spanning trees (Kruskal’s, Prim’s algorithm), and graph traversal (DFS, BFS). These algorithms help solve important graph-related problems in computing and operations research. - -A quick “do/don’t” that saves a lot of pain: always decide early whether your edges are **directed** and/or **weighted**. Many beginner mistakes come from using the right algorithm on the wrong kind of graph (for example, treating one-way roads as two-way, or ignoring weights when “shortest” actually means “cheapest,” not “fewest steps”). - -### Representation of Graphs in Computer Memory +This vocabulary isn’t just academic; it’s how you avoid confusion later. Graph problems often look similar on the surface, but one word, directed, weighted, connected, can completely change which algorithms work and which ones fail. Learning the terms is like learning traffic signs before driving. + +- A graph $G$ is a mathematical structure composed of vertices (also known as nodes or points), forming the set $V(G)$, and edges (also called links or lines), forming the set $E(G)$. In a simple undirected graph, an edge is an unordered pair $\{x,y\}$ of distinct vertices. Directed edges are ordered pairs $(x,y)$; more general graphs may allow self-loops and parallel edges. +- When two vertices, say $x$ and $y$, share an edge, they are termed adjacent. In a simple undirected graph, its degree counts its neighbors. More generally, degree counts incident edge ends, with a self-loop counted twice, so the degree sum is $2|E|$. Directed graphs instead have indegree and outdegree, each summing to $|E|$. +- A path of length $n$ is a sequence of vertices $v_1 \sim v_2 \sim \cdots \sim v_{n+1}$, each consecutively connected by edges, with no vertex repeated. +- A cycle is a closed sequence of edges where the first and last vertices are identical, forming a closed loop, and no other vertex is repeated within the cycle. +- Distance in a graph refers to the shortest path length between two vertices. In an unweighted graph this is the minimum number of edges. In a weighted graph, distance minimizes the sum of edge weights, when a finite shortest distance exists. +- A simple graph is characterized by the absence of self-loops, edges that connect vertices to themselves, and contains no more than one edge between any two distinct vertices. +- A directed graph (or digraph) includes edges with specific directions, known as arcs, represented as ordered pairs of vertices. Conversely, an undirected graph treats connections symmetrically, meaning the connection between vertices $A$ and $B$ is identical to the one from $B$ to $A$. +- A weighted graph assigns numerical values (weights) to edges, which may be positive, zero, or negative and may be nonintegral. An unweighted adjacency matrix uses 0 and 1 for absence and presence; these indicators differ from genuine edge weights of 0 and 1. Numeric weights quantify costs or strengths; and normalized weights scale outgoing connections from a vertex to sum to one, commonly used in probability models. +- Regarding connectivity, an undirected graph is connected if there's a path linking any pair of vertices. For directed graphs, weak connectivity means the graph becomes connected when edge directions are ignored, while strong connectivity demands a path in both directions between every vertex pair. +- Two vertices connected by an edge are called neighbors, and the edge itself is said to be incident to these vertices. Edges sharing a common vertex are described as adjacent edges. +- An isolated vertex refers to a vertex with no connecting edges, hence having a degree of zero and no neighbors. +- Subgraphs are smaller structures derived from selecting subsets of vertices and edges from a larger graph. Formally, a subgraph $H$ of a graph $G$ has vertex and edge sets entirely contained within those of $G$. +- A spanning tree of a graph $G$ is a subgraph that connects all vertices using the minimal number of edges required, specifically $|V| - 1$ edges, and contains no cycles. +- Bipartite graphs partition vertices into two distinct groups such that no edges connect vertices within the same group. Formally, a graph is bipartite if it can be split into sets $V_1$ and $V_2$ where each edge links vertices across the two sets only. +- A complete graph, denoted by $K_n$, is a simple graph with an edge between every possible pair of vertices. Thus, a complete graph with $n$ vertices contains exactly $\frac{n(n-1)}{2}$ edges. +- Planar graphs are graphs that can be drawn on a flat plane without edge intersections except at vertices. Such graphs can be embedded clearly onto a two-dimensional surface without visual confusion. +- An Eulerian trail (often called an Eulerian path) uses every edge exactly once and may revisit vertices. In an undirected graph with at least one edge, all non-isolated vertices must belong to one connected component. Under that condition, a circuit exists exactly when all degrees are even; an open trail exists exactly when two degrees are odd. Directed graphs require separate indegree/outdegree conditions. +- A Hamiltonian path traverses every vertex exactly once. If such a path returns to its starting vertex, it becomes a Hamiltonian circuit. The decision about the existence of Hamiltonian paths or circuits is notably difficult and classified as an NP-complete problem. +- Graph isomorphism describes two graphs, $G$ and $H$, that are structurally identical despite potentially different vertex or edge labels. Formally, an isomorphism is a one-to-one correspondence preserving adjacency between their vertex sets. +- The degree sequence of a graph lists all vertex degrees, typically ordered from highest to lowest. This sequence serves as a graph invariant, meaning it is consistent across isomorphic graphs. +- Graph coloring involves assigning different labels or "colors" to adjacent vertices. The minimum number of colors required to color a graph without adjacent vertices sharing the same color is called the chromatic number. +- A tree is a connected graph with no cycles. It possesses a unique path between every pair of vertices, and adding any extra edge will inevitably create a cycle. Trees are important in many computational applications, especially algorithms and data structures. +- A forest consists of multiple disconnected trees. It's essentially an acyclic graph that isn't necessarily connected, serving as a generalization of trees. +- Regarding connectivity, a vertex whose removal increases the number of disconnected parts is an articulation vertex (or cut vertex). Similarly, an edge whose removal disconnects the graph is called a bridge or cut edge. +- A matching in a graph is a set of edges with no shared vertices. A maximum matching is the largest possible such set. +- An independent set is a vertex set with no edges connecting any two vertices. The largest size of such a set is known as the independence number. +- A clique is a subset of vertices in which each vertex connects directly to all others. The largest clique size is called the clique number. +- A vertex cover is a collection of vertices such that every graph edge touches at least one vertex from the collection. +- These notions are related through graph complementation: a clique in $G$ is an independent set in its complement. A set is a vertex cover exactly when the remaining vertices form an independent set. +- Planarity testing checks if a graph can be drawn without intersecting edges. Kuratowski’s theorem states that non-planarity occurs precisely when a graph contains a subdivision of $K_5$ (complete graph of five vertices) or $K_{3,3}$ (complete bipartite graph). +- Common algorithms on graphs include methods for finding shortest paths (Bellman-Ford, Dijkstra's algorithm), discovering minimum spanning trees (Kruskal’s, Prim’s algorithm), and graph traversal (DFS, BFS). These algorithms help solve important graph-related problems in computing and operations research. + +A quick “do/don’t” that saves a lot of pain: always decide early whether your edges are directed and/or weighted. Many beginner mistakes come from using the right algorithm on the wrong kind of graph (for example, treating one-way roads as two-way, or ignoring weights when “shortest” actually means “cheapest,” not “fewest steps”). + +## Representation of Graphs in Computer Memory Graphs show how things are connected. To work with them on a computer, we need a good way to store and update those connections. The “right” choice depends on what the graph looks like (dense or sparse, directed or undirected, weighted or not) and what you plan to do with it. In practice, most people use one of two formats: an adjacency matrix or an adjacency list. This choice matters because it quietly determines performance. Two programs can run the same algorithm and get the same answers, but one finishes in milliseconds while the other crawls, purely because the graph was stored in a form that made common operations expensive. -#### Adjacency Matrix +### Adjacency Matrix An adjacency matrix represents a graph $G$ with $V$ vertices as a two-dimensional matrix $A$ of size $V \times V$. The rows and columns correspond to vertices, and each cell $A_{ij}$ holds: -* `1` if there is an edge between vertex $i$ and vertex $j$ (or specifically $i \to j$ in a directed graph) -* `0` if no such edge exists -* For weighted graphs, $A_{ij}$ contains the **weight** of the edge; often `0` or `∞` (or `None`) indicates “no edge” +- `1` if there is an edge between vertex $i$ and vertex $j$ (or specifically $i \to j$ in a directed graph) +- `0` if no such edge exists +- For weighted graphs, $A_{ij}$ contains the weight of the edge; use `∞` or `None` for “no edge” if zero-weight edges are allowed -**Same graph used throughout (undirected 4-cycle A–B–C–D–A):** +Same graph used throughout (undirected 4-cycle A–B–C–D–A): ``` (A)------(B) @@ -85,7 +87,7 @@ An adjacency matrix represents a graph $G$ with $V$ vertices as a two-dimensiona (D)------(C) ``` -**Matrix (table form):** +Matrix (table form): | | A | B | C | D | | - | - | - | - | - | @@ -96,7 +98,7 @@ An adjacency matrix represents a graph $G$ with $V$ vertices as a two-dimensiona Here, the matrix indicates a graph with vertices A to D. For instance, vertex A connects with vertices B and D, hence the respective 1s in the matrix. -**Matrix:** +Matrix: ``` 4x4 @@ -110,23 +112,23 @@ Row A | 0 | 1 | 0 | 1 | +---+---+---+---+ ``` -**Notes & Variants** +Notes & Variants -* When an *undirected graph* is represented, the adjacency matrix is symmetric because the connection from node $i$ to node $j$ also implies a connection from node $j$ to node $i$; if this property is omitted, the matrix will misrepresent mutual relationships, such as a road existing in both directions between two cities. -* In the case of a *directed graph*, the adjacency matrix does not need to be symmetric since an edge from node $i$ to node $j$ does not guarantee a reverse edge; without this rule, one might incorrectly assume bidirectional links, such as mistakenly treating a one-way street as two-way. -* A *self-loop* appears as a nonzero entry on the diagonal of the adjacency matrix, indicating that a node is connected to itself; if ignored, the representation will overlook scenarios like a website containing a hyperlink to its own homepage. +- When an undirected graph is represented, the adjacency matrix is symmetric because the connection from node $i$ to node $j$ also implies a connection from node $j$ to node $i$; if this property is omitted, the matrix will misrepresent mutual relationships, such as a road existing in both directions between two cities. +- In the case of a directed graph, the adjacency matrix does not need to be symmetric since an edge from node $i$ to node $j$ does not guarantee a reverse edge; without this rule, one might incorrectly assume bidirectional links, such as mistakenly treating a one-way street as two-way. +- A self-loop appears as a nonzero entry on the diagonal of the adjacency matrix, indicating that a node is connected to itself; if ignored, the representation will overlook scenarios like a website containing a hyperlink to its own homepage. -**Benefits** +Benefits -* An *edge existence check* in an adjacency matrix takes constant time $O(1)$ because the presence of an edge is determined by directly inspecting a single cell; if this property is absent, the lookup could require scanning a list, as in adjacency list representations where finding whether two cities are directly connected may take longer. -* With *simple, compact indexing*, the adjacency matrix aligns well with array-based structures, which makes it helpful for GPU optimizations or bitset operations; without this feature, algorithms relying on linear algebra techniques, such as computing paths with matrix multiplication, become less efficient. +- An edge existence check in an adjacency matrix takes constant time $O(1)$ because the presence of an edge is determined by directly inspecting a single cell; if this property is absent, the lookup could require scanning a list, as in adjacency list representations where finding whether two cities are directly connected may take longer. +- With simple, compact indexing, the adjacency matrix aligns well with array-based structures, which makes it helpful for GPU optimizations or bitset operations; without this feature, algorithms relying on linear algebra techniques, such as computing paths with matrix multiplication, become less efficient. -**Drawbacks** +Drawbacks -* The *space* requirement of an adjacency matrix is always $O(V^2)$, meaning memory usage grows with the square of the number of vertices even if only a few edges exist; if this property is overlooked, sparse networks such as social graphs with millions of users but relatively few connections will be stored inefficiently. -* For *neighbor iteration*, each vertex requires $O(V)$ time because the entire row of the matrix must be scanned to identify adjacent nodes; without recognizing this cost, tasks like finding all friends of a single user in a large social network could become unnecessarily slow. +- The space requirement of an adjacency matrix is always $O(V^2)$, meaning memory usage grows with the square of the number of vertices even if only a few edges exist; if this property is overlooked, sparse networks such as social graphs with millions of users but relatively few connections will be stored inefficiently. +- For neighbor iteration, each vertex requires $O(V)$ time because the entire row of the matrix must be scanned to identify adjacent nodes; without recognizing this cost, tasks like finding all friends of a single user in a large social network could become unnecessarily slow. -**Common Operations (Adjacency Matrix)** +Common Operations (Adjacency Matrix) | Operation | Time | | ---------------------------------- | -------- | @@ -136,18 +138,18 @@ Row A | 0 | 1 | 0 | 1 | | Compute degree of $u$ (undirected) | $O(V)$ | | Traverse all edges | $O(V^2)$ | -**Space Tips** +Space Tips -* Using a *boolean or bitset matrix* allows each adjacency entry to be stored in just one bit, which reduces memory consumption by a factor of eight compared to storing each entry as a byte; if this method is not applied, representing even moderately sized graphs, such as a network of 10,000 nodes, can require far more storage than necessary. -* The approach is most useful when the graph is *dense*, the number of vertices is relatively small, or constant-time edge queries are the primary operation; without these conditions, such as in a sparse graph with millions of vertices, the $V^2$ bit requirement remains wasteful and alternative representations like adjacency lists become more beneficial. +- Using a boolean or bitset matrix allows each adjacency entry to be stored in just one bit, which reduces memory consumption by a factor of eight compared to storing each entry as a byte; if this method is not applied, representing even moderately sized graphs, such as a network of 10,000 nodes, can require far more storage than necessary. +- The approach is most useful when the graph is dense, the number of vertices is relatively small, or constant-time edge queries are the primary operation; without these conditions, such as in a sparse graph with millions of vertices, the $V^2$ bit requirement remains wasteful and alternative representations like adjacency lists become more beneficial. A good way to “feel” adjacency matrices is to imagine a spreadsheet of all possible connections. That’s why they shine when you constantly ask “Is there an edge between u and v?”, but also why they get expensive when most of the spreadsheet is empty. -#### Adjacency List +### Adjacency List An adjacency list stores, for each vertex, the list of its neighbors. It’s usually implemented as an array/vector of lists (or vectors), hash sets, or linked structures. For weighted graphs, each neighbor entry also stores the weight. -**Same graph (A–B–C–D–A) as lists:** +Same graph (A–B–C–D–A) as lists: ``` A -> [B, D] @@ -156,7 +158,7 @@ C -> [B, D] D -> [A, C] ``` -**“In-memory” view (array of heads + per-vertex chains):** +“In-memory” view (array of heads + per-vertex chains): ``` Vertices (index) → 0 1 2 3 @@ -169,63 +171,63 @@ C-list: head -> [B] -> [D] -> NULL D-list: head -> [A] -> [C] -> NULL ``` -**Variants & Notes** +Variants & Notes -* In an *undirected graph* stored as adjacency lists, each edge is represented twice, once in the list of each endpoint, so that both directions can be traversed easily; if this duplication is omitted, traversing from one node to its neighbor may be possible in one direction but not in the other, as with a friendship relation that should be mutual but is stored only once. -* For a *directed graph*, only out-neighbors are recorded in each vertex’s list, meaning that edges can be followed in their given direction; without a separate structure for in-neighbors, tasks like finding all users who link to a webpage require inefficient scanning of every adjacency list. -* In a *weighted graph*, each adjacency list entry stores both the neighbor and the associated weight, such as $(\text{destination}, \text{distance})$; if weights are not included, algorithms like Dijkstra’s shortest path cannot be applied correctly. -* The *order of neighbors* in adjacency lists may be arbitrary, though keeping them sorted allows faster checks for membership; if left unsorted, testing whether two people are directly connected in a social network could require scanning the entire list rather than performing a quicker search. +- In an undirected graph stored as adjacency lists, each edge is represented twice, once in the list of each endpoint, so that both directions can be traversed easily; if this duplication is omitted, traversing from one node to its neighbor may be possible in one direction but not in the other, as with a friendship relation that should be mutual but is stored only once. +- For a directed graph, only out-neighbors are recorded in each vertex’s list, meaning that edges can be followed in their given direction; without a separate structure for in-neighbors, tasks like finding all users who link to a webpage require inefficient scanning of every adjacency list. +- In a weighted graph, each adjacency list entry stores both the neighbor and the associated weight, such as $(\text{destination}, \text{distance})$; if weights are not included, algorithms like Dijkstra’s shortest path cannot be applied correctly. +- The order of neighbors in adjacency lists may be arbitrary, though keeping them sorted allows faster checks for membership; if left unsorted, testing whether two people are directly connected in a social network could require scanning the entire list rather than performing a quicker search. -**Benefits** +Benefits -* The representation is *space-efficient for sparse graphs* because it requires $O(V+E)$ storage, growing only with the number of vertices and edges; without this property, a graph with millions of vertices but relatively few edges, such as a road network, would consume far more memory if stored as a dense matrix. -* For *neighbor iteration*, the time cost is $O(\deg(u))$, since only the actual neighbors of vertex $u$ are examined; if this benefit is absent, each query would need to scan through all possible vertices, as happens in adjacency matrices when identifying a node’s connections. -* In *edge traversals and searches*, adjacency lists support breadth-first search and depth-first search efficiently on sparse graphs because only existing edges are processed; without this design, traversals would involve wasted checks on non-edges, making exploration of large but sparsely connected networks, like airline routes, much slower. +- The representation is space-efficient for sparse graphs because it requires $O(V+E)$ storage, growing only with the number of vertices and edges; without this property, a graph with millions of vertices but relatively few edges, such as a road network, would consume far more memory if stored as a dense matrix. +- For neighbor iteration, the time cost is $O(\deg(u))$, since only the actual neighbors of vertex $u$ are examined; if this benefit is absent, each query would need to scan through all possible vertices, as happens in adjacency matrices when identifying a node’s connections. +- In edge traversals and searches, adjacency lists support breadth-first search and depth-first search efficiently on sparse graphs because only existing edges are processed; without this design, traversals would involve wasted checks on non-edges, making exploration of large but sparsely connected networks, like airline routes, much slower. -**Drawbacks** +Drawbacks -* An *edge existence check* in adjacency lists requires $O(\deg(u))$ time in the worst case because the entire neighbor list may need to be scanned; if a hash set is used for each vertex, the expected time improves to $O(1)$, though at the cost of extra memory and overhead, as seen in fast membership tests within large social networks. -* With respect to *cache locality*, adjacency lists often rely on pointers or scattered memory, which reduces their efficiency on modern hardware; without this drawback, as in dense matrix storage, sequential memory access patterns make repeated operations such as matrix multiplication more beneficial. +- An edge existence check in adjacency lists requires $O(\deg(u))$ time in the worst case because the entire neighbor list may need to be scanned; if a hash set is used for each vertex, the expected time improves to $O(1)$, though at the cost of extra memory and overhead, as seen in fast membership tests within large social networks. +- With respect to cache locality, adjacency lists often rely on pointers or scattered memory, which reduces their efficiency on modern hardware; without this drawback, as in dense matrix storage, sequential memory access patterns make repeated operations such as matrix multiplication more beneficial. -**Common Operations (Adjacency List)** +Common Operations (Adjacency List) | Operation | Time (typical) | | ---------------------------------- | ------------------------------------------- | | Check if edge $u\leftrightarrow v$ | $O(\deg(u))$ (or expected $O(1)$ with hash) | | Add edge | Amortized $O(1)$ (append to list(s)) | -| Remove edge | $O(\deg(u))$ (find & delete) | +| Remove edge | $O(\deg(u))$ directed; $O(\deg(u)+\deg(v))$ undirected | | Iterate neighbors of $u$ | $O(\deg(u))$ | | Traverse all edges | $O(V + E)$ | The choice between these (and other) representations often depends on the graph's characteristics and the specific tasks or operations envisioned. -* Choosing an *adjacency matrix* is helpful when the graph is dense, the number of vertices is moderate, and constant-time edge queries or linear-algebra formulations are beneficial; if this choice is ignored, operations such as repeatedly checking flight connections in a fully connected air network may become slower or harder to express mathematically. -* Opting for an *adjacency list* is useful when the graph is sparse or when neighbor traversal dominates, as in breadth-first search or shortest-path algorithms; without this structure, exploring a large but lightly connected road network would waste time scanning nonexistent edges. +- Choosing an adjacency matrix is helpful when the graph is dense, the number of vertices is moderate, and constant-time edge queries or linear-algebra formulations are beneficial; if this choice is ignored, operations such as repeatedly checking flight connections in a fully connected air network may become slower or harder to express mathematically. +- Opting for an adjacency list is useful when the graph is sparse or when neighbor traversal dominates, as in breadth-first search or shortest-path algorithms; without this structure, exploring a large but lightly connected road network would waste time scanning nonexistent edges. -A simple “do”: pick the representation that makes your *most common operation* cheap. If your algorithm spends most of its time iterating neighbors, adjacency lists are usually the happy path. If it spends most of its time checking whether edges exist, matrices can be surprisingly effective. +A simple “do”: pick the representation that makes your most common operation cheap. If your algorithm spends most of its time iterating neighbors, adjacency lists are usually the happy path. If it spends most of its time checking whether edges exist, matrices can be surprisingly effective. -**Hybrids/Alternatives:** +Hybrids/Alternatives: -* With *CSR/CSC (Compressed Sparse Row/Column)* formats, all neighbors of a vertex are stored contiguously in memory, which improves cache locality and enables fast traversals; without this layout, as in basic pointer-based adjacency lists, high-performance analytics on graphs like web link networks would suffer from slower memory access. -* An *edge list* stores edges simply as $(u,v)$ pairs, making it convenient for graph input, output, and algorithms like Kruskal’s minimum spanning tree; if used for queries such as checking whether two nodes are adjacent, the lack of structure forces scanning the entire list, which becomes inefficient in large graphs. -* In *hash-based adjacency* structures, each vertex’s neighbor set is managed as a hash table, enabling expected $O(1)$ membership tests; without this tradeoff, checking connections in dense social networks requires linear scans, while the hash-based design accelerates lookups at the cost of extra memory. +- With CSR/CSC (Compressed Sparse Row/Column) formats, all neighbors of a vertex are stored contiguously in memory, which improves cache locality and enables fast traversals; without this layout, as in basic pointer-based adjacency lists, high-performance analytics on graphs like web link networks would suffer from slower memory access. +- An edge list stores edges simply as $(u,v)$ pairs, making it convenient for graph input, output, and algorithms like Kruskal’s minimum spanning tree; if used for queries such as checking whether two nodes are adjacent, the lack of structure forces scanning the entire list, which becomes inefficient in large graphs. +- In hash-based adjacency structures, each vertex’s neighbor set is managed as a hash table, enabling expected $O(1)$ membership tests; without this tradeoff, checking connections in dense social networks requires linear scans, while the hash-based design accelerates lookups at the cost of extra memory. -### Planarity +## Planarity -**Planarity** asks: can a graph be drawn on a flat plane so that edges only meet at their endpoints (no crossings)? +Planarity asks: can a graph be drawn on a flat plane so that edges only meet at their endpoints (no crossings)? Why it matters: layouts of circuits, road networks, maps, and data visualizations often rely on planar drawings. Planarity is one of those ideas that sounds like “just drawing,” but it has real consequences. If a graph is non-planar, then any attempt to draw it cleanly in 2D will eventually hit unavoidable clutter. For things like circuit design, that clutter can mean extra layers, extra cost, and harder debugging, so planarity becomes a practical constraint, not a visual preference. -#### What is a planar graph? +### What is a planar graph? -A graph is **planar** if it has **some** drawing in the plane with **no edge crossings**. A messy drawing with crossings doesn’t disqualify it, if you can **redraw** it without crossings, it’s planar. +A graph is planar if it has some drawing in the plane with no edge crossings. A messy drawing with crossings doesn’t disqualify it, if you can redraw it without crossings, it’s planar. -* A crossing-free drawing of a planar graph is called a **planar embedding** (or **plane graph** once embedded). -* In a planar embedding, the plane is divided into **faces** (regions), including the unbounded **outer face**. +- A crossing-free drawing of a planar graph is called a planar embedding (or plane graph once embedded). +- In a planar embedding, the plane is divided into faces (regions), including the unbounded outer face. -**Euler’s Formula (connected planar graphs):** +Euler’s Formula (connected planar graphs): $$ |V| - |E| + |F| = 2 @@ -237,32 +239,32 @@ $$ Euler’s formula is more than a neat identity: it’s a quick reality check. If your counts can’t possibly satisfy it (given the constraints for planar graphs), you already know something must give, either your assumption of planarity is wrong, or your representation is incomplete. -#### Kuratowski’s & Wagner’s characterizations +### Kuratowski’s & Wagner’s characterizations -* According to *Kuratowski’s Theorem*, a graph is planar if and only if it does not contain a subgraph that is a subdivision of $K_5$ or $K_{3,3}$; if this condition is not respected, as in a network with five nodes all mutually connected, the graph cannot be drawn on a plane without edge crossings. -* By *Wagner’s Theorem*, a graph is planar if and only if it has no $K_5$ or $K_{3,3}$ minor, meaning such structures cannot be formed through edge deletions, vertex deletions, or edge contractions; without ruling out these minors, a graph like the complete bipartite structure of three stations each linked to three others cannot be embedded in the plane without overlaps. +- According to Kuratowski’s Theorem, a graph is planar if and only if it does not contain a subgraph that is a subdivision of $K_5$ or $K_{3,3}$; if this condition is not respected, as in a network with five nodes all mutually connected, the graph cannot be drawn on a plane without edge crossings. +- By Wagner’s Theorem, a graph is planar if and only if it has no $K_5$ or $K_{3,3}$ minor, meaning such structures cannot be formed through edge deletions, vertex deletions, or edge contractions; without ruling out these minors, a graph like the complete bipartite structure of three stations each linked to three others cannot be embedded in the plane without overlaps. These are equivalent “forbidden pattern” views. -The “do” here is to look for *structure*, not literal pictures. You don’t need to see a perfect $K_5$ sitting in your graph; you need to spot the ways a graph can *collapse* into one via contractions or subdivisions. That’s why these theorems are powerful: they tell you exactly what kind of complexity ruins planarity. +The “do” here is to look for structure, not literal pictures. You don’t need to see a perfect $K_5$ sitting in your graph; you need to spot the ways a graph can collapse into one via contractions or subdivisions. That’s why these theorems are powerful: they tell you exactly what kind of complexity ruins planarity. -#### Handy planar edge bounds (quick tests) +### Handy planar edge bounds (quick tests) -For a **simple** planar graph with $|V|\ge 3$: +For a simple planar graph with $|V|\ge 3$: -* $|E| \le 3|V| - 6$. -* If the graph is **bipartite**, then $|E| \le 2|V| - 4$. +- $|E| \le 3|V| - 6$. +- If the graph is bipartite, then $|E| \le 2|V| - 4$. These give fast non-planarity proofs: -* $K_5$: $|V|=5, |E|=10 > 3\cdot5-6=9$ ⇒ **non-planar**. -* $K_{3,3}$: $|V|=6, |E|=9 > 2\cdot6-4=8$ ⇒ **non-planar**. +- $K_5$: $|V|=5, |E|=10 > 3\cdot5-6=9$ ⇒ non-planar. +- $K_{3,3}$: $|V|=6, |E|=9 > 2\cdot6-4=8$ ⇒ non-planar. -These bounds are the “quick detective test.” They won’t prove a graph *is* planar, but they can often prove it’s *not* planar instantly, especially when graphs get dense. That’s incredibly useful when you want a fast answer before investing time in more complex checks. +These bounds are the “quick detective test.” They won’t prove a graph is planar, but they can often prove it’s not planar instantly, especially when graphs get dense. That’s incredibly useful when you want a fast answer before investing time in more complex checks. -#### Examples +### Examples -**I. Cycle graphs $C_n$ (always planar)** +I. Cycle graphs $C_n$ (always planar) A 4-cycle $C_4$: @@ -274,42 +276,35 @@ D───C No crossings; faces: 2 (inside + outside). -**II. Complete graph on four vertices $K_4$ (planar)** +II. Complete graph on four vertices $K_4$ (planar) A planar embedding places one vertex inside a triangle: -``` -# - A - / \ -B───C - \ / - D +```text + A + /|\ + / D \ + / / \ \ + B-------C ``` All edges meet only at vertices; no crossings. -**III. Complete graph on five vertices $K_5$ (non-planar)** +III. Complete graph on five vertices $K_5$ (non-planar) -No drawing avoids crossings. Even a “best effort” forces at least one: +No drawing of all ten edges avoids crossings. An explicit edge list prevents a partial sketch from hiding missing edges: -``` -A───B -│╲ ╱│ -│ ╳ │ (some crossing is unavoidable) -│╱ ╲│ -D───C - \ / - E +```text +AB, AC, AD, AE, BC, BD, BE, CD, CE, DE ``` The edge bound $10>9$ (above) certifies non-planarity. -**IV. Complete bipartite graph $K_{3,3}$ (non-planar)** +IV. Complete bipartite graph $K_{3,3}$ (non-planar) The complete bipartite graph $K_{3,3}$ consists of two disjoint vertex sets ${u_1, u_2, u_3}$ and ${v_1, v_2, v_3}$. -Every vertex $u_i$ is connected to every vertex $v_j$, and there are **no edges within a set**. +Every vertex $u_i$ is connected to every vertex $v_j$, and there are no edges within a set. Structure of (K_{3,3}) @@ -323,9 +318,9 @@ v1 v2 v3 This diagram represents all 9 edges: -* (u_1)–(v_1, v_2, v_3) -* (u_2)–(v_1, v_2, v_3) -* (u_3)–(v_1, v_2, v_3) +- (u_1)–(v_1, v_2, v_3) +- (u_2)–(v_1, v_2, v_3) +- (u_3)–(v_1, v_2, v_3) (Edge crossings are unavoidable in any planar drawing.) @@ -343,8 +338,8 @@ $$ e = 9 $$ -Because $K_{3,3}$ is **bipartite**, it has no odd cycles. -For any **planar bipartite graph** with $v \ge 3$, the edge bound is: +Because $K_{3,3}$ is bipartite, it has no odd cycles. +For any planar bipartite graph with $v \ge 3$, the edge bound is: $$ e \le 2v - 4 @@ -360,64 +355,64 @@ $$9 > 8$$ These examples also build intuition: some graphs “want” to live on a plane (cycles, $K_4$), while others contain too much cross-connection pressure ($K_5$, $K_{3,3}$). When you feel that pressure, you start recognizing non-planarity before doing any formal proof. -#### How to check planarity in practice +### How to check planarity in practice -**For small graphs** +For small graphs 1. Rearrange vertices and try to remove crossings. 2. Look for $K_5$ / $K_{3,3}$ (or their subdivisions/minors). 3. Apply the edge bounds above for quick eliminations. -**For large graphs (efficient algorithms)** +For large graphs (efficient algorithms) -* The *Hopcroft–Tarjan* algorithm uses a depth-first search approach to decide planarity in linear time; without such an efficient method, testing whether a circuit layout can be drawn without wire crossings would take longer on large graphs. -* The *Boyer–Myrvold* algorithm also runs in linear time but, in addition to deciding planarity, it produces a planar embedding when one exists; if this feature is absent, as in Hopcroft–Tarjan, a separate procedure would be required to actually construct a drawing of a planar transportation network. +- The Hopcroft–Tarjan algorithm uses a depth-first search approach to decide planarity in linear time; without such an efficient method, testing whether a circuit layout can be drawn without wire crossings would take longer on large graphs. +- Boyer–Myrvold can test planarity and produce a combinatorial embedding or a non-planarity certificate in linear time for simple graphs. An embedding specifies cyclic edge order around vertices; assigning drawing coordinates is a separate task. See the [Boost Graph documentation](https://www.boost.org/doc/libs/latest/libs/graph/doc/html/graph/algorithms/planar/boyer_myrvold.html). Both are widely used in graph drawing, EDA (circuit layout), GIS, and network visualization. The key practical point is that planarity testing isn’t just theoretical, it’s something real tools rely on. If software can quickly decide planarity (and even produce an embedding), it can automatically generate cleaner diagrams and layouts instead of leaving you to manually untangle crossings. -### Traversals +## Traversals What does it mean to traverse a graph? -Graph traversal **can** be done in a way that visits *all* vertices and edges (like a full DFS/BFS), but it doesn’t *have to*. +Graph traversal can be done in a way that visits all vertices and edges (like a full DFS/BFS), but it doesn’t have to. -* If you start DFS or BFS from a single source vertex, you’ll only reach the **connected component** containing that vertex. Any vertices in other components won’t be visited. -* Some algorithms (like shortest path searches, A*, or even partial DFS) intentionally stop early, meaning not all vertices or edges are visited. -* In weighted or directed graphs, you may also skip certain edges depending on the traversal rules. +- From one source, DFS or BFS visits the vertices reachable from it. In an undirected graph this is its connected component; in a directed graph it need not be an entire weak or strong component. +- Some algorithms (like shortest path searches, A*, or even partial DFS) intentionally stop early, meaning not all vertices or edges are visited. +- In weighted or directed graphs, you may also skip certain edges depending on the traversal rules. So the precise way to answer that question is: -> **Graph traversal is a systematic way of exploring vertices and edges, often ensuring complete coverage of the reachable part of the graph , but whether all vertices/edges are visited depends on the algorithm and stopping conditions.** +> Graph traversal is a systematic way of exploring vertices and edges, often ensuring complete coverage of the reachable part of the graph: but whether all vertices/edges are visited depends on the algorithm and stopping conditions. Traversal is the heartbeat of graph algorithms. Almost everything you do on a graph, finding components, shortest paths, detecting cycles, building spanning trees, starts with “walk the graph in a disciplined way.” If graphs are the map, traversal is how you actually move through the territory. -* Graphs, unlike **trees**, don’t have a single starting point like a root. This means we either need to be given a starting vertex or pick one randomly. -* Let’s say we start from a specific vertex, like **$i$**. From there, the traversal explores all connected vertices according to the rules of the chosen method. -* In both **breadth-first search (BFS)** and **depth-first search (DFS)**, the order of visiting vertices depends on how the algorithm is implemented. -* For example, if the starting vertex **A** has three neighbors, like $C, F,$ and $G$, the algorithm doesn’t have a fixed rule for which neighbor to visit first. It could choose any of them based on the way it’s programmed. -* Because of this flexibility, we talk about **a result** of BFS or DFS, rather than **the result**, since different implementations might visit vertices in different orders. +- Graphs, unlike trees, don’t have a single starting point like a root. This means we either need to be given a starting vertex or pick one randomly. +- Let’s say we start from a specific vertex, like $i$. From there, the traversal explores all connected vertices according to the rules of the chosen method. +- In both breadth-first search (BFS) and depth-first search (DFS), the order of visiting vertices depends on how the algorithm is implemented. +- For example, if the starting vertex A has three neighbors, like $C, F,$ and $G$, the algorithm doesn’t have a fixed rule for which neighbor to visit first. It could choose any of them based on the way it’s programmed. +- Because of this flexibility, we talk about a result of BFS or DFS, rather than the result, since different implementations might visit vertices in different orders. -A useful “do” before you choose BFS vs DFS: decide what “close” means in your problem. If “close” means *fewest edges*, BFS naturally fits. If “close” means “go deep and explore structure,” DFS often fits better. Choosing the traversal is choosing the shape of your exploration. +A useful “do” before you choose BFS vs DFS: decide what “close” means in your problem. If “close” means fewest edges, BFS naturally fits. If “close” means “go deep and explore structure,” DFS often fits better. Choosing the traversal is choosing the shape of your exploration. -#### Breadth-First Search (BFS) +### Breadth-First Search (BFS) -Breadth-First Search (BFS) is a graph traversal algorithm that explores a graph **level by level** from a specified start vertex. It first visits all vertices at distance 1 from the start, then all vertices at distance 2, and so on. This makes BFS the natural choice whenever “closest in number of edges” matters. +Breadth-First Search (BFS) is a graph traversal algorithm that explores a graph level by level from a specified start vertex. It first visits all vertices at distance 1 from the start, then all vertices at distance 2, and so on. This makes BFS the natural choice whenever “closest in number of edges” matters. To efficiently keep track of the traversal, BFS employs two primary data structures: -* A **queue** (often named `queue` or `unexplored`) that stores vertices pending exploration in **first-in, first-out (FIFO)** order. -* A **`visited` set** (or boolean array) that records which vertices have already been discovered to prevent revisiting. +- A queue (often named `queue` or `unexplored`) that stores vertices pending exploration in first-in, first-out (FIFO) order. +- A `visited` set (or boolean array) that records which vertices have already been discovered to prevent revisiting. -*Useful additions in practice:* +Useful additions in practice: -* An optional **`parent` map** to reconstruct shortest paths (store `parent[child] = current` when you first discover `child`). -* An optional **`dist` map** to record the edge-distance from the start (`dist[start] = 0`, and when discovering `v` from `u`, set `dist[v] = dist[u] + 1`). +- An optional `parent` map to reconstruct shortest paths (store `parent[child] = current` when you first discover `child`). +- An optional `dist` map to record the edge-distance from the start (`dist[start] = 0`, and when discovering `v` from `u`, set `dist[v] = dist[u] + 1`). BFS feels like ripples in water: start at one node and expand outward in rings. That “ripple” behavior is exactly why BFS gives shortest paths in unweighted graphs, because the first time you reach a node is guaranteed to be via the fewest edges. -**Algorithm Steps** +Algorithm Steps: 1. Pick a start vertex $i$. 2. Set `visited = {i}`, `parent[i] = None`, optionally `dist[i] = 0`, and enqueue $i$ into `queue`. @@ -426,21 +421,21 @@ BFS feels like ripples in water: start at one node and expand outward in rings. 5. For each neighbor `v` of `u`, if `v` is not in `visited`, add it to `visited`, set `parent[v] = u` (and `dist[v] = dist[u] + 1` if tracking), and enqueue `v`. 6. Stop when the queue is empty. -Marking nodes as **visited at the moment they are enqueued** (not when dequeued) is crucial: it prevents the same node from being enqueued multiple times in graphs with cycles or multiple incoming edges. +Marking nodes as visited at the moment they are enqueued (not when dequeued) is crucial: it prevents the same node from being enqueued multiple times in graphs with cycles or multiple incoming edges. -*Reference pseudocode (adjacency-list graph):* +Reference pseudocode (adjacency-list graph): ``` BFS(G, i): visited = {i} parent = {i: None} dist = {i: 0} # optional - queue = [i] + queue = deque([i]) order = [] # optional: visitation order while queue: - u = queue.pop(0) # dequeue + u = queue.popleft() order.append(u) for v in G[u]: # iterate neighbors @@ -453,15 +448,15 @@ BFS(G, i): return order, parent, dist ``` -*Sanity notes:* +Sanity notes: -* The *time* complexity of breadth-first search is $O(V+E)$ because each vertex is enqueued once and each edge is examined once; if this property is overlooked, one might incorrectly assume that exploring a large social graph requires quadratic time rather than scaling efficiently with its size. -* The *space* requirement is $O(V)$ since the algorithm maintains a queue and a visited array, with optional parent or distance arrays if needed; without accounting for this, applying BFS to a network of millions of nodes could be underestimated in memory cost. -* The order in which BFS visits vertices depends on the *neighbor iteration order*, meaning that traversal results can vary between implementations; if this variation is not recognized, two runs on the same graph, such as exploring a road map, may appear inconsistent even though both are correct BFS traversals. +- The time complexity of breadth-first search is $O(V+E)$ because each vertex is enqueued once and each edge is examined once; if this property is overlooked, one might incorrectly assume that exploring a large social graph requires quadratic time rather than scaling efficiently with its size. +- The space requirement is $O(V)$ since the algorithm maintains a queue and a visited array, with optional parent or distance arrays if needed; without accounting for this, applying BFS to a network of millions of nodes could be underestimated in memory cost. +- The order in which BFS visits vertices depends on the neighbor iteration order, meaning that traversal results can vary between implementations; if this variation is not recognized, two runs on the same graph, such as exploring a road map, may appear inconsistent even though both are correct BFS traversals. -**Example** +Example -Graph (undirected) with start at **A**: +Graph (undirected) with start at A: ``` # @@ -479,7 +474,7 @@ Graph (undirected) with start at **A**: Edges: A–B, A–C, B–D, C–E ``` -*Queue/Visited evolution (front → back):* +Queue/Visited evolution (front → back): ``` Step | Dequeued | Action | Queue | Visited @@ -492,7 +487,7 @@ Step | Dequeued | Action | Queue | 5 | E | no new neighbors | [] | {A, B, C, D, E} ``` -*BFS tree and distances from A:* +BFS tree and distances from A: ``` dist[A]=0 @@ -503,43 +498,43 @@ Parents: parent[B]=A, parent[C]=A, parent[D]=B, parent[E]=C Shortest path A→E: backtrack E→C→A ⇒ A - C - E ``` -**Applications** +Applications -* In *shortest path computation on unweighted graphs*, BFS finds the minimum number of edges from a source to all reachable nodes and allows path reconstruction via a parent map; without this approach, one might incorrectly use Dijkstra’s algorithm, which is slower for unweighted networks such as social connections. -* For identifying *connected components in undirected graphs*, BFS is run repeatedly from unvisited vertices, with each traversal discovering one full component; without this method, components in a road map or friendship network may remain undetected. -* When modeling *broadcast or propagation*, BFS naturally mirrors wavefront-like spreading, such as message distribution or infection spread; ignoring this property makes it harder to simulate multi-hop communication in networks. -* During BFS-based *cycle detection in undirected graphs*, encountering a visited neighbor that is not the current vertex’s parent signals a cycle; without this check, cycles in structures like utility grids may be overlooked. -* For *bipartite testing*, BFS alternates colors by level, and the appearance of an edge connecting same-colored nodes disproves bipartiteness; without this strategy, verifying whether a task-assignment graph can be split into two groups becomes more complicated. -* In *multi-source searches*, initializing the queue with several start nodes at distance zero allows efficient nearest-facility queries, such as finding the closest hospital from multiple candidate sites; without this, repeated single-source BFS runs would be less efficient. -* In *topological sorting of DAGs*, a BFS-like procedure processes vertices of indegree zero using a queue, producing a valid ordering; without this method, scheduling tasks with dependency constraints may require less efficient recursive DFS approaches. +- In shortest path computation on unweighted graphs, BFS finds the minimum number of edges from a source to all reachable nodes and allows path reconstruction via a parent map; without this approach, one might incorrectly use Dijkstra’s algorithm, which adds unnecessary priority-queue overhead for unweighted networks such as social connections. +- For identifying connected components in undirected graphs, BFS is run repeatedly from unvisited vertices, with each traversal discovering one full component; without this method, components in a road map or friendship network may remain undetected. +- When modeling broadcast or propagation, BFS naturally mirrors wavefront-like spreading, such as message distribution or infection spread; ignoring this property makes it harder to simulate multi-hop communication in networks. +- During BFS-based cycle detection in undirected graphs, encountering a visited neighbor that is not the current vertex’s parent signals a cycle; without this check, cycles in structures like utility grids may be overlooked. +- For bipartite testing, BFS alternates colors by level, and the appearance of an edge connecting same-colored nodes disproves bipartiteness; without this strategy, verifying whether a task-assignment graph can be split into two groups becomes more complicated. +- In multi-source searches, initializing the queue with several start nodes at distance zero allows efficient nearest-facility queries, such as finding the closest hospital from multiple candidate sites; without this, repeated single-source BFS runs would be less efficient. +- In topological sorting of DAGs, a BFS-like procedure processes vertices of indegree zero using a queue, producing a valid ordering; a DFS-based topological sort has the same $O(V+E)$ asymptotic time. -**Implementation** +Implementation: -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/graphs/cpp/bfs) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/graphs/python/bfs) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/graphs/cpp/bfs) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/graphs/python/bfs) -*Implementation tip:* For dense graphs or when memory locality matters, an adjacency **matrix** can be used, but the usual adjacency **list** representation is more space- and time-efficient for sparse graphs. +Implementation tip: For dense graphs or when memory locality matters, an adjacency matrix can be used, but the usual adjacency list representation is more space- and time-efficient for sparse graphs. A practical “do” for BFS code: use a real queue (like `collections.deque` in Python) instead of `pop(0)` on a list, because list dequeues are $O(n)$. The algorithm stays BFS either way, but performance can change drastically on large graphs. -#### Depth-First Search (DFS) +### Depth-First Search (DFS) -Depth-First Search (DFS) is a graph traversal algorithm that explores **as far as possible** along each branch before backtracking. Starting from a source vertex, it dives down one neighbor, then that neighbor’s neighbor, and so on, only backing up when it runs out of new vertices to visit. +Depth-First Search (DFS) is a graph traversal algorithm that explores as far as possible along each branch before backtracking. Starting from a source vertex, it dives down one neighbor, then that neighbor’s neighbor, and so on, only backing up when it runs out of new vertices to visit. To track the traversal efficiently, DFS typically uses: -* A **call stack** via **recursion** *or* an explicit **stack** data structure (LIFO). -* A **`visited` set** (or boolean array) to avoid revisiting vertices. +- A call stack via recursion or an explicit stack data structure (LIFO). +- A `visited` set (or boolean array) to avoid revisiting vertices. -*Useful additions in practice:* +Useful additions in practice: -* A **`parent` map** to reconstruct paths and build the DFS tree (`parent[child] = current` on discovery). -* Optional **timestamps** (`tin[u]` on entry, `tout[u]` on exit) to reason about edge types, topological order, and low-link computations. -* Optional **`order` lists**: pre-order (on entry) and post-order (on exit). +- A `parent` map to reconstruct paths and build the DFS tree (`parent[child] = current` on discovery). +- Optional timestamps (`tin[u]` on entry, `tout[u]` on exit) to reason about edge types, topological order, and low-link computations. +- Optional `order` lists: pre-order (on entry) and post-order (on exit). DFS is the explorer’s walk: pick a direction, go until you can’t, then backtrack and try the next path. That “go deep first” behavior is why DFS is so useful for structure: it naturally builds trees, reveals cycles, and powers classic algorithms like topological sort and strongly connected components. -**Algorithm Steps** +Algorithm Steps: 1. Pick a start vertex $i$. 2. Initialize `visited[v]=False` for all $v$; optionally set `parent[v]=None`; set a global timer `t=0`. @@ -550,14 +545,15 @@ DFS is the explorer’s walk: pick a direction, go until you can’t, then backt 7. When the DFS from $i$ finishes, if any vertex remains unvisited, choose one and repeat steps 4–6 to cover disconnected components. 8. Stop when no unvisited vertices remain. -Mark vertices **when first discovered** (on entry/push) to prevent infinite loops in cyclic graphs. +The recursive version marks on entry. The stack-only version below marks on pop and discards duplicate pending entries. To reproduce recursive discovery and finish events with $O(V)$ stack space, store neighbor iterators in explicit stack frames. -*Pseudocode (recursive, adjacency list):* +The following DFS examples explore the source's reachable vertices. To cover the entire graph, repeat from unvisited vertices while keeping the same visited set. In a directed graph, this produces a DFS forest rather than identifying strongly connected components by itself. -``` -time = 0 +Pseudocode (recursive, adjacency list): +``` DFS(G, i): + time = 0 visited = set() parent = {i: None} tin = {} @@ -585,39 +581,39 @@ DFS(G, i): return pre, post, parent, tin, tout ``` -*Pseudocode (iterative, traversal order only):* +Pseudocode (iterative, traversal order only): ``` DFS_iter(G, i): visited = set() parent = {i: None} order = [] - stack = [i] + stack = [(i, None)] while stack: - u = stack.pop() # take the top - if u in visited: + u, predecessor = stack.pop() + if u in visited: continue visited.add(u) + parent[u] = predecessor order.append(u) # Push neighbors in reverse of desired visiting order for v in reversed(G[u]): if v not in visited: - parent[v] = u - stack.append(v) + stack.append((v, u)) return order, parent ``` -*Sanity notes:* +Sanity notes: -* The *time* complexity of DFS is $O(V+E)$ because every vertex and edge is processed a constant number of times; if this property is ignored, one might incorrectly assume exponential growth when analyzing networks like citation graphs. -* The *space* complexity is $O(V)$, coming from the visited array and the recursion stack (or an explicit stack in iterative form); without recognizing this, applying DFS to very deep structures such as long linked lists could risk stack overflow unless the iterative approach is used. +- The time complexity of DFS is $O(V+E)$ because every vertex and edge is processed a constant number of times; if this property is ignored, one might incorrectly assume exponential growth when analyzing networks like citation graphs. +- Recursive DFS uses $O(V)$ auxiliary space for marks and the call stack. The simple iterative version shown can hold $O(E)$ pending entries, so its bound is $O(V+E)$. A stack of neighbor iterators avoids duplicates and uses $O(V)$ space. Deep recursive traversals can exceed the language's stack limit. -**Example** +Example -Same graph as the BFS section, start at **A**; assume neighbor order: `B` before `C`, and for `B` the neighbor `D`; for `C` the neighbor `E`. +Same graph as the BFS section, start at A; assume neighbor order: `B` before `C`, and for `B` the neighbor `D`; for `C` the neighbor `E`. ``` # @@ -639,7 +635,7 @@ Same graph as the BFS section, start at **A**; assume neighbor order: `B` before Edges: A–B, A–C, B–D, C–E (undirected) ``` -*Recursive DFS trace (pre-order):* +Recursive DFS trace (pre-order): ``` call DFS(A) @@ -659,7 +655,7 @@ call DFS(A) return A ``` -*Discovery/finish times (one valid outcome):* +Discovery/finish times (one valid outcome): ``` Vertex | tin | tout | parent @@ -671,7 +667,7 @@ C | 6 | 9 | A E | 7 | 8 | C ``` -*Stack/Visited evolution (iterative DFS, top = right):* +Stack/Visited evolution (iterative DFS, top = right): ``` Step | Action | Stack | Visited @@ -687,7 +683,7 @@ Step | Action | Stack | Visited 5 | pop E; visit | [] | {A, B, D, C, E} ``` -*DFS tree (tree edges shown), with preorder: A, B, D, C, E* +DFS tree (tree edges shown), with preorder: A, B, D, C, E ``` A @@ -697,50 +693,50 @@ A └── E ``` -**Applications** +Applications -* In *path existence and reconstruction*, DFS records parent links so that after reaching a target node, the path can be backtracked to the source; without this, finding an explicit route through a maze-like graph would require re-running the search. -* For *topological sorting of DAGs*, running DFS and outputting vertices in reverse postorder yields a valid order; if this step is omitted, dependencies in workflows such as build systems cannot be properly sequenced. -* During *cycle detection*, DFS in undirected graphs reports a cycle when a visited neighbor is not the parent, while in directed graphs the discovery of a back edge to an in-stack node reveals a cycle; without these checks, feedback loops in control systems or task dependencies may go unnoticed. -* To identify *connected components in undirected graphs*, DFS is launched from every unvisited vertex, with each traversal discovering one component; without this method, clusters in social or biological networks remain hidden. -* Using *low-link values* in DFS enables detection of bridges (edges whose removal disconnects the graph) and articulation points (vertices whose removal increases components); if these are not identified, critical links in communication or power networks may be overlooked. -* In *strongly connected components* of directed graphs, algorithms like Tarjan’s and Kosaraju’s use DFS to group vertices where every node is reachable from every other; ignoring this method prevents reliable partitioning of web link graphs or citation networks. -* For *backtracking and state-space search*, DFS systematically explores decision trees and reverses when hitting dead ends, as in solving puzzles like Sudoku or N-Queens; without DFS, these problems would be approached less efficiently with blind trial-and-error. -* With *edge classification in directed graphs*, DFS timestamps allow edges to be labeled as tree, back, forward, or cross, which helps analyze structure and correctness; without this classification, reasoning about graph algorithms such as detecting cycles or proving properties becomes more difficult. +- In path existence and reconstruction, DFS records parent links so that after reaching a target node, the path can be backtracked to the source; without this, finding an explicit route through a maze-like graph would require re-running the search. +- For topological sorting of DAGs, running DFS and outputting vertices in reverse postorder yields a valid order; if this step is omitted, dependencies in workflows such as build systems cannot be properly sequenced. +- During cycle detection, DFS in undirected graphs reports a cycle when a visited neighbor is not the parent, while in directed graphs the discovery of a back edge to an in-stack node reveals a cycle; without these checks, feedback loops in control systems or task dependencies may go unnoticed. +- To identify connected components in undirected graphs, DFS is launched from every unvisited vertex, with each traversal discovering one component; without this method, clusters in social or biological networks remain hidden. +- Using low-link values in DFS enables detection of bridges (edges whose removal disconnects the graph) and articulation points (vertices whose removal increases components); if these are not identified, critical links in communication or power networks may be overlooked. +- In strongly connected components of directed graphs, algorithms like Tarjan’s and Kosaraju’s use DFS to group vertices where every node is reachable from every other; ignoring this method prevents reliable partitioning of web link graphs or citation networks. +- For backtracking and state-space search, DFS systematically explores decision trees and reverses when hitting dead ends, as in solving puzzles like Sudoku or N-Queens; without DFS, these problems would be approached less efficiently with blind trial-and-error. +- With edge classification in directed graphs, DFS timestamps allow edges to be labeled as tree, back, forward, or cross, which helps analyze structure and correctness; without this classification, reasoning about graph algorithms such as detecting cycles or proving properties becomes more difficult. -**Implementation** +Implementation: -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/graphs/cpp/dfs) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/graphs/python/dfs) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/graphs/cpp/dfs) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/graphs/python/dfs) -*Implementation tips:* +Implementation tips: -* For **very deep** or skewed graphs, prefer the **iterative** form to avoid recursion limits. -* If neighbor order matters (e.g., lexicographic traversal), control push order (push in reverse for stacks) or sort adjacency lists. -* For sparse graphs, adjacency **lists** are preferred over adjacency matrices for time/space efficiency. +- For very deep or skewed graphs, prefer the iterative form to avoid recursion limits. +- If neighbor order matters (e.g., lexicographic traversal), control push order (push in reverse for stacks) or sort adjacency lists. +- For sparse graphs, adjacency lists are preferred over adjacency matrices for time/space efficiency. Once you can represent a problem as a graph and you’re fluent in storing it (matrix/list) and traversing it (BFS/DFS), you’ve unlocked the foundation for almost every other graph technique, shortest paths, minimum spanning trees, flows, matchings, connectivity analysis, and more. Graphs are the doorway; traversal is the first step through it. -### Shortest paths +## Shortest paths A common task when dealing with weighted graphs is to find the shortest route between two vertices, such as from vertex $A$ to vertex $B$. Note that there might not be a unique shortest path, since several paths could have the same length. -#### Dijkstra’s Algorithm +### Dijkstra’s Algorithm -Dijkstra’s algorithm computes **shortest paths** from a specified start vertex in a graph with **non-negative edge weights**. It grows a “settled” region outward from the start, always choosing the unsettled vertex with the **smallest known distance** and relaxing its outgoing edges to improve neighbors’ distances. +Dijkstra’s algorithm computes shortest paths from a specified start vertex in a graph with non-negative edge weights. It grows a “settled” region outward from the start, always choosing the unsettled vertex with the smallest known distance and relaxing its outgoing edges to improve neighbors’ distances. To efficiently keep track of the traversal, Dijkstra’s algorithm employs two primary data structures: -* A **min-priority queue** (often named `pq`, `open`, or `unexplored`) keyed by each vertex’s current best known distance from the start. -* A **`dist` map** storing the best known distance to each vertex (∞ initially, except the start), a **`visited`/`finalized` set** to mark vertices whose shortest distance is proven, and a **`parent` map** to reconstruct paths. +- A min-priority queue (often named `pq`, `open`, or `unexplored`) keyed by each vertex’s current best known distance from the start. +- A `dist` map storing the best known distance to each vertex (∞ initially, except the start), a `visited`/`finalized` set to mark vertices whose shortest distance is proven, and a `parent` map to reconstruct paths. -*Useful additions in practice:* +Useful additions in practice: -* A *target-aware early stop* allows Dijkstra’s algorithm to halt once the target vertex is extracted from the priority queue, saving work compared to continuing until all distances are finalized; without this optimization, computing the shortest route between two cities would require processing the entire network unnecessarily. -* With *decrease-key or lazy insertion* strategies, priority queues that lack a decrease-key operation can still work by inserting updated entries and discarding outdated ones when popped; without this adjustment, distance updates in large road networks would be inefficient or require a more complex data structure. -* Adding optional *predecessor lists* enables reconstruction of multiple optimal paths or counting the number of shortest routes; if these lists are not maintained, applications like enumerating all equally fast routes between transit stations cannot be supported. +- A target-aware early stop allows Dijkstra’s algorithm to halt once the target vertex is extracted from the priority queue, saving work compared to continuing until all distances are finalized; without this optimization, computing the shortest route between two cities would require processing the entire network unnecessarily. +- With decrease-key or lazy insertion strategies, priority queues that lack a decrease-key operation can still work by inserting updated entries and discarding outdated ones when popped; without this adjustment, distance updates in large road networks would be inefficient or require a more complex data structure. +- Adding optional predecessor lists can support reconstruction of multiple optimal paths; if these lists are not maintained, applications like enumerating all equally fast routes between transit stations cannot be supported. -**Algorithm Steps** +Algorithm Steps: 1. Pick a start vertex $i$. 2. Set `dist[i] = 0` and `parent[i] = None`; for all other vertices $v \ne i$, set `dist[v] = ∞`. @@ -753,9 +749,9 @@ To efficiently keep track of the traversal, Dijkstra’s algorithm employs two p 9. Stop when the queue is empty (all reachable vertices finalized) or, if you have a target, when that target is finalized. 10. Reconstruct any shortest path by following `parent[·]` backward from the target to $i$. -Vertices are **finalized when they are dequeued** (popped) from the priority queue. With **non-negative** weights, once a vertex is popped the recorded `dist` is **provably optimal**. +Vertices are finalized when their first current entry is removed from the priority queue. Later stale entries are ignored. Non-negative weights make the finalized distance optimal. Before reconstructing a target path, check that the target has finite distance; an unreachable target has no path to reconstruct. -*Reference pseudocode (adjacency-list graph):* +Reference pseudocode (adjacency-list graph): ``` Dijkstra(G, i, target=None): @@ -797,37 +793,28 @@ reconstruct(parent, t): return list(reversed(path)) ``` -*Sanity notes:* +Sanity notes: -* The *time* complexity of Dijkstra’s algorithm depends on the priority queue: $O((V+E)\log V)$ with a binary heap, $O(E+V\log V)$ with a Fibonacci heap, and $O(V^2)$ with a plain array; without this distinction, one might wrongly assume that all implementations scale equally on dense versus sparse road networks. -* The *space* complexity is $O(V)$, needed to store distance values, parent pointers, and priority queue bookkeeping; if underestimated, running Dijkstra on very large graphs such as nationwide transit systems may exceed available memory. -* The *precondition* is that all edge weights must be nonnegative, since the algorithm assumes distances only improve as edges are relaxed; if negative weights exist, as in certain financial models with losses, the computed paths can be incorrect and Bellman–Ford must be used instead. -* In terms of *ordering*, the sequence in which neighbors are processed does not affect correctness, only the handling of ties and slight performance differences; without recognizing this, variations in output order between implementations might be mistakenly interpreted as errors. +- The time complexity of Dijkstra’s algorithm depends on the priority queue: $O((V+E)\log V)$ with a binary heap, $O(E+V\log V)$ with a Fibonacci heap, and $O(V^2)$ with a plain array; without this distinction, one might wrongly assume that all implementations scale equally on dense versus sparse road networks. +- An indexed priority queue uses $O(V)$ auxiliary space. The lazy duplicate-entry version shown can use $O(V+E)$ space and $O(V+E\log(E+2))$ time. On a simple graph this gives the usual $O((V+E)\log V)$ upper bound. +- The precondition is that all edge weights must be nonnegative, because non-negative weights guarantee that a finalized distance cannot later be improved; if negative weights exist, as in certain financial models with losses, the computed paths can be incorrect and Bellman–Ford must be used instead. +- In terms of ordering, the sequence in which neighbors are processed does not affect correctness, only the handling of ties and slight performance differences; without recognizing this, variations in output order between implementations might be mistakenly interpreted as errors. -**Example** +Example -Weighted, undirected graph; start at **A**. Edge weights are on the links. +Weighted, undirected graph; start at A. Edge weights are on the links. -``` -# - ┌────────┐ - │ A │ - └─┬──┬──┬┘ - 4 │ │ │1 - ┌────┘ │ └───┐ - ┌─────▼──┐ │ ┌▼──────┐ - │ B │────┘2 │ C │ - └───┬────┘ └──┬────┘ - 1 │ 4 │ - │ │ - ┌───▼────┐ 3 ┌──▼────┐ - │ E │──────────│ D │ - └────────┘ └───────┘ +```text +A --4-- B --1-- E + \ / | + 1 2 3 + \ / | + C ----4---- D Edges: A–B(4), A–C(1), C–B(2), B–E(1), C–D(4), D–E(3) ``` -*Priority queue / Finalized evolution (front = smallest key):* +Priority queue / Finalized evolution (front = smallest key): ``` Step | Pop (u,dist) | Relaxations (v: new dist, parent) | PQ after push | Finalized @@ -841,7 +828,7 @@ Step | Pop (u,dist) | Relaxations (v: new dist, parent) | PQ after push 6 | (D,5) | , | [] | {A,C,B,E,D} ``` -*Distances and parents (final):* +Distances and parents (final): ``` dist[A]=0 (, ) @@ -853,7 +840,7 @@ dist[D]=5 (C) Shortest path A→E: A → C → B → E (total cost 4) ``` -*Big-picture view of the expanding frontier:* +Big-picture view of the expanding frontier: ``` Settled set grows outward from A by increasing distance. @@ -864,48 +851,48 @@ Shortest path A→E: A → C → B → E (total cost 4) After Step 6: {A, C, B, E, D} (all reachable nodes done) ``` -**Applications** +Applications -* In *single-source shortest paths* with non-negative edge weights, Dijkstra’s algorithm efficiently finds minimum-cost routes in settings like roads, communication networks, or transit systems; without it, travel times or costs could not be computed reliably when distances vary. -* For *navigation and routing*, stopping the search as soon as the destination is extracted from the priority queue avoids unnecessary work; without this early stop, route planning in a road map continues exploring irrelevant regions of the network. -* In *network planning and quality of service (QoS)*, Dijkstra selects minimum-latency or minimum-cost routes when weights are additive and non-negative; without this, designing efficient data or logistics paths becomes more error-prone. -* As a *building block*, Dijkstra underlies algorithms like A* (with zero heuristic), Johnson’s algorithm for all-pairs shortest paths in sparse graphs, and $k$-shortest path variants; without it, these higher-level methods would lack a reliable core procedure. -* In *multi-source Dijkstra*, initializing the priority queue with several starting nodes at distance zero solves nearest-facility queries, such as finding the closest hospital; without this extension, repeated single-source runs would waste time. -* As a *label-setting baseline*, Dijkstra provides the reference solution against which heuristics like A*, ALT landmarks, or contraction hierarchies are compared; without this baseline, heuristic correctness and performance cannot be properly evaluated. -* For *grid pathfinding with terrain costs*, Dijkstra handles non-negative cell costs when no admissible heuristic is available; without it, finding a least-effort path across weighted terrain would require less efficient exhaustive search. +- In single-source shortest paths with non-negative edge weights, Dijkstra’s algorithm efficiently finds minimum-cost routes in settings like roads, communication networks, or transit systems; without it, travel times or costs could not be computed reliably when distances vary. +- For navigation and routing, stopping the search as soon as the destination is extracted from the priority queue avoids unnecessary work; without this early stop, route planning in a road map continues exploring irrelevant regions of the network. +- In network planning and quality of service (QoS), Dijkstra selects minimum-latency or minimum-cost routes when weights are additive and non-negative; without this, designing efficient data or logistics paths becomes more error-prone. +- As a building block, Dijkstra underlies algorithms like A* (with zero heuristic), Johnson’s algorithm for all-pairs shortest paths in sparse graphs, and $k$-shortest path variants; without it, these higher-level methods would lack a reliable core procedure. +- In multi-source Dijkstra, initializing the priority queue with several starting nodes at distance zero solves nearest-facility queries, such as finding the closest hospital; without this extension, repeated single-source runs would waste time. +- As a label-setting baseline, Dijkstra provides the reference solution against which heuristics like A*, ALT landmarks, or contraction hierarchies are compared; without this baseline, heuristic correctness and performance cannot be properly evaluated. +- For grid pathfinding with terrain costs, Dijkstra handles non-negative cell costs when no admissible heuristic is available; without it, finding a least-effort path across weighted terrain would require less efficient exhaustive search. -**Implementation** +Implementation: -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/graphs/cpp/dijkstra) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/graphs/python/dijkstra) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/graphs/cpp/dijkstra) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/graphs/python/dijkstra) -*Implementation tip:* If your PQ has no decrease-key, **push duplicates** on improvement and, when popping a vertex, **skip it** if it’s already finalized or if the popped key doesn’t match `dist[u]`. This “lazy” approach is simple and fast in practice. +Implementation tip: If your PQ has no decrease-key, push duplicates on improvement and, when popping a vertex, skip it if it’s already finalized or if the popped key doesn’t match `dist[u]`. This “lazy” approach is simple and fast in practice. -#### Bellman–Ford Algorithm +### Bellman–Ford Algorithm -Bellman–Ford computes **shortest paths** from a start vertex in graphs that may have **negative edge weights** (but no negative cycles reachable from the start). It works by repeatedly **relaxing** every edge; each full pass can reduce some distances until they stabilize. A final check detects **negative cycles**: if an edge can still be relaxed after $(V-1)$ passes, a reachable negative cycle exists. +Bellman–Ford computes shortest paths from a start vertex in graphs that may have negative edge weights (but no negative cycles reachable from the start). It works by repeatedly relaxing every edge; each full pass can reduce some distances until they stabilize. A final check detects negative cycles: if an edge can still be relaxed after $(V-1)$ passes, a reachable negative cycle exists. To efficiently keep track of the computation, Bellman–Ford employs two primary data structures: -* A **`dist` map** (or array) with the best-known distance to each vertex (initialized to ∞ except the start). -* A **`parent` map** to reconstruct shortest paths (store `parent[v] = u` when relaxing edge $u\!\to\!v$). +- A `dist` map (or array) with the best-known distance to each vertex (initialized to ∞ except the start). +- A `parent` map to reconstruct shortest paths (store `parent[v] = u` when relaxing edge $u\!\to\!v$). -*Useful additions in practice:* +Useful additions in practice: -* With an *edge list*, iterating directly over edges simplifies implementation and keeps updates fast, even if the graph is stored in adjacency lists; without this practice, repeatedly scanning adjacency structures adds unnecessary overhead in each relaxation pass. -* Using an *early exit* allows termination once a full iteration over edges yields no updates, improving efficiency; without this check, the algorithm continues all $V-1$ passes even on graphs like road networks where distances stabilize early. -* For *negative-cycle extraction*, if an update still occurs on the $V$-th pass, backtracking through parent links reveals a cycle; without this step, applications such as financial arbitrage detection cannot identify opportunities caused by negative cycles. -* Adding a *reachability guard* skips edges from vertices with infinite distance, avoiding wasted work on unreached nodes; without this filter, the algorithm needlessly inspects irrelevant edges in disconnected parts of the graph. +- With an edge list, iterating directly over edges simplifies implementation and keeps updates fast, even if the graph is stored in adjacency lists; without this practice, repeatedly scanning adjacency structures adds unnecessary overhead in each relaxation pass. +- Using an early exit allows termination once a full iteration over edges yields no updates, improving efficiency; without this check, the algorithm continues all $V-1$ passes even on graphs like road networks where distances stabilize early. +- For negative-cycle extraction, if an update still occurs on the $V$-th pass, backtracking through parent links reveals a cycle; without this step, applications such as financial arbitrage detection cannot identify opportunities caused by negative cycles. +- Adding a reachability guard skips edges from vertices with infinite distance, avoiding wasted work on unreached nodes; without this filter, the algorithm needlessly inspects irrelevant edges in disconnected parts of the graph. -**Algorithm Steps** +Algorithm Steps: 1. Pick a start vertex $i$. 2. Set `dist[i] = 0` and `parent[i] = None`; for all other vertices $v \ne i$, set `dist[v] = ∞`. 3. Do up to $V-1$ passes: in each pass, scan every directed edge $(u,v,w)$; if `dist[u] + w < dist[v]`, set `dist[v] = dist[u] + w` and `parent[v] = u`. If a full pass makes no changes, stop early. -4. (Optional) Detect negative cycles: if any edge $(u,v,w)$ still satisfies `dist[u] + w < dist[v]`, a reachable negative cycle exists. To extract one, follow `parent` from $v$ for $V$ steps to enter the cycle, then continue until a vertex repeats, collecting the cycle. +4. Detect reachable negative cycles before trusting finite distances: if an edge from a finite-distance vertex can still be relaxed, a reachable negative cycle exists. To extract a cycle, perform a full additional relaxation pass, updating distances and parents, remember a vertex updated on that pass, and follow its parent chain for $V$ steps before collecting the repeated cycle. The detection-only pseudocode below does not perform that extraction pass. 5. To get a shortest path to a target $t$ (when no relevant negative cycle exists), follow `parent[t]` backward to $i$. -*Reference pseudocode (edge list):* +Reference pseudocode (edge list): ``` BellmanFord(V, E, i): # V: set/list of vertices @@ -943,45 +930,29 @@ reconstruct(parent, t): return list(reversed(path)) ``` -*Sanity notes:* +Sanity notes: -* The *time* complexity of Bellman–Ford is $O(VE)$ because each of the $V-1$ relaxation passes scans all edges; without this understanding, one might underestimate the cost of running it on dense graphs with many edges. -* The *space* complexity is $O(V)$, needed for storing distance estimates and parent pointers; if this is not accounted for, memory use may be underestimated in large-scale applications such as road networks. -* The algorithm *handles negative weights* correctly and can also *detect negative cycles* that are reachable from the source; without this feature, Dijkstra’s algorithm would produce incorrect results on graphs with negative edge costs. -* When a reachable *negative cycle* exists, shortest paths to nodes that can be reached from it are undefined, effectively taking value $-\infty$; without recognizing this, results such as infinitely decreasing profit in arbitrage graphs would be misinterpreted as valid finite paths. +- The time complexity of Bellman–Ford is $O(VE)$ because each of the $V-1$ relaxation passes scans all edges; without this understanding, one might underestimate the cost of running it on dense graphs with many edges. +- The space complexity is $O(V)$, needed for storing distance estimates and parent pointers; if this is not accounted for, memory use may be underestimated in large-scale applications such as road networks. +- The algorithm handles negative weights correctly and can also detect negative cycles that are reachable from the source; without this feature, Dijkstra’s algorithm would produce incorrect results on graphs with negative edge costs. +- When a reachable negative cycle exists, shortest paths to nodes that can be reached from it are undefined, effectively taking value $-\infty$; without recognizing this, results such as infinitely decreasing profit in arbitrage graphs would be misinterpreted as valid finite paths. -**Example** +Example -Directed, weighted graph; start at **A**. (Negative edges allowed; **no** negative cycles here.) +Directed, weighted graph; start at A. (Negative edges allowed; no negative cycles here.) +```text +A -> B (4) A -> C (2) +B -> C (-1) B -> D (2) +C -> B (1) C -> D (5) C -> E (3) +D -> E (-3) ``` -# - ┌─────────┐ - │ A │ - └──┬───┬──┘ - 4 │ │ 2 - │ └───────────────┐ -┌──────────▼───┐ -1 ┌───▼─────┐ -│ B │ ─────────►│ C │ -└──────┬───────┘ └──┬──────┘ - 2 │ 5 │ - │ │ -┌──────▼──────┐ -3 ┌───▼─────┐ -│ D │ ◄─────── │ E │ -└──────────────┘ └─────────┘ - -Also: -A → C (2) -C → B (1) -C → E (3) -(Edges shown with weights on arrows) -``` - -*Edges list:* + +Edges list: `A→B(4), A→C(2), B→C(-1), B→D(2), C→B(1), C→D(5), C→E(3), D→E(-3)` -*Relaxation trace (dist after each full pass; start A):* +Relaxation trace (dist after each full pass; start A): ``` Init (pass 0): @@ -989,7 +960,7 @@ Init (pass 0): After pass 1: A=0, B=3, C=2, D=6, E=3 - (A→B=4, A→C=2; C→B improved B to 3; B→D=5? (via B gives 6); D→E=-3 gives E=3) + (B→D sets D=6 before C→B lowers B to 3; D→E then sets E=3) After pass 2: A=0, B=3, C=2, D=5, E=2 @@ -999,7 +970,7 @@ After pass 3: A=0, B=3, C=2, D=5, E=2 (no changes → early stop) ``` -*Parents / shortest paths (one valid set):* +Parents / shortest paths (one valid set): ``` parent[A]=None @@ -1012,28 +983,28 @@ Example shortest path A→E: A → C → B → D → E with total cost 2 + 1 + 2 + (-3) = 2 ``` -*Negative-cycle detection (illustration):* -If we **add** an extra edge `E→C(-4)`, the cycle `C → D → E → C` has total weight `5 + (-3) + (-4) = -2` (negative). -Bellman–Ford would perform a $V$-th pass and still find an improvement (e.g., relaxing `E→C(-4)`), so it reports a **reachable negative cycle**. +Negative-cycle detection (illustration): +If we add an extra edge `E→C(-4)`, the cycle `C → D → E → C` has total weight `5 + (-3) + (-4) = -2` (negative). +Bellman–Ford would perform a $V$-th pass and still find an improvement (e.g., relaxing `E→C(-4)`), so it reports a reachable negative cycle. -**Applications** +Applications -* In *shortest path problems with negative edges*, Bellman–Ford is applicable where Dijkstra or A* fail, such as road networks with toll credits; without this method, these graphs cannot be handled correctly. -* For *arbitrage detection* in currency or financial markets, converting exchange rates into $\log$ weights makes profit loops appear as negative cycles; without Bellman–Ford, such opportunities cannot be systematically identified. -* In solving *difference constraints* of the form $x_v - x_u \leq w$, the algorithm checks feasibility by detecting whether any negative cycles exist; without this check, inconsistent scheduling or timing systems may go unnoticed. -* As a *robust baseline*, Bellman–Ford verifies results of faster algorithms or initializes methods like Johnson’s for all-pairs shortest paths; without it, correctness guarantees in sparse-graph all-pairs problems would be weaker. -* For *graphs with penalties or credits*, where some transitions decrease accumulated cost, Bellman–Ford models these adjustments accurately; without it, such systems, like transport discounts or energy recovery paths, cannot be represented properly. +- In shortest path problems with negative edges, Bellman–Ford is applicable where Dijkstra or A* fail, such as road networks with toll credits; without this method, these graphs cannot be handled correctly. +- For arbitrage detection in currency or financial markets, converting an exchange rate $r$ into weight $-\log r$ makes profit loops appear as negative cycles; without Bellman–Ford, such opportunities cannot be systematically identified. +- In solving difference constraints of the form $x_v - x_u \leq w$, the algorithm checks feasibility by detecting whether any negative cycles exist; without this check, inconsistent scheduling or timing systems may go unnoticed. +- As a robust baseline, Bellman–Ford verifies results of faster algorithms or initializes methods like Johnson’s for all-pairs shortest paths; without it, correctness guarantees in sparse-graph all-pairs problems would be weaker. +- For graphs with penalties or credits, where some transitions decrease accumulated cost, Bellman–Ford models these adjustments accurately; without it, such systems, like transport discounts or energy recovery paths, cannot be represented properly. -##### Implementation +Implementation: -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/graphs/cpp/bellman_ford) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/graphs/python/bellman_ford) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/graphs/cpp/bellman_ford) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/graphs/python/bellman_ford) -*Implementation tip:* For **all-pairs** on sparse graphs with possible negative edges, use **Johnson’s algorithm**: run Bellman–Ford once from a super-source to reweight edges (no negatives), then run **Dijkstra** from each vertex. +Implementation tip: For all-pairs on sparse graphs with possible negative edges, use Johnson’s algorithm: run Bellman–Ford once from a super-source to reweight edges (no negatives), then run Dijkstra from each vertex. -#### A* (A-Star) Algorithm +### A* (A-Star) Algorithm -A* is a best-first search that finds a **least-cost path** from a start to a goal by minimizing +A* is a best-first search that finds a least-cost path from a start to a goal by minimizing $$ f(n) = g(n) + h(n), @@ -1041,20 +1012,20 @@ $$ where: -* $g(n)$ = cost from start to $n$ (so far), -* $h(n)$ = heuristic estimate of the remaining cost from $n$ to the goal. +- $g(n)$ = cost from start to $n$ (so far), +- $h(n)$ = heuristic estimate of the remaining cost from $n$ to the goal. -If $h$ is **admissible** (never overestimates) and **consistent** (triangle inequality), A* is **optimal** and never needs to “reopen” closed nodes. +Assume a finite graph, non-negative edge costs, and $h(goal)=0$. With a consistent heuristic, A* is optimal without reopening expanded nodes. With an admissible but inconsistent heuristic, reopen a node whenever a better path reaches it. These guarantees apply when the goal is removed from the queue with a current entry, not when it is first discovered. [Berkeley CS 188 notes](https://inst.eecs.berkeley.edu/~cs188/fa22/assets/notes/cs188-fa22-note02.pdf). -**Core data structures** +Core data structures -* The *open set* is a min-priority queue keyed by the evaluation function $f=g+h$, storing nodes pending expansion; without it, selecting the next most promising state in pathfinding would require inefficient linear scans. -* The *closed set* contains nodes already expanded and finalized, preventing reprocessing; if omitted, the algorithm may revisit the same grid cells or graph states repeatedly, wasting time. -* The *$g$ map* tracks the best known cost-so-far to each node, ensuring paths are only updated when improvements are found; without it, the algorithm cannot correctly accumulate and compare path costs. -* The *parent map* stores predecessors so that a complete path can be reconstructed once the target is reached; if absent, the algorithm would output only a final distance without the actual route. -* An optional *heuristic cache* and *tie-breaker* (such as preferring larger $g$ or smaller $h$ when $f$ ties) can improve efficiency and consistency; without these, the search may expand more nodes than necessary or return different paths under equivalent conditions. +- The open set is a min-priority queue keyed by the evaluation function $f=g+h$, storing nodes pending expansion; without it, selecting the next most promising state in pathfinding would require inefficient linear scans. +- The closed set contains nodes already expanded and finalized, preventing reprocessing; if omitted, the algorithm may revisit the same grid cells or graph states repeatedly, wasting time. +- The *$g$ map* tracks the best known cost-so-far to each node, ensuring paths are only updated when improvements are found; without it, the algorithm cannot correctly accumulate and compare path costs. +- The parent map stores predecessors so that a complete path can be reconstructed once the target is reached; if absent, the algorithm would output only a final distance without the actual route. +- An optional heuristic cache and tie-breaker (such as preferring larger $g$ or smaller $h$ when $f$ ties) can improve efficiency and consistency; without these, the search may expand more nodes than necessary or return different paths under equivalent conditions. -**Algorithm Steps** +Algorithm Steps: 1. Put `start` in `open` (a min-priority queue by `f`); set `g[start]=0`, `f[start]=h(start)`, `parent[start]=None`; initialize `closed = ∅`. 2. While `open` is nonempty, repeat steps 3–7. @@ -1065,37 +1036,38 @@ If $h$ is **admissible** (never overestimates) and **consistent** (triangle ineq 7. If `v` not in `g` or `tentative < g[v]`, set `parent[v]=u`, `g[v]=tentative`, `f[v]=g[v]+h(v)`, and push `v` into `open` (even if it was already there with a worse key). 8. If the loop ends because `open` is empty, no path exists. -*Mark neighbors **when you enqueue them** (by storing their best `g`) to avoid duplicate work; with **consistent** $h$, any node popped from `open` is final and will not improve later.* +Store the best known `g` when enqueuing a neighbor, but allow later improvements. Skip stale queue entries before expanding a node or accepting the goal. With a consistent heuristic, the first current entry popped for a node finalizes its distance. -**Reference pseudocode** +Reference pseudocode -``` -A_star(G, start, goal, h): - open = MinPQ() # keyed by f = g + h - open.push(start, h(start)) - g = {start: 0} +```text +A_star(graph, start, goal, heuristic): + frontier = MinPriorityQueue() + frontier.push((start, 0), heuristic(start)) + best_cost = {start: 0} parent = {start: None} closed = set() - while open: - u = open.pop_min() # node with smallest f - if u == goal: - return reconstruct_path(parent, goal), g[goal] - - closed.add(u) - - for (v, w_uv) in G.neighbors(u): # w_uv >= 0 - tentative = g[u] + w_uv - if v in closed and tentative >= g.get(v, +inf): - continue - - if tentative < g.get(v, +inf): - parent[v] = u - g[v] = tentative - f_v = tentative + h(v) - open.push(v, f_v) # decrease-key OR push new entry - - return None, +inf + while frontier: + (current, stored_cost), priority = frontier.pop_min() + if stored_cost != best_cost[current] or current in closed: + continue + if current == goal: + return reconstruct_path(parent, goal), stored_cost + + closed.add(current) + for neighbor, edge_cost in graph.neighbors(current): + candidate = stored_cost + edge_cost + if candidate < best_cost.get(neighbor, +infinity): + best_cost[neighbor] = candidate + parent[neighbor] = current + closed.discard(neighbor) + frontier.push( + (neighbor, candidate), + candidate + heuristic(neighbor) + ) + + return None, +infinity reconstruct_path(parent, t): path = [] @@ -1105,13 +1077,13 @@ reconstruct_path(parent, t): return list(reversed(path)) ``` -*Sanity notes:* +Sanity notes: -* The *time* complexity of A* is worst-case exponential, though in practice it runs much faster when the heuristic $h$ provides useful guidance; without an informative heuristic, the search can expand nearly the entire graph, as in navigating a large grid without directional hints. -* The *space* complexity is $O(V)$, covering the priority queue and bookkeeping maps, which makes A* memory-intensive; without recognizing this, applications such as robotics pathfinding may exceed available memory on large maps. -* In *special cases*, A* reduces to Dijkstra’s algorithm when $h \equiv 0$, and further reduces to BFS when all edges have cost 1 and $h \equiv 0$; without this perspective, one might overlook how A* generalizes these familiar shortest-path algorithms. +- On an explicit finite graph with a consistent heuristic, each vertex is expanded at most once; a binary heap gives bounds comparable to Dijkstra. For implicit state spaces, the number of states can be exponential in solution depth. Inconsistent heuristics can also cause repeated expansions, so these are different complexity models. +- With consistency and an indexed queue, auxiliary space is $O(V)$; lazy duplicate entries can raise it to $O(V+E)$. With reopening, pending entries depend on the number of improvements. Large implicit state spaces can make memory the limiting resource. +- In special cases, A* reduces to Dijkstra’s algorithm when $h \equiv 0$, and further reduces to BFS when all edges have cost 1 and $h \equiv 0$; without this perspective, one might overlook how A* generalizes these familiar shortest-path algorithms. -**Visual walkthrough (grid with 4-neighborhood, Manhattan $h$)** +Visual walkthrough (grid with 4-neighborhood, Manhattan $h$) Legend: `S` start, `G` goal, `#` wall, `.` free, `◉` expanded (closed), `•` frontier (open), `×` final path @@ -1127,7 +1099,7 @@ Row/Col → 1 2 3 4 5 6 7 8 9 Movement cost = 1 per step; 4-dir moves; h = Manhattan distance ``` -**Early expansion snapshot (conceptual):** +Early expansion snapshot (conceptual): ``` Step 0: @@ -1144,113 +1116,113 @@ A* keeps popping the lowest f, steering toward G. Nodes near the straight line to G are preferred over detours around '#'. ``` -**When goal is reached, reconstruct the path:** +When goal is reached, reconstruct the path: -``` -Final path (example rendering): -Row/Col → 1 2 3 4 5 6 7 8 9 - ┌─────────────────────────────┐ - 1 │ × × × × . # . . . │ - 2 │ × # # × × # . # . │ - 3 │ × × × × × × × # . │ - 4 │ # . # # × # × × × │ - 5 │ . . . # × × × # G │ - └─────────────────────────────┘ -Path length (g at G) equals number of × steps (optimal with admissible/consistent h). -``` +```text +One shortest path (row, column; one-based coordinates): +(1,1) -> (1,2) -> (1,3) -> (1,4) -> (2,4) -> (3,4) + -> (3,5) -> (3,6) -> (3,7) -> (4,7) -> (4,8) + -> (4,9) -> (5,9) -**Priority queue evolution (toy example)** +S × × × . # . . . +. # # × . # . # . +. . . × × × × # . +# . # # . # × × × +. . . # . . . # G +There are 13 path vertices and 12 moves, so g(G) = 12. ``` -Step | Popped u | Inserted neighbors (v: g,h,f) | Note ------+----------+-------------------------------------------------+--------------------------- -0 | , | push S: g=0, h=14, f=14 | S at (1,1), G at (5,9) -1 | S | (1,2): g=1,h=13,f=14 ; (2,1): g=1,h=12,f=13 | pick (2,1) next -2 | (2,1) | (3,1): g=2,h=11,f=13 ; (2,2) blocked | ... -3 | (3,1) | (4,1) wall; (3,2): g=3,h=10,f=13 | still f=13 band -… | … | frontier slides along the corridor toward G | A* hugs the beeline + +Priority queue evolution (toy example) + +```text +Step | Popped | New entries (g, h, f) | Tie choice +0 | — | S: (0, 12, 12) | S = (1,1) +1 | S | (1,2): (1,11,12), (2,1): (1,11,12) | choose (2,1) +2 | (2,1) | (3,1): (2,10,12) | choose (3,1) +3 | (3,1) | (3,2): (3,9,12); (4,1) is a wall | both branches remain possible ``` -(Exact numbers depend on the specific grid and walls; shown for intuition.) +These Manhattan values follow directly from the displayed coordinates. Equal priorities leave the expansion order dependent on tie-breaking. -**Heuristic design** +Heuristic design -For **grids**: +For grids: -* *4-dir moves:* $h(n)=|x_n-x_g|+|y_n-y_g|$ (Manhattan). -* *8-dir (diag cost √2):* **Octile**: $h=\Delta_{\max} + (\sqrt{2}-1)\Delta_{\min}$. -* *Euclidean* when motion is continuous and diagonal is allowed. +- 4-dir moves: $h(n)=|x_n-x_g|+|y_n-y_g|$ (Manhattan). +- 8-dir (diag cost √2): Octile: $h=\Delta_{\max} + (\sqrt{2}-1)\Delta_{\min}$. +- Euclidean when motion is continuous and diagonal is allowed. -For **sliding puzzles (e.g., 8/15-puzzle)**: +For sliding puzzles (e.g., 8/15-puzzle): -**Misplaced tiles* (admissible, weak). -* *Manhattan sum* (stronger). -* *Linear conflict / pattern databases* (even stronger). +- Misplaced tiles, excluding the blank (admissible, weak). +- Manhattan sum over numbered tiles, excluding the blank (stronger). +- Linear conflict / pattern databases (even stronger). -**Admissible vs. consistent** +Admissible vs. Consistent -* An *admissible* heuristic satisfies $h(n) \leq h^*(n)$, meaning it never overestimates the true remaining cost, which guarantees that A* finds an optimal path; without admissibility, the algorithm may return a suboptimal route, such as a longer-than-necessary driving path. -* A *consistent (monotone)* heuristic obeys $h(u) \leq w(u,v) + h(v)$ for every edge, ensuring that $f$-values do not decrease along paths and that once a node is removed from the open set, its $g$-value is final; without consistency, nodes may need to be reopened, increasing complexity in searches like grid navigation. +- An admissible heuristic satisfies $h(n) \leq h^*(n)$, meaning it never overestimates the true remaining cost, which guarantees optimality for the graph-search version only with an appropriate reopening policy and goal-pop termination; without admissibility, the algorithm may return a suboptimal route, such as a longer-than-necessary driving path. +- A consistent (monotone) heuristic obeys $h(u) \leq w(u,v) + h(v)$ for every edge, ensuring that $f$-values do not decrease along paths and that once a node is removed from the open set, its $g$-value is final; without consistency, nodes may need to be reopened, increasing complexity in searches like grid navigation. -**Applications** +Applications -* In *pathfinding* for maps, games, and robotics, A* computes shortest or least-risk routes by combining actual travel cost with heuristic guidance; without it, movement planning in virtual or physical environments becomes slower or less efficient. -* For *route planning* with road metrics such as travel time, distance, or tolls, A* incorporates these costs and constraints into its evaluation; without heuristic search, navigation systems must fall back to slower methods like plain Dijkstra. -* In *planning and scheduling* tasks, A* serves as a general shortest-path algorithm in abstract state spaces, supporting AI decision-making; without it, solving resource allocation or task sequencing problems may require less efficient exhaustive search. -* In *puzzle solving* domains such as the 8-puzzle or Sokoban, A* uses problem-specific heuristics to guide the search efficiently; without heuristics, the state space may grow exponentially and become impractical to explore. -* For *network optimization* problems with nonnegative edge costs, A* applies whenever a useful heuristic is available to speed convergence; without heuristics, computations on communication or logistics networks may take longer than necessary. +- In pathfinding for maps, games, and robotics, A* computes shortest or least-risk routes by combining actual travel cost with heuristic guidance; without it, movement planning in virtual or physical environments becomes slower or less efficient. +- For route planning with road metrics such as travel time, distance, or tolls, A* incorporates these costs and constraints into its evaluation; without heuristic search, navigation systems must fall back to slower methods like plain Dijkstra. +- In planning and scheduling tasks, A* serves as a general shortest-path algorithm in abstract state spaces, supporting AI decision-making; without it, solving resource allocation or task sequencing problems may require less efficient exhaustive search. +- In puzzle solving domains such as the 8-puzzle or Sokoban, A* uses problem-specific heuristics to guide the search efficiently; without heuristics, the state space may grow exponentially and become impractical to explore. +- For network optimization problems with nonnegative edge costs, A* applies whenever a useful heuristic is available to speed convergence; without heuristics, computations on communication or logistics networks may take longer than necessary. -**Variants & practical tweaks** +Variants & practical tweaks -* Viewing *Dijkstra* as A* with $h \equiv 0$ shows that A* generalizes the classic shortest-path algorithm; without this equivalence, the connection between uninformed and heuristic search may be overlooked. -* In *Weighted A**, the evaluation function becomes $f = g + \varepsilon h$ with $\varepsilon > 1$, trading exact optimality for faster performance with bounded suboptimality; without this variant, applications needing quick approximate routing, like logistics planning, would run slower. -* The *A*ε / Anytime A** approach begins with $\varepsilon > 1$ for speed and gradually reduces it to converge toward optimal paths; without this strategy, incremental refinement in real-time systems like navigation aids is harder to achieve. -* With *IDA** (Iterative Deepening A*), the search is conducted by gradually increasing an $f$-cost threshold, greatly reducing memory usage but sometimes increasing runtime; without it, problems like puzzle solving could exceed memory limits. -* *RBFS and Fringe Search* are memory-bounded alternatives that manage recursion depth or fringe sets more carefully; without these, large state spaces in AI planning can overwhelm storage. -* In *tie-breaking*, preferring larger $g$ or smaller $h$ when $f$ ties reduces unnecessary re-expansions; without careful tie-breaking, searches on uniform-cost grids may explore more nodes than needed. -* For the *closed-set policy*, when heuristics are inconsistent, nodes must be reopened if a better $g$ value is found; without allowing this, the algorithm may miss shorter paths, as in road networks with varying travel times. +- Viewing Dijkstra as A* with $h \equiv 0$ shows that A* generalizes the classic shortest-path algorithm; without this equivalence, the connection between uninformed and heuristic search may be overlooked. +- In Weighted A\*, the evaluation function becomes $f = g + \varepsilon h$ with $\varepsilon > 1$, trading exact optimality for faster performance with bounded suboptimality; without this variant, applications needing quick approximate routing, like logistics planning, would run slower. +- An anytime A\* approach begins with $\varepsilon > 1$ for speed and gradually reduces it to converge toward optimal paths; without this strategy, incremental refinement in real-time systems like navigation aids is harder to achieve. +- With IDA\* (Iterative Deepening A\*), the search is conducted by gradually increasing an $f$-cost threshold, greatly reducing memory usage but sometimes increasing runtime; without it, problems like puzzle solving could exceed memory limits. +- RBFS and Fringe Search are memory-bounded alternatives that manage recursion depth or fringe sets more carefully; without these, large state spaces in AI planning can overwhelm storage. +- In tie-breaking, preferring larger $g$ or smaller $h$ when $f$ ties reduces unnecessary re-expansions; without careful tie-breaking, searches on uniform-cost grids may explore more nodes than needed. +- For the closed-set policy, when heuristics are inconsistent, nodes must be reopened if a better $g$ value is found; without allowing this, the algorithm may miss shorter paths, as in road networks with varying travel times. -**Pitfalls & tips** +Pitfalls & tips -* The algorithm requires *non-negative edge weights* because A* assumes $w(u,v) \ge 0$; without this, negative costs can cause nodes to be expanded too early, breaking correctness in applications like navigation. -* If the heuristic *overestimates* actual costs, A* loses its guarantee of optimality; without enforcing admissibility, a routing system may return a path that is faster to compute but longer in distance. -* With *floating-point precision issues*, comparisons of $f$-values should include small epsilons to avoid instability; without this safeguard, two nearly equal paths may lead to inconsistent queue ordering in large-scale searches. -* In *state hashing*, equivalent states must hash identically so duplicates are merged properly; without this, search in puzzles or planning domains may blow up due to treating the same state as multiple distinct ones. -* While *neighbor order* does not affect correctness, it influences performance and the aesthetics of the returned path trace; without considering this, two identical problems might yield very different expansion sequences or outputs. +- The algorithm requires non-negative edge weights because A* assumes $w(u,v) \ge 0$; without this, negative costs can cause nodes to be expanded too early, breaking correctness in applications like navigation. +- If the heuristic overestimates actual costs, A* loses its guarantee of optimality; without enforcing admissibility, a routing system may return a path that is faster to compute but longer in distance. +- For floating-point weights, keep a consistent total ordering in the priority queue, for example by adding a unique tie-break counter. An epsilon-based heap comparator can violate transitivity. Relaxation tolerances may discard small improvements and must be treated as an approximation policy, not an automatic correctness fix. +- In state hashing, equivalent states must hash identically so duplicates are merged properly; without this, search in puzzles or planning domains may blow up due to treating the same state as multiple distinct ones. +- While neighbor order does not affect correctness, it influences performance and the aesthetics of the returned path trace; without considering this, two identical problems might yield very different expansion sequences or outputs. -**Implementation** +Implementation: -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/graphs/cpp/a_star) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/graphs/python/a_star) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/graphs/cpp/a_star) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/graphs/python/a_star) -*Implementation tip:* If your PQ lacks decrease-key, **push duplicates** with improved keys and ignore stale entries when popped (check if popped `g` matches current `g[u]`). This is simple and fast in practice. +Implementation tip: If your PQ lacks decrease-key, push duplicates with improved keys and ignore stale entries when popped (check if popped `g` matches current `g[u]`). This is simple and fast in practice. -### Minimal Spanning Trees +## Minimum Spanning Trees Suppose we have a graph that represents a network of houses. Weights represent the distances between vertices, which each represent a single house. All houses must have water, electricity, and internet, but we want the cost of installation to be as low as possible. We need to identify a subgraph of our graph with the following properties: -* There are no cycles in the graph. -* All vertices are connected. -* The total sum of weights is minimum. +- There are no cycles in the graph. +- All vertices are connected. +- The total sum of weights is minimum. -Such a subgraph is called a minimal spanning tree. +Such a subgraph is called a minimum spanning tree (MST). It minimizes total edge weight, not the distance from a chosen source to every other vertex. A disconnected graph has a minimum spanning forest, with one tree for each connected component. -#### Prim’s Algorithm +### Prim’s Algorithm -Prim’s algorithm builds a **minimum spanning tree (MST)** of a **weighted, undirected** graph by growing a tree from a start vertex. At each step it adds the **cheapest edge** that connects a vertex **inside** the tree to a vertex **outside** the tree. +Prim’s algorithm builds a minimum spanning tree (MST) of a weighted, undirected graph by growing a tree from a start vertex. At each step it adds the cheapest edge that connects a vertex inside the tree to a vertex outside the tree. To efficiently keep track of the construction, Prim’s algorithm employs two primary data structures: -* A **min-priority queue** (often named `pq`, `open`, or `unexplored`) keyed by a vertex’s **best known connection cost** to the current tree. -* A **`in_mst`/`visited` set** to mark vertices already added to the tree, plus a **`parent` map** to record the chosen incoming edge for each vertex. +- A min-priority queue (often named `pq`, `open`, or `unexplored`) keyed by a vertex’s best known connection cost to the current tree. +- A `in_mst`/`visited` set to mark vertices already added to the tree, plus a `parent` map to record the chosen incoming edge for each vertex. -*Useful additions in practice:* +Useful additions in practice: -* A *key map* stores, for each vertex, the lightest edge weight connecting it to the current spanning tree, initialized to infinity except for the starting vertex at zero; without this, Prim’s algorithm cannot efficiently track which edges should be added next to grow the tree. -* With *lazy updates*, when the priority queue lacks a decrease-key operation, improved entries are simply pushed again and outdated ones are skipped upon popping; without this adjustment, priority queues become harder to manage, slowing down minimum spanning tree construction. -* For *component handling*, if the graph is disconnected, Prim’s algorithm must either restart from each unvisited vertex or seed multiple starts with key values of zero to produce a spanning forest; without this, the algorithm would stop after one component, leaving parts of the graph unspanned. +- A key map stores, for each vertex, the lightest edge weight connecting it to the current spanning tree, initialized to infinity except for the starting vertex at zero; without this, Prim’s algorithm cannot efficiently track which edges should be added next to grow the tree. +- With lazy updates, when the priority queue lacks a decrease-key operation, improved entries are simply pushed again and outdated ones are skipped upon popping; without this adjustment, priority queues become harder to manage, slowing down minimum spanning tree construction. +- For component handling, if the graph is disconnected, Prim’s algorithm must either restart from each unvisited vertex or seed exactly one representative per connected component to produce a spanning forest; without this, the algorithm would stop after one component, leaving parts of the graph unspanned. -**Algorithm Steps** +Algorithm Steps: 1. Pick a start vertex $i$. 2. Set `key[i] = 0`, `parent[i] = None`; for all other vertices $v \ne i$, set `key[v] = ∞`; push $i$ into a min-priority queue keyed by `key`. @@ -1261,9 +1233,9 @@ To efficiently keep track of the construction, Prim’s algorithm employs two pr 7. Stop when the queue is empty or when all vertices are in the MST (for a connected graph). 8. The edges $\{(parent[v], v) : v \ne i\}$ form an MST; the MST total weight equals $\sum key[v]$ at the moments when each $v$ is added. -Vertices are **finalized when they are dequeued**: at that moment, `key[u]` is the **minimum** cost to connect `u` to the growing tree (by the **cut property**). +Vertices are finalized when they are dequeued: at that moment, `key[u]` is the minimum cost to connect `u` to the growing tree (by the cut property). -*Reference pseudocode (adjacency-list graph):* +Reference pseudocode (adjacency-list graph): ``` Prim(G, i): @@ -1296,16 +1268,16 @@ Prim(G, i): return mst_edges, parent, sum(w for (_,_,w) in mst_edges) ``` -*Sanity notes:* +Sanity notes: -* The *time* complexity of Prim’s algorithm is $O(E \log V)$ with a binary heap, $O(E + V \log V)$ with a Fibonacci heap, and $O(V^2)$ for the dense-graph adjacency-matrix variant; without knowing this, one might apply the wrong implementation and get poor performance on sparse or dense networks. -* The *space* complexity is $O(V)$, required for storing the key values, parent pointers, and bookkeeping to build the minimum spanning tree; without this allocation, the algorithm cannot track which edges belong to the MST. -* The *graph type* handled is a weighted, undirected graph with no restrictions on edge weights being positive; without this flexibility, graphs with negative costs, such as energy-saving transitions, could not be processed. -* In terms of *uniqueness*, if all edge weights are distinct, the minimum spanning tree is unique; without distinct weights, multiple MSTs may exist, such as in networks where two equally light connections are available. +- The time complexity of Prim’s algorithm is $O(E \log V)$ with a binary heap, $O(E + V \log V)$ with a Fibonacci heap, and $O(V^2)$ for the dense-graph adjacency-matrix variant; without knowing this, one might apply the wrong implementation and get poor performance on sparse or dense networks. +- An indexed queue requires $O(V)$ auxiliary space. The lazy duplicate-entry version shown can require $O(V+E)$ space and $O(V+E\log(E+2))$ time. On a connected simple graph, its time simplifies to $O(E\log V)$. +- The graph type handled is a weighted, undirected graph with no restrictions on edge weights being positive; without this flexibility, graphs with negative costs, such as energy-saving transitions, could not be processed. +- In terms of uniqueness, if all edge weights are distinct, the minimum spanning tree is unique; without distinct weights, multiple MSTs may exist, such as in networks where two equally light connections are available. -**Example** +Example -Undirected, weighted graph; start at **A**. Edge weights shown on links. +Undirected, weighted graph; start at A. Edge weights shown on links. ``` # @@ -1328,7 +1300,7 @@ Undirected, weighted graph; start at **A**. Edge weights shown on links. Edges: A–B(4), A–C(1), C–B(2), B–E(1), C–D(4), D–E(3) ``` -*Frontier (keys) / In-tree evolution (min at front):* +Frontier (keys) / In-tree evolution (min at front): ``` Legend: key[v] = cheapest known connection to tree; parent[v] = chosen neighbor @@ -1343,14 +1315,14 @@ Step | Action | PQ (key:vertex) after push | In 5 | pop D(3) → add | [4:D, 4:B] | {A,C,B,E,D} | done ``` -*MST edges chosen (with weights):* +MST edges chosen (with weights): ``` -A, C(1), C, B(2), B, E(1), E, D(3) +A–C(1), C–B(2), B–E(1), E–D(3) Total weight = 1 + 2 + 1 + 3 = 7 ``` -*Resulting MST (tree edges only):* +Resulting MST (tree edges only): ``` A @@ -1360,40 +1332,40 @@ A └── D (3) ``` -**Applications** +Applications -* In *network design*, Prim’s or Kruskal’s MST construction connects all sites such as offices, cities, or data centers with the least total cost of wiring, piping, or fiber; without using MSTs, infrastructure plans risk including redundant and more expensive links. -* As an *approximation for the traveling salesman problem (TSP)*, building an MST and performing a preorder walk of it yields a tour within twice the optimal length for metric TSP; without this approach, even approximate solutions for large instances may be much harder to obtain. -* In *clustering with single linkage*, removing the $k-1$ heaviest edges of the MST partitions the graph into $k$ clusters; without this technique, hierarchical clustering may require recomputing pairwise distances repeatedly. -* For *image processing and segmentation*, constructing an MST over pixels or superpixels highlights low-contrast boundaries as cut edges; without MST-based grouping, segmentations may fail to respect natural intensity or color edges. -* In *map generalization and simplification*, the MST preserves a connectivity backbone with minimal redundancy, reducing complexity while maintaining essential routes; without this, simplified maps may show excessive or unnecessary detail. -* In *circuit design and VLSI*, MSTs minimize interconnect length under simple wiring models, supporting efficient layouts; without this method, chip designs may consume more area and power due to avoidable wiring overhead. +- In network design, Prim’s or Kruskal’s MST construction connects all sites such as offices, cities, or data centers with the least total cost of wiring, piping, or fiber; without using MSTs, infrastructure plans risk including redundant and more expensive links. +- As an approximation for the traveling salesman problem (TSP), building an MST and performing a preorder walk of it yields a tour within twice the optimal length for metric TSP; without this approach, even approximate solutions for large instances may be much harder to obtain. +- In clustering with single linkage, removing the $k-1$ heaviest edges of the MST partitions the graph into $k$ clusters; without this technique, hierarchical clustering may require recomputing pairwise distances repeatedly. +- For image processing and segmentation, constructing an MST over pixels or superpixels connects similar regions with low-weight edges, while cutting selected high-weight edges can separate contrasting regions; without MST-based grouping, segmentations may fail to respect natural intensity or color edges. +- In map generalization and simplification, the MST preserves a connectivity backbone with minimal redundancy, reducing complexity while maintaining essential routes; without this, simplified maps may show excessive or unnecessary detail. +- In circuit design and VLSI, MSTs minimize interconnect length under simple wiring models, supporting efficient layouts; without this method, chip designs may consume more area and power due to avoidable wiring overhead. -##### Implementation +Implementation: -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/graphs/cpp/prim) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/graphs/python/prim) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/graphs/cpp/prim) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/graphs/python/prim) -*Implementation tip:* -For **dense graphs** ($E \approx V^2$), skip heaps: store `key` in an array and, at each step, scan all non-MST vertices to pick the minimum `key` in $O(V)$. Overall $O(V^2)$ but often **faster in practice** on dense inputs due to low overhead. +Implementation tip: +For dense graphs ($E \approx V^2$), skip heaps: store `key` in an array and, at each step, scan all non-MST vertices to pick the minimum `key` in $O(V)$. Overall $O(V^2)$ but often faster in practice on dense inputs due to low overhead. -#### Kruskal’s Algorithm +### Kruskal’s Algorithm -Kruskal’s algorithm builds a **minimum spanning tree (MST)** for a **weighted, undirected** graph by sorting all edges by weight (lightest first) and repeatedly adding the next lightest edge that **does not create a cycle**. It grows the MST as a forest of trees that gradually merges until all vertices are connected. +Kruskal’s algorithm builds a minimum spanning tree (MST) for a weighted, undirected graph by sorting all edges by weight (lightest first) and repeatedly adding the next lightest edge that does not create a cycle. It grows the MST as a forest of trees that gradually merges until all vertices are connected. To efficiently keep track of the construction, Kruskal’s algorithm employs two primary data structures: -* A *sorted edge list* arranged in ascending order of weights ensures that Kruskal’s algorithm always considers the lightest available edge next; without this ordering, the method cannot guarantee that the resulting spanning tree has minimum total weight. -* A *Disjoint Set Union (DSU)*, or Union–Find structure, tracks which vertices belong to the same tree and prevents cycles by only uniting edges from different sets; without this mechanism, the algorithm could inadvertently form cycles instead of building a spanning tree. +- A sorted edge list arranged in ascending order of weights ensures that Kruskal’s algorithm always considers the lightest available edge next; without this ordering, the method cannot guarantee that the resulting spanning tree has minimum total weight. +- A Disjoint Set Union (DSU), or Union–Find structure, tracks which vertices belong to the same tree and prevents cycles by only uniting edges from different sets; without this mechanism, the algorithm could inadvertently form cycles instead of building a spanning tree. -*Useful additions in practice:* +Useful additions in practice: -* Using *Union–Find with path compression and union by rank/size* enables near-constant-time merge and find operations, making Kruskal’s algorithm efficient; without these optimizations, edge processing in large graphs such as communication networks would slow down significantly. -* Applying an *early stop* allows the algorithm to terminate once $V-1$ edges have been added in a connected graph, since the MST is then complete; without this, unnecessary edges are still considered, adding avoidable work. -* Enforcing *deterministic tie-breaking* ensures that when multiple edges share equal weights, the same MST is consistently produced; without this, repeated runs on the same weighted graph might yield different but equally valid spanning trees, complicating reproducibility. -* On *disconnected graphs*, Kruskal’s algorithm naturally outputs a minimum spanning forest with one tree per component; without this property, handling graphs such as multiple separate road systems would require additional adjustments. +- Using Union–Find with path compression and union by rank/size enables near-constant-time merge and find operations, making Kruskal’s algorithm efficient; without these optimizations, edge processing in large graphs such as communication networks would slow down significantly. +- Applying an early stop allows the algorithm to terminate once $V-1$ edges have been added in a connected graph, since the MST is then complete; without this, unnecessary edges are still considered, adding avoidable work. +- Enforcing deterministic tie-breaking ensures that when multiple edges share equal weights, the same MST is consistently produced; without this, repeated runs on the same weighted graph might yield different but equally valid spanning trees, complicating reproducibility. +- On disconnected graphs, Kruskal’s algorithm naturally outputs a minimum spanning forest with one tree per component; without this property, handling graphs such as multiple separate road systems would require additional adjustments. -**Algorithm Steps** +Algorithm Steps: 1. Gather all edges $E=\{(u,v,w)\}$ and sort them by weight $w$ in ascending order. 2. Initialize a DSU with each vertex in its own set: `parent[v]=v`, `rank[v]=0`. @@ -1404,9 +1376,9 @@ To efficiently keep track of the construction, Kruskal’s algorithm employs two 7. Continue until $V-1$ edges have been chosen (connected graph) or until all edges are processed (forest). 8. The chosen edges form the MST; the total weight is the sum of their weights. -By the **cycle** and **cut** properties of MSTs, selecting the minimum-weight edge that crosses any cut between components is always safe; rejecting edges that close a cycle preserves optimality. +By the cycle and cut properties of MSTs, selecting the minimum-weight edge that crosses any cut between components is always safe; rejecting edges that close a cycle preserves optimality. -*Reference pseudocode (edge list + DSU):* +Reference pseudocode (edge list + DSU): ``` Kruskal(V, E): @@ -1448,14 +1420,14 @@ union(x, y): rank[rx] += 1 ``` -*Sanity notes:* +Sanity notes: -* The *time* complexity of Kruskal’s algorithm is dominated by sorting edges, which takes $O(E \log E)$, or equivalently $O(E \log V)$, while DSU operations run in near-constant amortized time; without recognizing this, one might wrongly attribute the main cost to the union–find structure rather than sorting. -* The *space* complexity is $O(V)$ for the DSU arrays and $O(E)$ to store the edges; without this allocation, the algorithm cannot track connectivity or efficiently access candidate edges. -* With respect to *weights*, Kruskal’s algorithm works on undirected graphs with either negative or positive weights; without this flexibility, cases like networks where some connections represent cost reductions could not be handled. -* Regarding *uniqueness*, if all edge weights are distinct, the MST is guaranteed to be unique; without distinct weights, multiple equally valid minimum spanning trees may exist, such as in graphs where two different links have identical costs. +- The time complexity of Kruskal’s algorithm is dominated by sorting edges, which takes $O(E \log E)$, which is $O(E\log V)$ for simple graphs, while DSU operations run in near-constant amortized time; without recognizing this, one might wrongly attribute the main cost to the union–find structure rather than sorting. +- The space complexity is $O(V)$ for the DSU arrays and $O(E)$ to store the edges; without this allocation, the algorithm cannot track connectivity or efficiently access candidate edges. +- With respect to weights, Kruskal’s algorithm works on undirected graphs with either negative or positive weights; without this flexibility, cases like networks where some connections represent cost reductions could not be handled. +- Regarding uniqueness, if all edge weights are distinct, the MST is guaranteed to be unique; without distinct weights, multiple equally valid minimum spanning trees may exist, such as in graphs where two different links have identical costs. -**Example** +Example Undirected, weighted graph (we’ll draw the key edges clearly and list the rest). Start with all vertices as separate sets: `{A} {B} {C} {D} {E} {F}`. @@ -1472,10 +1444,10 @@ Other edges (not all drawn to keep the picture clean): A–C(4), B–D(5), C–E(5), D–E(6), D–F(2) ``` -*Sorted edge list (ascending):* +Sorted edge list (ascending): `E–F(1), B–C(2), D–F(2), B–E(3), A–B(4), A–C(4), B–D(5), C–D(5), C–E(5), D–E(6), A–F(7)` -*Union–Find / MST evolution (take the edge if it connects different sets):* +Union–Find / MST evolution (take the edge if it connects different sets): ``` Step | Edge (w) | Find(u), Find(v) | Action | Components after union | MST so far | Total @@ -1488,13 +1460,13 @@ Step | Edge (w) | Find(u), Find(v) | Action | Components after union | (stop: we have V−1 = 5 edges for 6 vertices) ``` -*Resulting MST edges and weight:* +Resulting MST edges and weight: ``` E–F(1), B–C(2), D–F(2), B–E(3), A–B(4) ⇒ Total = 1 + 2 + 2 + 3 + 4 = 12 ``` -*Clean MST view (tree edges only):* +Clean MST view (tree edges only): ``` A @@ -1505,40 +1477,40 @@ A └── D (2) ``` -**Applications** +Applications -* In *network design*, Kruskal’s algorithm builds the least-cost backbone, such as roads, fiber, or pipelines, that connects all sites with minimal total expense; without MST construction, the resulting infrastructure may include redundant and costlier links. -* For *clustering with single linkage*, constructing the MST and then removing the $k-1$ heaviest edges partitions the graph into $k$ clusters; without this method, grouping data points into clusters may require repeated and slower distance recalculations. -* In *image segmentation*, applying Kruskal’s algorithm to pixel or superpixel graphs groups regions by intensity or feature similarity through MST formation; without MST-based grouping, boundaries between regions may be less well aligned with natural contrasts. -* As an *approximation for the metric traveling salesman problem*, building an MST and performing a preorder walk (with shortcutting) yields a tour at most twice the optimal length; without this approach, near-optimal solutions would be harder to compute efficiently. -* In *circuit and VLSI layout*, Kruskal’s algorithm finds minimal interconnect length under simplified wiring models; without this, designs may require more area and energy due to unnecessarily long connections. -* For *maze generation*, a randomized Kruskal process selects edges in random order while maintaining acyclicity, producing mazes that remain connected without loops; without this structure, generated mazes could contain cycles or disconnected regions. +- In network design, Kruskal’s algorithm builds the least-cost backbone, such as roads, fiber, or pipelines, that connects all sites with minimal total expense; without MST construction, the resulting infrastructure may include redundant and costlier links. +- For clustering with single linkage, constructing the MST and then removing the $k-1$ heaviest edges partitions the graph into $k$ clusters; without this method, grouping data points into clusters may require repeated and slower distance recalculations. +- In image segmentation, applying Kruskal’s algorithm to pixel or superpixel graphs groups regions by intensity or feature similarity through MST formation; without MST-based grouping, boundaries between regions may be less well aligned with natural contrasts. +- As an approximation for the metric traveling salesman problem, building an MST and performing a preorder walk (with shortcutting) yields a tour at most twice the optimal length; without this approach, near-optimal solutions would be harder to compute efficiently. +- In circuit and VLSI layout, Kruskal’s algorithm finds minimal interconnect length under simplified wiring models; without this, designs may require more area and energy due to unnecessarily long connections. +- For maze generation, a randomized Kruskal process selects edges in random order while maintaining acyclicity, producing mazes that remain connected without loops; without this structure, generated mazes could contain cycles or disconnected regions. -**Implementation** +Implementation: -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/graphs/cpp/kruskal) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/graphs/python/kruskal) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/graphs/cpp/kruskal) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/graphs/python/kruskal) -*Implementation tip:* -On huge graphs that **stream from disk**, you can **external-sort** edges by weight, then perform a single pass with DSU. For reproducibility across platforms, **stabilize** sorting by `(weight, min(u,v), max(u,v))`. +Implementation tip: +On huge graphs that stream from disk, you can external-sort edges by weight, then perform a single pass with DSU. For reproducibility across platforms, stabilize sorting by `(weight, min(u,v), max(u,v))`. -### Topological Sort +## Topological Sort -Topological sort orders the vertices of a **directed acyclic graph (DAG)** so that **every directed edge** $u \rightarrow v$ goes **from left to right** in the order (i.e., $u$ appears before $v$). It’s the canonical tool for scheduling tasks with dependencies. +Topological sort orders the vertices of a directed acyclic graph (DAG) so that every directed edge $u \rightarrow v$ goes from left to right in the order (i.e., $u$ appears before $v$). It’s the canonical tool for scheduling tasks with dependencies. To efficiently keep track of the process (Kahn’s algorithm), we use: -* A **queue** (or min-heap if you want lexicographically smallest order) holding all vertices with **indegree = 0** (no unmet prerequisites). -* An **`indegree` map/array** that counts for each vertex how many prerequisites remain. -* An **`order` list** to append vertices as they are “emitted.” +- A queue (or min-heap if you want lexicographically smallest order) holding all vertices with indegree = 0 (no unmet prerequisites). +- An `indegree` map/array that counts for each vertex how many prerequisites remain. +- An `order` list to append vertices as they are “emitted.” -*Useful additions in practice:* +Useful additions in practice: -* Maintaining a *visited count* or tracking the length of the output order lets you detect cycles, since producing fewer than $V$ vertices indicates that some could not be placed due to a cycle; without this check, algorithms like Kahn’s may silently return incomplete results on cyclic task graphs. -* Using a *min-heap* instead of a simple FIFO queue ensures that, among available candidates, the smallest-indexed vertex is always chosen, yielding the lexicographically smallest valid topological order; without this modification, the output order depends on arbitrary queueing, which may vary between runs. -* A *DFS-based alternative* computes a valid topological order by recording vertices in reverse postorder, also in $O(V+E)$ time, while detecting cycles via a three-color marking or recursion stack; without DFS, cycle detection must be handled separately in Kahn’s algorithm. +- Maintaining a visited count or tracking the length of the output order lets you detect cycles, since producing fewer than $V$ vertices indicates that some could not be placed due to a cycle; without this check, algorithms like Kahn’s may silently return incomplete results on cyclic task graphs. +- Using a min-heap instead of a simple FIFO queue ensures that, among available candidates, the smallest-indexed vertex is always chosen, yielding the lexicographically smallest valid topological order; without this modification, the output order depends on arbitrary queueing, which may vary between runs. +- A DFS-based alternative computes a valid topological order by recording vertices in reverse postorder, also in $O(V+E)$ time, while detecting cycles via a three-color marking or recursion stack; without DFS, cycle detection must be handled separately in Kahn’s algorithm. -**Algorithm Steps (Kahn’s algorithm)** +Algorithm Steps (Kahn’s algorithm) 1. Compute `indegree[v]` for every vertex $v$; set `order = []`. 2. Initialize a queue `Q` with all vertices of indegree 0. @@ -1548,7 +1520,7 @@ To efficiently keep track of the process (Kahn’s algorithm), we use: 6. If `indegree[v]` becomes 0, enqueue `v` into `Q`. 7. If `len(order) < V` at the end, a cycle exists and no topological order; otherwise `order` is a valid topological ordering. -*Reference pseudocode (adjacency-list graph):* +Reference pseudocode (adjacency-list graph): ``` TopoSort_Kahn(G): @@ -1579,48 +1551,35 @@ TopoSort_Kahn(G): return order ``` -*Sanity notes:* +Sanity notes: -* The *time* complexity of topological sorting is $O(V+E)$ because each vertex is enqueued exactly once and every edge is processed once when its indegree decreases; without this efficiency, ordering tasks in large dependency graphs would be slower. -* The *space* complexity is $O(V)$, required for storing indegree counts, the processing queue, and the final output order; without allocating this space, the algorithm cannot track which vertices are ready to be placed. -* The required *input* is a directed acyclic graph (DAG), since if a cycle exists, no valid topological order is possible; without this restriction, attempts to schedule cyclic dependencies, such as tasks that mutually depend on each other, will fail. +- With a FIFO queue, time is $O(V+E)$. The min-heap variant used in the example gives the lexicographically smallest order in $O(E+V\log V)$ time. Both require $O(V)$ auxiliary storage. +- The space complexity is $O(V)$, required for storing indegree counts, the processing queue, and the final output order; without allocating this space, the algorithm cannot track which vertices are ready to be placed. +- The required input is a directed acyclic graph (DAG), since if a cycle exists, no valid topological order is possible; without this restriction, attempts to schedule cyclic dependencies, such as tasks that mutually depend on each other, will fail. -**Example** +Example DAG; we’ll start with all indegree-0 vertices. (Edges shown as arrows.) -``` -DAG - ┌───────┐ - │ A │ - └───┬───┘ - │ - │ - ┌───────┐ ┌───▼───┐ ┌───────┐ - │ B │──────────│ C │──────────│ D │ - └───┬───┘ └───┬───┘ └───┬───┘ - │ │ │ - │ │ │ - │ ┌───▼───┐ │ - │ │ E │──────────────┘ - │ └───┬───┘ - │ │ - │ │ - ┌───▼───┐ ┌───▼───┐ - │ G │ │ F │ - └───────┘ └───────┘ +```text +A ----> C ----> D + ^ \ ^ + | v | +B ------+ E --+ +| +v +G F (isolated) -Edges: -A→C, B→C, C→D, C→E, E→D, B→G +Edges: A→C, B→C, C→D, C→E, E→D, B→G ``` -*Initial indegrees:* +Initial indegrees: ``` indeg[A]=0, indeg[B]=0, indeg[C]=2, indeg[D]=2, indeg[E]=1, indeg[F]=0, indeg[G]=1 ``` -*Queue/Indegree evolution (front → back; assume we keep the queue **lexicographically** by using a min-heap):* +Queue/Indegree evolution (front → back; assume we keep the queue lexicographically by using a min-heap): ``` Step | Pop u | Emit order | Decrease indeg[...] | Newly 0 → Enqueue | Q after @@ -1635,41 +1594,37 @@ Step | Pop u | Emit order | Decrease indeg[...] | Newly 0 → Enque 7 | G | [A, B, C, E, D, F, G] | , | , | [] ``` -*A valid topological order:* +A valid topological order: `A, B, C, E, D, F, G` (others like `B, A, C, E, D, F, G` are also valid.) -**Cycle detection (why it fails on cycles)** +Cycle detection (why it fails on cycles) -If there’s a cycle, some vertices **never** reach indegree 0. Example: +If there’s a cycle, some vertices never reach indegree 0. Example: -``` -# -┌─────┐ ┌─────┐ -│ X │ ───► │ Y │ -└──┬──┘ └──┬──┘ - └───────────►┘ - (Y ───► X creates a cycle) +```text +X --> Y +X <-- Y ``` -Here `indeg[X]=indeg[Y]=1` initially; `Q` starts empty ⇒ `order=[]` and `len(order) < V` ⇒ **cycle reported**. +Here `indeg[X]=indeg[Y]=1` initially; `Q` starts empty ⇒ `order=[]` and `len(order) < V` ⇒ cycle reported. -**Applications** +Applications -* In *build systems and compilation*, topological sorting ensures that each file is compiled only after its prerequisites are compiled; without it, a build may fail by trying to compile a module before its dependencies are available. -* For *course scheduling*, topological order provides a valid sequence in which to take courses respecting prerequisite constraints; without it, students may be assigned courses they are not yet eligible to take. -* In *data pipelines and DAG workflows* such as Airflow or Spark, tasks are executed when their inputs are ready by following a topological order; without this, pipeline stages might run prematurely and fail due to missing inputs. -* For *dependency resolution* in package managers or container systems, topological sorting installs components in an order that respects their dependencies; without it, software may be installed in the wrong sequence and break. -* In *dynamic programming on DAGs*, problems like longest path, shortest path, or path counting are solved efficiently by processing vertices in topological order; without this ordering, subproblems may be computed before their dependencies are solved. -* For *circuit evaluation or spreadsheets*, topological order ensures that each cell or net is evaluated only after its referenced inputs; without it, computations could use undefined or incomplete values. +- In build systems and compilation, topological sorting ensures that each file is compiled only after its prerequisites are compiled; without it, a build may fail by trying to compile a module before its dependencies are available. +- For course scheduling, topological order provides a valid sequence in which to take courses respecting prerequisite constraints; without it, students may be assigned courses they are not yet eligible to take. +- In data pipelines and DAG workflows such as Airflow or Spark, tasks are executed when their inputs are ready by following a topological order; without this, pipeline stages might run prematurely and fail due to missing inputs. +- For dependency resolution in package managers or container systems, topological sorting installs components in an order that respects their dependencies; without it, software may be installed in the wrong sequence and break. +- In dynamic programming on DAGs, problems like longest path, shortest path, or path counting are solved efficiently by processing vertices in topological order; without this ordering, subproblems may be computed before their dependencies are solved. +- For circuit evaluation or spreadsheets, topological order ensures that each cell or net is evaluated only after its referenced inputs; without it, computations could use undefined or incomplete values. -**Implementation** +Related exercises enumerate topological orderings using backtracking; the Kahn procedure above computes one ordering directly. -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/graphs/cpp/topological_sort) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/graphs/topological_sort/kruskal) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/backtracking/cpp/topological_sort) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/tree/master/src/backtracking/python/topological_sort) -*Implementation tips:* +Implementation tips: -* Use a **deque** for FIFO behavior; use a **min-heap** to get the **lexicographically smallest** topological order. -* When the graph is large and sparse, store adjacency as **lists** and compute indegrees in one pass for $O(V+E)$. -* **DFS variant** (brief): color states `0=unseen,1=visiting,2=done`; on exploring `u`, mark `1`; DFS to neighbors; if you see `1` again, there’s a cycle; on finish, push `u` to a stack. Reverse the stack for the order. +- Use a deque for FIFO behavior; use a min-heap to get the lexicographically smallest topological order. +- When the graph is large and sparse, store adjacency as lists and compute indegrees in one pass for $O(V+E)$. +- DFS variant (brief): color states `0=unseen,1=visiting,2=done`; on exploring `u`, mark `1`; DFS to neighbors; if you see `1` again, there’s a cycle; on finish, push `u` to a stack. Reverse the stack for the order. diff --git a/notes/greedy_algorithms.md b/notes/greedy_algorithms.md index fc53648..ec4348e 100644 --- a/notes/greedy_algorithms.md +++ b/notes/greedy_algorithms.md @@ -1,96 +1,100 @@ -## Greedy algorithms +# Greedy Algorithms -Greedy algorithms are the “make progress now” strategy: build a solution one step at a time, and at each step take the option that looks best *right now* according to a simple rule (highest value, earliest finish, smallest weight, smallest distance label, etc.). You keep the choice only if it doesn’t break the problem’s rules. +Greedy algorithms are the “make progress now” strategy: build a solution one step at a time, and at each step take the option that looks best right now according to a simple rule (highest value, earliest finish, smallest weight, smallest distance label, etc.). You keep the choice only if it doesn’t break the problem’s rules. -The story you should keep in your head is **greedy is fast because it refuses to look ahead**, but that means it earns its correctness only when you can prove that local choices are *safe*. The fun part is that the *proof patterns* repeat across problems, so once you learn the toolkit, new greedy problems stop feeling like magic tricks and start feeling like “apply the template.” +Greedy methods avoid exploring alternative combinations after each choice. That can make them efficient, but correctness requires proving that each commitment preserves an optimal solution. Exchange arguments and loop invariants provide reusable ways to establish this. -**The general pattern:** +A common pattern for selection problems is: 1. Sort by your rule (the “key”). 2. Scan items in that order. 3. If adding this item keeps the partial answer valid, keep it. 4. Otherwise skip it. -Picking the best “now” doesn’t obviously give the best “overall.” **Greedy isn’t magic; you must prove the choice is safe.** The real work is showing that these local choices still lead to a globally best answer. +Other greedy algorithms repeatedly update a priority queue or a running frontier instead of sorting once. The common feature is making a locally justified commitment without backtracking. + +A locally best option need not produce a globally optimal answer. The proof must connect the local rule to the complete objective, rather than merely show that each partial choice is feasible. A good do/don’t to internalize right away: -* **Do** write down the constraint you must never violate (your *feasibility invariant*). -* **Don’t** trust a greedy rule until you can justify it with a reusable proof pattern. +- Do write down the constraint you must never violate (your feasibility invariant). +- Don’t trust a greedy rule until you can justify it with a reusable proof pattern. -### The Greedy Proof Toolkit +## The Greedy Proof Toolkit This section is your reusable toolbox. The goal is to stop re-inventing proofs from scratch: the same 2–3 ideas show up again and again, just wearing different costumes. -#### The Universal Greedy Checklist +### The Universal Greedy Checklist For every greedy algorithm, answer these five questions: -1. **Greedy choice**: What local choice do we make at each step? -2. **Feasibility**: What constraint must always remain true? (the invariant) -3. **Why safe**: Why does this choice never hurt optimality? -4. **Implementation**: What data structure makes the choice fast? -5. **Complexity**: What dominates runtime? +1. Greedy choice: What local choice do we make at each step? +2. Feasibility: What constraint must always remain true? (the invariant) +3. Why safe: Why does this choice never hurt optimality? +4. Implementation: What data structure makes the choice fast? +5. Complexity: What dominates runtime? Why you should care: if you can’t fill in #3, you don’t have an algorithm yet, you have a hopeful heuristic. The checklist is your “am I actually done?” filter. -#### Exchange Argument Template +### Exchange Argument Template The exchange argument is the workhorse of greedy correctness proofs. It’s the “I can transform any optimal solution into greedy without paying more” trick. Generic template: -1. Take any optimal solution **OPT** that disagrees with greedy **G** at the first position. +1. Take any optimal solution OPT that disagrees with greedy G at the first position. 2. Show you can swap in the greedy choice at that position without making the solution worse or breaking feasibility. 3. The modified solution is still optimal and now agrees with greedy for one more position. -4. Repeat until **OPT** becomes **G** → greedy is optimal. +4. Repeat until OPT becomes G → greedy is optimal. -*Picture it like this:* +Picture it like this: ``` position → 0 1 2 3 4 greedy: [✓] [✗] [✓] [✓] [✗] some optimal: ✓ ✓ ✗ ? ? -First mismatch at position 2 → swap in greedy's pick without harm. +First mismatch at position 1 → swap in greedy's pick without harm. Repeat until both rows match → greedy is optimal. ``` -Do/don’t: **do** look for the *first* disagreement (it keeps the swap localized), and **don’t** try to “compare whole solutions at once”, you usually only need one clean swap. +Do/don’t: do look for the first disagreement (it keeps the swap localized), and don’t try to “compare whole solutions at once”, you usually only need one clean swap. -#### Loop Invariant Template +### Loop Invariant Template Sometimes it’s easier to prove that your partial solution is “the best possible so far,” iteration by iteration. Template: -1. **State the invariant**: a sentence true after every iteration. -2. **Initialization**: show it holds before the loop starts. -3. **Maintenance**: show one iteration preserves it. -4. **Termination**: show the invariant implies optimality when the loop ends. +1. State the invariant: a sentence true after every iteration. +2. Initialization: show it holds before the loop starts. +3. Maintenance: show one iteration preserves it. +4. Termination: show the invariant implies optimality when the loop ends. This is especially nice when greedy isn’t “choose and lock forever,” but “maintain a best frontier/label” (reachability scans, shortest paths, etc.). -#### Cut and Cycle Rules (Graph Specializations) +### Cut and Cycle Rules (Graph Specializations) For graph problems, exchange arguments often become two famous “rules”: -* **Cut rule (safe to add)**: For any partition $(S, V\setminus S)$, the cheapest edge crossing that cut can safely be included in an MST. -* **Cycle rule (safe to skip)**: In any cycle, the most expensive edge is never in an MST. +- Cut rule: a minimum-weight edge crossing a cut belongs to some MST. To extend a forest already chosen by a greedy algorithm, use a cut that no chosen edge crosses; then the new edge can coexist with those choices in an MST. +- Cycle rule: an edge strictly heavier than every other edge on a cycle belongs to no MST. If several edges tie for maximum weight, any one can be excluded from some MST, but it may belong to another. Do not discard all tied edges at once. + +These are exchange arguments with graph language. Equal-weight edges can produce multiple optimal trees; the rule guarantees a compatible optimum, not that every optimum uses the same edges. [Princeton Algorithms: minimum spanning trees](https://algs4.cs.princeton.edu/43mst/). -These are exchange arguments with graph language. The cut rule says “swapping in the cheapest crossing edge can’t hurt,” and the cycle rule says “dropping the heaviest edge in a cycle can’t hurt.” + The cut rule says “swapping in the cheapest crossing edge can’t hurt,” and the cycle rule says “dropping the heaviest edge in a cycle can’t hurt.” -### Examples Grouped by Pattern +## Examples Grouped by Pattern -Each example follows the same structure: **Problem → Greedy rule → Algorithm → Proof sketch → Complexity → Edge cases**. +Each example follows the same structure: Problem → Greedy rule → Algorithm → Proof sketch → Complexity → Edge cases. That structure is not cosmetic. It’s the point: if you can tell the story cleanly, you understand the algorithm, and you can re-derive it under pressure. -#### Reachability on a line +### Reachability on a line -**Pattern**: Frontier/Reachability Greedy +Pattern: Frontier/Reachability Greedy -This pattern is the “don’t overthink it” cousin of graph reachability. Instead of tracking *every* reachable square, you track a *frontier*: the furthest place your current knowledge guarantees you can get to. The *do* is: compress lots of reachability facts into one summary number. The *don’t* is: simulate every jump or mark every square repeatedly when you only need to know how far the wave has pushed. +This pattern is the “don’t overthink it” cousin of graph reachability. Instead of tracking every reachable square, you track a frontier: the furthest place your current knowledge guarantees you can get to. Compress lots of reachability facts into one summary number. Do not simulate every jump or mark every square repeatedly when you only need to know how far the wave has pushed. ``` Mental model: a growing flashlight beam @@ -100,13 +104,13 @@ Mental model: a growing flashlight beam [=======] <- everything up to F is "lit" (reachable) ``` -**Problem** +Problem: -* You stand at square $0$ on squares $0,1,\ldots,n-1$. -* Each square $i$ has jump power $a[i]$. From $i$ you may land on any of $i+1,i+2,\dots,i+a[i]$. -* Goal: decide if you can reach $n-1$; if not, report the furthest reachable square. +- Assume a nonempty array of non-negative integer jump lengths. You stand at square $0$ on squares $0,1,\ldots,n-1$. +- Each square $i$ has jump power $a[i]$. From $i$ you may land on any of $i+1,i+2,\dots,i+a[i]$. +- Goal: decide if you can reach $n-1$; if not, report the furthest reachable square. -Why this matters: it’s the simplest form of “can I keep progressing?” problems (game levels, packet routing along hops, scheduling with ranges). The *do* is: treat each `a[i]` as “potential energy” you can spend once you actually reach `i`. The *don’t* is: assume a large jump power somewhere helps unless you can *get there*. +Why this matters: it’s the simplest form of “can I keep progressing?” problems (game levels, packet routing along hops, scheduling with ranges). Treat each `a[i]` as “potential energy” you can spend once you actually reach `i`. Do not assume a large jump power somewhere helps unless you can get there. ``` Key gotcha: @@ -117,11 +121,11 @@ Big power at 10 is useless if you get stuck at 6. X gap (a[10] = 100 doesn't matter) ``` -**Example** +Example Input: $a=[3,1,0,0,4,1]$, so $n=6$. -This example is perfect because it contains both “early progress” and a “dead zone.” It forces the algorithm to prove it can detect getting stuck without backtracking. The *do* is: watch for the first index that lies beyond your frontier. That’s the exact moment the game ends. +This example is perfect because it contains both “early progress” and a “dead zone.” It forces the algorithm to prove it can detect getting stuck without backtracking. Watch for the first index that lies beyond your frontier. That’s the exact moment the game ends. ``` indices: 0 1 2 3 4 5 @@ -134,8 +138,8 @@ From any $i$, the allowed landings are a range: ``` i=0 (a[0]=3): 1..3 i=1 (a[1]=1): 2 -i=2 (a[2]=0): , -i=3 (a[3]=0): , +i=2 (a[2]=0): , +i=3 (a[3]=0): , i=4 (a[4]=4): 5..8 (board ends at 5) ``` @@ -152,7 +156,7 @@ Visualizing the trap: So if frontier never reaches 4, you're done. ``` -**Baseline idea** +Baseline idea: “Paint everything reachable, one wave at a time.” @@ -164,7 +168,7 @@ Correct, but it can reprocess many squares. This is a classic “why greedy exists” moment: the baseline is correct but wasteful because it keeps rediscovering the same reachability facts. Greedy will compress “everything I’ve learned so far” into one number. -The baseline is basically a slow-motion flood fill. It’s conceptually comforting because it mirrors how you’d explain reachability to a human: “from here I can go there, and from there I can go…” But it’s also the classic performance mistake: you keep scanning regions you already know are reachable. The *do* is: steal the baseline’s correctness idea, then compress it. The *don’t* is: keep “painting” when the only question is “how far can the paint possibly spread?” +The baseline is basically a slow-motion flood fill. It’s conceptually comforting because it mirrors how you’d explain reachability to a human: “from here I can go there, and from there I can go…” But it’s also the classic performance mistake: you keep scanning regions you already know are reachable. Steal the baseline’s correctness idea, then compress it. Do not keep “painting” when the only question is “how far can the paint possibly spread?” ``` Baseline flood (inefficient): @@ -176,18 +180,18 @@ Round 4: from 2 add {} ... ... lots of re-checking for no new info ``` -**Greedy rule** +Greedy rule: Carry one number while scanning left→right: the furthest frontier $F$ seen so far. -Why one number is enough: on a line, if you know you can reach everything up to `F`, then you’ve implicitly learned *all* the starting points that matter for future jumps. Every index ≤ F is a possible launchpad. So instead of tracking them individually, you treat them as a single “reachable prefix.” The *do* is: think “reachable prefix length.” The *don’t* is: store per-square reachability unless you’re asked for paths or counts. +Why one number is enough: on a line, if you know you can reach everything up to `F`, then you’ve implicitly learned all the starting points that matter for future jumps. Every index ≤ F is a possible launchpad. So instead of tracking them individually, you treat them as a single “reachable prefix.” think “reachable prefix length.” Do not store per-square reachability unless you’re asked for paths or counts. Rules: -* If you are at $i$ with $i>F$, you hit a gap → stuck forever. -* Otherwise extend $F\leftarrow \max(F,i+a[i])$ and continue. +- If you are at $i$ with $i>F$, you hit a gap → stuck forever. +- Otherwise extend $F\leftarrow \max(F,i+a[i])$ and continue. -This “gap” rule is the entire punchline. If you arrive at an index you can’t even stand on, then nothing to the right can be reached either, because all jumps move forward, and you’ve already accounted for all possible forward reach from all reachable places. The *do* is: stop immediately on the first gap. The *don’t* is: keep scanning hoping a later jump fixes it (it can’t). +This “gap” rule is the entire punchline. If you arrive at an index you can’t even stand on, then nothing to the right can be reached either, because all jumps move forward, and you’ve already accounted for all possible forward reach from all reachable places. Stop immediately on the first gap. Do not keep scanning hoping a later jump fixes it (it can’t). ``` Gap logic (why stopping is correct): @@ -201,12 +205,12 @@ So nothing beyond i can be reached either. At the end: -* Can reach last iff $F\ge n-1$. -* Furthest reachable square is $F$ (capped by $n-1$). +- Can reach last iff $F\ge n-1$. +- Furthest reachable square is $F$ (capped by $n-1$). -This gives you both answers “for free”: a yes/no (did we cover the last index?) and a diagnostic (how far did we get before the road ended?). The *do* is: return both in interviews / debugging, “no, and here’s where it fails.” The *don’t* is: just return false and throw away the useful boundary. +This gives you both answers “for free”: a yes/no (did we cover the last index?) and a diagnostic (how far did we get before the road ended?). Return both in interviews / debugging, “no, and here’s where it fails.” Do not just return false and throw away the useful boundary. -**Algorithm** +Algorithm ``` F = 0 @@ -218,7 +222,7 @@ can_reach_last = (F >= n-1) furthest = min(F, n-1) ``` -What makes this feel “human” is that it matches how you’d actually play: you keep walking forward as long as you’re within what you already know you can reach, and every time you land somewhere with jump power, you update your best possible future. The *do* is: read it like “keep upgrading my maximum reach.” The *don’t* is: interpret it as “I must jump at every square”, you’re not choosing a specific jump sequence; you’re summarizing all possible sequences. +What makes this feel “human” is that it matches how you’d actually play: you keep walking forward as long as you’re within what you already know you can reach, and every time you land somewhere with jump power, you update your best possible future. Read it like “keep upgrading my maximum reach.” Do not interpret it as “I must jump at every square”, you’re not choosing a specific jump sequence; you’re summarizing all possible sequences. ``` Frontier evolution on the example a=[3,1,0,0,4,1] @@ -234,16 +238,16 @@ i=4 > F -> GAP -> stop Result: furthest reachable = 3 (can't reach 5) ``` -**Proof sketch** +Proof sketch: Loop invariant: “After processing index $i$, $F$ equals the furthest position reachable using jumps that start at some truly reachable square $\le i$.” -This invariant is doing a lot of work, and it’s worth appreciating why it’s phrased that way: you’re only allowed to use jump power from squares you can actually reach, and you’re only considering launch points up to the current scan index. That matches the algorithm’s left-to-right processing. The *do* is: tie the invariant to what the loop has “seen.” The *don’t* is: claim $F$ is “furthest reachable overall” mid-loop; it’s “furthest reachable given processed launchpads.” +This invariant is doing a lot of work, and it’s worth appreciating why it’s phrased that way: you’re only allowed to use jump power from squares you can actually reach, and you’re only considering launch points up to the current scan index. That matches the algorithm’s left-to-right processing. Tie the invariant to what the loop has “seen.” Do not claim $F$ is “furthest reachable overall” mid-loop; it’s “furthest reachable given processed launchpads.” -* Initialization: $F=0$ is correct at the start. -* Maintenance: if $i\le F$, then $i$ is reachable, so $i+a[i]$ is a valid new reach and updating $F$ is safe. -* If $i>F$, no earlier reachable square can jump to $i$ (all such reach was already summarized in $F$), so stopping is correct. -* Termination: the final $F$ is exactly the furthest reachable position. +- Initialization: $F=0$ is correct at the start. +- Maintenance: if $i\le F$, then $i$ is reachable, so $i+a[i]$ is a valid new reach and updating $F$ is safe. +- If $i>F$, no earlier reachable square can jump to $i$ (all such reach was already summarized in $F$), so stopping is correct. +- Termination: $\min(F,n-1)$ is the furthest reachable board position. The uncapped frontier may extend beyond the end of the board. A nice way to visualize the proof is to imagine the scan as “unlocking” jump powers. You can only use `a[i]` after you confirm `i` is inside the unlocked zone (`i ≤ F`). Every time you unlock a square, you possibly expand the unlocked zone. If you ever encounter a locked square (`i > F`), the unlocking process cannot continue. @@ -258,16 +262,16 @@ If i steps beyond F, we found the first unreachable launchpad. Forward-only jumps => no future square can fix that. ``` -**Complexity**: time $O(n)$, space $O(1)$. +Complexity: time $O(n)$, space $O(1)$. -This is the reward for choosing the right summary statistic. Instead of simulating many possibilities, you scan once and keep one frontier value. The *do* is: recognize this as “linear scan with a running max.” The *don’t* is: accidentally reintroduce extra work (nested loops) when implementing. +This is the reward for choosing the right summary statistic. Instead of simulating many possibilities, you scan once and keep one frontier value. Recognize this as “linear scan with a running max.” Do not accidentally reintroduce extra work (nested loops) when implementing. -**Edge cases** +Edge cases: -* If $n=1$, you’re already at the goal. -* If many $a[i]=0$, the scan correctly stops at the first gap. +- If $n=1$, you’re already at the goal. +- If many $a[i]=0$, the scan correctly stops at the first gap. -These edge cases are basically the algorithm’s personality check. For `n=1`, the frontier starts at the goal. For many zeros, you quickly find where progress ends, and you stop without wasted work. The *do* is: handle `n=1` cleanly. The *don’t* is: try to “jump from nowhere” past a gap, forward-only movement makes gaps final. +These edge cases are basically the algorithm’s personality check. For `n=1`, the frontier starts at the goal. For many zeros, you quickly find where progress ends, and you stop without wasted work. Handle `n=1` cleanly. Do not try to “jump from nowhere” past a gap, forward-only movement makes gaps final. ``` Zeros create cliffs: @@ -281,11 +285,11 @@ i=3 => i>F => stop at 2 Cliff detected exactly where it should be. ``` -#### Minimum spanning trees +### Minimum spanning trees -**Pattern**: Cut-Based Greedy on Graphs +Pattern: Cut-Based Greedy on Graphs -Before we even touch the algorithms, it helps to picture the “job” an MST is doing: you want *all* the vertices connected, you want to spend as little total weight as possible, and you’re not allowed to waste edges making loops. This shows up everywhere, laying fiber between cities, wiring circuits on a board, connecting servers in a data center, clustering points in ML, any time “connect everything cheaply” matters. The *do* here is: think “infrastructure budget.” The *don’t* is: treat it like a random graph puzzle; it’s really about spending weight efficiently. +Before we even touch the algorithms, it helps to picture the “job” an MST is doing: you want all the vertices connected, you want to spend as little total weight as possible, and you’re not allowed to waste edges making loops. This shows up everywhere, laying fiber between cities, wiring circuits on a board, connecting servers in a data center, clustering points in ML, any time “connect everything cheaply” matters. The do here is: think “infrastructure budget.” Do not treat it like a random graph puzzle; it’s really about spending weight efficiently. ``` Goal (MST): connect all nodes, no cycles, minimum total cost @@ -300,11 +304,11 @@ Bad: has a cycle (extra spend) Good: just enough edges to connect all nodes (|V|-1 edges) ``` -**Problem** +Problem: Given a connected, undirected, weighted graph, connect all vertices with minimum total edge weight without cycles: an MST. -A useful way to “feel” this constraint is: a tree with `|V|` vertices always has exactly `|V|-1` edges. So if you ever add an edge that creates a cycle, you’ve basically bought an unnecessary cable. The *do* is: keep asking “does this edge actually help me reach a new vertex/component?” The *don’t* is: add edges just because they look cheap, cheap and *useful* is what matters. +A useful way to “feel” this constraint is: a tree with `|V|` vertices always has exactly `|V|-1` edges. So if you ever add an edge that creates a cycle, you’ve basically bought an unnecessary cable. Keep asking “does this edge actually help me reach a new vertex/component?” Do not add edges just because they look cheap, cheap and useful is what matters. ``` Why cycles are waste (intuition) @@ -318,11 +322,11 @@ and still keep everything connected. So that heaviest edge was never necessary. ``` -**Baseline** +Baseline: Enumerate all spanning trees and pick the lightest, correct but combinatorially impossible for real graphs. This baseline is important conceptually: it reminds you what “optimal” even means, and it highlights why a greedy shortcut is valuable. -This is the “brute-force North Star.” It’s the version you’d run if graphs were tiny and time was infinite. The *do* is: keep it in your head as the definition of correctness (“we’re trying to match what brute force would choose”). The *don’t* is: ever implement it outside of toy examples, its only real job is to motivate why we need smarter structure. +Exhaustive search defines a useful correctness baseline on small graphs: compare all feasible spanning trees and select the lightest. Its rapidly growing search space motivates a proof that greedy selection reaches the same optimum. ``` Brute-force vs. greedy (scale intuition) @@ -335,14 +339,14 @@ Real graph: "nope" Greedy works because we can prove certain choices are "safe". ``` -**Two facts that power the greedy proofs** +Two facts that power the greedy proofs -* Cut rule (safe to add) -* Cycle rule (safe to skip) +- Cut rule (safe to add) +- Cycle rule (safe to skip) These are the “why safe” pieces that let greedy commit early without regret. -These two rules are the emotional support system for greedy algorithms. Greedy is scary because it “locks in” choices without seeing the future, but MST is one of the lucky problems where we can prove some choices can’t hurt. The *do* is: when you read a proof, hunt for “cut” or “cycle” language. The *don’t* is: memorize steps without internalizing what makes them safe; the rules are the whole reason the algorithms work. +Cut and cycle properties explain why MST algorithms can commit to particular edges. Identify the relevant cut or cycle at each step, and check that the rule preserves an optimum compatible with the edges already selected. ``` Cut rule (picture a cut) @@ -368,15 +372,15 @@ Heaviest edge on the cycle is safe to skip. (Here: weight 5) ``` -#### Kruskal’s method +### Kruskal’s method -Kruskal feels like “shopping with a strict budget.” You walk through edges from cheapest to priciest, and you only buy an edge if it doesn’t create a loop. The algorithm is simple; the magic is that the *rules above* guarantee you’re not painting yourself into a corner. The *do* is: think “merge components.” The *don’t* is: think “build one expanding blob” (that’s Prim). +Kruskal processes edges globally from lightest to heaviest, adding an edge only when it joins distinct components. The cut and cycle properties justify those choices; disjoint-set union tracks the components efficiently. -**Greedy rule** +Greedy rule: Sort edges by weight. Scan from lightest to heaviest and keep an edge if it connects two different components (i.e., doesn’t form a cycle). Stop when you have $|V|-1$ edges. -The “two different components” test is the entire game. Early on, every vertex is its own component; every accepted edge fuses two components into a bigger one. The *do* is: visualize components as islands and edges as bridges. The *don’t* is: accept a bridge that starts and ends on the same island, congrats, you just paid for a scenic loop. +The “two different components” test is the entire game. Early on, every vertex is its own component; every accepted edge fuses two components into a bigger one. Visualize components as islands and edges as bridges. Do not accept a bridge that starts and ends on the same island, congrats, you just paid for a scenic loop. ``` Kruskal = "connect islands cheaply" @@ -390,11 +394,11 @@ Pick smallest edges that connect DIFFERENT sets: {C}-{D} => {ACBD} (done) ``` -**Implementation idea** +Implementation idea Use a disjoint-set union-find structure to test whether two vertices are already connected. -Union-find is basically your “island tracker.” It answers: “Are u and v already in the same component?” quickly, and if not, it merges them. The *do* is: rely on union-find for speed. The *don’t* is: do connectivity checks with full graph searches per edge, you’ll turn a fast algorithm into a slow one. +Union-find is basically your “island tracker.” It answers: “Are u and v already in the same component?” quickly, and if not, it merges them. Rely on union-find for speed. Do not do connectivity checks with full graph searches per edge, you’ll turn a fast algorithm into a slow one. ``` Union-Find (DSU) mental model @@ -406,13 +410,13 @@ If find(u) == find(v): adding (u,v) makes a cycle -> skip Else: safe to add -> union(u,v) ``` -**Proof sketch** +Proof sketch: -* If an edge would create a cycle, the cycle rule says it’s safe to skip. -* If an edge is the lightest that connects two components, it is the lightest crossing edge of some cut, and the cut rule says it’s safe to include. -* Exchange argument: transform any optimal MST to include Kruskal’s chosen edge without increasing weight. +- If an edge would create a cycle, the cycle rule says it’s safe to skip. +- If an edge is the lightest that connects two components, it is the lightest crossing edge of some cut, and the cut rule says it’s safe to include. +- Exchange argument: transform any optimal MST to include Kruskal’s chosen edge without increasing weight. -Here’s the flow that makes this proof feel human: every time Kruskal picks an edge, it’s either (a) obviously wasteful (it would create a cycle) so you skip it, or (b) it’s the cheapest way to connect two currently-separated groups, meaning it’s the cheapest edge crossing the cut between those groups. The exchange argument is the “no hard feelings” clause: even if the optimal MST you imagined didn’t include your chosen edge, you can swap edges and not increase cost. The *do* is: connect each bullet to either “cut” or “cycle.” The *don’t* is: treat “exchange argument” like a spell, see it as a controlled swap that keeps the tree valid and not heavier. +Here’s the flow that makes this proof feel human: every time Kruskal picks an edge, it’s either (a) obviously wasteful (it would create a cycle) so you skip it, or (b) it’s the cheapest way to connect two currently-separated groups, meaning it’s the cheapest edge crossing the cut between those groups. The exchange argument is the “no hard feelings” clause: even if the optimal MST you imagined didn’t include your chosen edge, you can swap edges and not increase cost. Connect each bullet to either “cut” or “cycle.” Do not treat “exchange argument” like a spell, see it as a controlled swap that keeps the tree valid and not heavier. ``` Exchange argument (tiny sketch) @@ -427,20 +431,20 @@ Result is still a spanning tree and no heavier. So you can "exchange" f for e safely. ``` -**Complexity** +Complexity: -* Sorting: $O(E\log E)$ (often written $O(E\log V)$). -* Union-find operations: near-constant amortized ($\alpha(V)$). -* Space: $O(V)$. +- Sorting: $O(E\log E)$ (often written $O(E\log V)$). +- Union-find operations: near-constant amortized ($\alpha(V)$). +- Auxiliary space: $O(V)$ for DSU, plus the output and any storage required by edge sorting. Storing the input edge list uses $O(E)$. -Interpretation: most of the time is spent sorting edges; union-find is the lightweight bouncer at the club door. The *do* is: remember “sort dominates.” The *don’t* is: overthink $\alpha(V)$, it’s effectively constant for any practical input size. +Interpretation: most of the time is spent sorting edges; union-find is the lightweight bouncer at the club door. Remember “sort dominates.” Do not overthink $\alpha(V)$, it’s effectively constant for any practical input size. -**Edge cases** +Edge cases: -* Disconnected graph → minimum spanning forest. -* Ties are fine: multiple MSTs may exist with the same total weight. +- Disconnected graph → minimum spanning forest. +- Ties are fine: multiple MSTs may exist with the same total weight. -Practical takeaway: in real data, disconnection is common (clusters, communities, separated regions). Kruskal doesn’t panic; it just builds one MST per connected component. And if weights tie, the graph is basically saying “you have multiple equally good designs.” The *do* is: accept that MST may not be unique. The *don’t* is: expect identical edge sets across runs if tie-breaking differs. +Practical takeaway: in real data, disconnection is common (clusters, communities, separated regions). Kruskal doesn’t panic; it just builds one MST per connected component. And if weights tie, the graph is basically saying “you have multiple equally good designs.” accept that MST may not be unique. Do not expect identical edge sets across runs if tie-breaking differs. ``` Disconnected -> forest @@ -451,15 +455,15 @@ Component 2: D--E MST2 Output = {MST1, MST2} ``` -#### Prim’s method +### Prim’s method -Prim feels like “growing a single organism.” You start from one vertex and keep attaching the cheapest edge that reaches something new. If Kruskal is “merge islands everywhere,” Prim is “expand one territory.” The *do* is: think “frontier/boundary.” The *don’t* is: think “global cheapest edge anywhere” (that’s Kruskal’s vibe). +Prim grows one tree from a starting vertex. Its candidates are edges crossing from the current tree to an outside vertex, and it takes a minimum-weight candidate at each step. -**Greedy rule** +Greedy rule: Grow one tree: repeatedly add the lightest edge leaving the current tree to bring in a new vertex. -This “leaving the current tree” phrase is the key limiter that makes Prim different: you’re only allowed to choose edges that cross from inside to outside. The *do* is: picture a boundary fence around your current tree. The *don’t* is: pick an edge fully outside the tree (even if it’s super cheap), because it doesn’t help your current structure grow. +This “leaving the current tree” phrase is the key limiter that makes Prim different: you’re only allowed to choose edges that cross from inside to outside. Picture a boundary fence around your current tree. Do not pick an edge fully outside the tree (even if it’s super cheap), because it doesn’t help your current structure grow. ``` Prim boundary picture @@ -475,11 +479,11 @@ Only consider edges that cross the boundary: Pick the cheapest crossing edge, add that outside vertex. ``` -**Implementation idea** +Implementation idea Use a min-heap (priority queue) keyed by the cheapest “boundary edge” to each outside vertex. -The heap is your “best next deal” list: for every outside vertex, keep track of the cheapest known edge that would bring it into the tree. The *do* is: update keys when you find a cheaper connection. The *don’t* is: try to keep the heap perfectly clean at all times, real implementations often allow duplicates and ignore stale entries later. +The heap is your “best next deal” list: for every outside vertex, keep track of the cheapest known edge that would bring it into the tree. Update keys when you find a cheaper connection. Do not try to keep the heap perfectly clean at all times, real implementations often allow duplicates and ignore stale entries later. ``` Min-heap idea (keys = best known connection cost) @@ -493,11 +497,11 @@ When you add E, you may discover: D can be reached with weight 1 instead of 4 -> decrease-key (or push new entry) ``` -**Proof sketch** +Proof sketch: -At every step, Prim picks the lightest edge crossing the cut (tree vs. outside). By the cut rule, that edge belongs to some MST, so committing to it is safe. +At every step, Prim picks the lightest edge crossing the cut (tree vs. Outside). By the cut rule, that edge belongs to some MST, so committing to it is safe. -This proof is basically one clean sentence: Prim always picks the lightest edge crossing the specific cut “current tree vs. everything else,” and the cut rule says that’s safe. The *do* is: explicitly name the cut each step. The *don’t* is: get lost in implementation details (heap, adjacency lists) when you’re trying to understand correctness. +This proof is basically one clean sentence: Prim always picks the lightest edge crossing the specific cut “current tree vs. Everything else,” and the cut rule says that’s safe. Explicitly name the cut each step. Do not get lost in implementation details (heap, adjacency lists) when you’re trying to understand correctness. ``` Prim correctness in one diagram @@ -510,25 +514,25 @@ e* = lightest edge crossing this cut Cut rule => e* is safe => you can keep growing. ``` -**Complexity** +Complexity: -* Binary heap + adjacency lists: $O(E\log V)$. -* Space: $O(V)$ plus adjacency representation. +- Binary heap + adjacency lists: $O(E\log V)$. +- Auxiliary space: $O(V)$ with an indexed queue, or $O(V+E)$ when improved entries are pushed lazily. Input adjacency lists use $O(V+E)$. -Meaning: every edge might cause a heap update-ish operation, and the heap costs log V per meaningful push/pop. The *do* is: use adjacency lists (especially for sparse graphs). The *don’t* is: use an adjacency matrix for huge sparse graphs unless you really mean it. +Meaning: every edge might cause a heap update-ish operation, and the heap costs log V per meaningful push/pop. Use adjacency lists (especially for sparse graphs). Do not use an adjacency matrix for huge sparse graphs unless you really mean it. -**Edge cases** +Edge cases: -* Start vertex doesn’t affect total MST weight (though the chosen edges may differ). -* Heaps may contain stale entries; skip when popped. +- Start vertex doesn’t affect total MST weight (though the chosen edges may differ). +- Heaps may contain stale entries; skip when popped. -If you start at a different node, you may build a different-looking MST but with the same optimal total weight (assuming ties/structure allow multiple). And if your heap has outdated offers, just ignore them when you notice they no longer match the best-known state. The *do* is: code defensively with a visited set / current best key checks. The *don’t* is: assume the heap always reflects the current truth without verification. +If you start at a different node, you may build a different-looking MST but with the same optimal total weight (assuming ties/structure allow multiple). And if your heap has outdated offers, just ignore them when you notice they no longer match the best-known state. Code defensively with a visited set / current best key checks. Do not assume the heap always reflects the current truth without verification. -#### Shortest paths with non-negative weights (Dijkstra) +### Shortest paths with non-negative weights (Dijkstra) -**Pattern**: Cut-Based Greedy on Graphs +Pattern: Cut-Based Greedy on Graphs -Dijkstra is the “no regrets” version of shortest paths: you keep a boundary between what you *know for sure* and what you’re still *guessing*, and you repeatedly promote the safest-looking guess into certainty. The whole reason this works is non-negativity, roads don’t give you refunds. The *do* is: treat the algorithm as expanding a region of confirmed shortest distances. The *don’t* is: use it when edges can be negative; a negative edge is exactly the kind of “refund detour” that breaks the safety logic. +Dijkstra is the “no regrets” version of shortest paths: you keep a boundary between what you know for sure and what you’re still guessing, and you repeatedly promote the safest-looking guess into certainty. The whole reason this works is non-negativity, roads don’t give you refunds. Treat the algorithm as expanding a region of confirmed shortest distances. Do not use it when edges can be negative; a negative edge is exactly the kind of “refund detour” that breaks the safety logic. ``` Two-zone view: @@ -540,11 +544,11 @@ Dijkstra repeatedly moves the boundary rightward: it "locks in" one node at a time. ``` -**Problem** +Problem: From a start node $s$, compute shortest distances $d(\cdot)$ to all nodes (and routes via parents). Edge weights must be non-negative. -Why you should care: shortest paths is the backbone of routing, navigation, dependency planning, game AI, and “minimum cost to reach X” problems. The output isn’t just numbers, it’s also *parents* (how you actually get there). The *do* is: maintain both distance and predecessor for reconstruction. The *don’t* is: only compute distances and then wonder how to produce the path later. +Why you should care: shortest paths is the backbone of routing, navigation, dependency planning, game AI, and “minimum cost to reach X” problems. The output isn’t just numbers, it’s also parents (how you actually get there). Maintain both distance and predecessor for reconstruction. Do not only compute distances and then wonder how to produce the path later. ``` Distances + parents => path reconstruction @@ -556,11 +560,11 @@ t <- parent[t] <- parent[parent[t]] <- ... <- s (reverse it) ``` -**Baseline** +Baseline: Bellman–Ford-style repeated relaxation: about $O(|V||E|)$. It handles negatives; Dijkstra doesn’t need that generality, so greedy buys speed. -The baseline is “keep trying to improve everyone until nothing changes.” It’s powerful because it works even with negative edges, but it spends time rechecking improvements that can’t possibly matter when all edges are non-negative. Greedy says: “instead of endlessly revisiting, let’s *finalize* nodes when we’re sure.” The *do* is: see Dijkstra as a faster specialization of relaxation. The *don’t* is: forget that Bellman–Ford exists when negatives appear. +The baseline is “keep trying to improve everyone until nothing changes.” It’s powerful because it works even with negative edges, but it spends time rechecking improvements that can’t possibly matter when all edges are non-negative. Greedy says: “instead of endlessly revisiting, let’s finalize nodes when we’re sure.” see Dijkstra as a faster specialization of relaxation. Do not forget that Bellman–Ford exists when negatives appear. ``` Relaxation idea (shared by both): @@ -572,11 +576,11 @@ Difference: - Dijkstra chooses an order that makes some nodes final early. ``` -**Greedy rule: “settle the smallest label”** +Greedy rule: “settle the smallest label” -Repeatedly pick the unsettled node $u$ with smallest tentative distance $d(u)$ and **settle** it (declare its distance final). +Repeatedly pick the unsettled node $u$ with smallest tentative distance $d(u)$ and settle it (declare its distance final). -“Settle” is the key word: once settled, a node never changes again. That’s the greedy commitment. The *do* is: interpret “smallest label” as “closest frontier point.” The *don’t* is: settle nodes in arbitrary order; the safety proof depends on always picking the minimum tentative distance next. +“Settle” is the key word: once settled, a node never changes again. That’s the greedy commitment. Interpret “smallest label” as “closest frontier point.” Do not settle nodes in arbitrary order; the safety proof depends on always picking the minimum tentative distance next. ``` Label-setting picture: @@ -589,11 +593,11 @@ A: 7 B: 3 C: 11 D: 5 Pick the smallest (B:3), settle it, relax its outgoing edges. ``` -**Why safe** +Why safe -With non-negative weights, any path to $u$ that detours through other unsettled nodes can only add non-negative cost, so it cannot beat the current smallest label. +Consider a shortest path to the minimum-label unsettled vertex $u$, and let $y$ be its first unsettled vertex. Its predecessor is settled, so relaxing that predecessor's edges has already given $y$ its optimal prefix cost. Therefore $d(u) \le d(y) \le \delta(u)$, where $\delta(u)$ is the true shortest distance. Since any tentative label is a real path cost, $d(u) \ge \delta(u)$, proving equality. -This is the intuition that makes the algorithm feel “obviously right”: if `u` is currently the cheapest unsettled node, then going from `s` to some other unsettled node `x` (which is already ≥ `d(u)`) and then traveling more edges to reach `u` can’t magically become cheaper, because those extra edges can’t subtract cost. The *do* is: remember “detours only add.” The *don’t* is: forget that a single negative edge is enough to make a detour beneficial. +Non-negativity matters because the rest of the path cannot reduce the cost of that first unsettled prefix. Tentative labels are upper bounds, so it would be incorrect to assume that every route through an arbitrary unsettled node costs at least that node's current label. The proof relies specifically on the first crossing from settled to unsettled vertices. ``` Why non-negative matters: @@ -602,27 +606,27 @@ If all edges >= 0, then detour cost >= 0. So if u is the cheapest tentative node, any route that reaches u via another unsettled node -must be >= that other node's tentative distance >= d[u]. +must cost at least the optimal first-crossing prefix >= d[u]. No detour can undercut d[u]. ``` -**Proof sketch** +Proof sketch: Loop invariant: “All settled nodes have correct shortest-path distances.” -The invariant is your safety harness: we’re gradually building a set `S` of nodes whose distances are finalized and correct. Each step must preserve that truth. The *do* is: keep the invariant in mind while reading the algorithm. The *don’t* is: think Dijkstra is “just BFS with weights”, the correctness comes from this settle-min step plus non-negativity. +The invariant is your safety harness: we’re gradually building a set `S` of nodes whose distances are finalized and correct. Each step must preserve that truth. Keep the invariant in mind while reading the algorithm. Do not think Dijkstra is “just BFS with weights”, the correctness comes from this settle-min step plus non-negativity. -* Initialization: $d(s)=0$ is correct. -* Maintenance: settling the minimum-label node is safe by non-negativity. -* Termination: when all reachable nodes are settled, distances are correct. +- Initialization: $d(s)=0$ is correct. +- Maintenance: settling the minimum-label node is safe by non-negativity. +- Termination: when all reachable nodes are settled, distances are correct. A cut-based way to picture the maintenance step: -* Let `S` be settled nodes, `V \ S` unsettled. -* Dijkstra chooses `u` in `V \ S` with minimum `d(u)`. -* Any path to an unsettled node must cross the cut at some edge out of `S`. -* `d(u)` is already the best possible among those, so it’s final. +- Let `S` be settled nodes, `V \ S` unsettled. +- Dijkstra chooses `u` in `V \ S` with minimum `d(u)`. +- Any path to an unsettled node must cross the cut at some edge out of `S`. +- `d(u)` is already the best possible among those, so it’s final. ``` Cut view: @@ -634,12 +638,12 @@ All known best routes to outside go through this boundary. Pick the smallest boundary-reachable node u, lock it in. ``` -**Complexity** +Complexity: -* With binary heap: $O((V+E)\log V)$. -* Space: $O(V)$. +- With binary heap: $O((V+E)\log V)$. +- Auxiliary space: $O(V)$ with an indexed queue, or $O(V+E)$ for the lazy duplicate-entry heap below. The lazy heap takes $O(V+E\log(E+2))$ time in general, simplifying to the usual bound for simple graphs. -Heap intuition: you need a fast way to repeatedly extract the smallest tentative label and to push improved distances. The *do* is: use a priority queue keyed by `d`. The *don’t* is: linearly scan for the minimum each time unless `V` is tiny. +Heap intuition: you need a fast way to repeatedly extract the smallest tentative label and to push improved distances. Use a priority queue keyed by `d`. Do not linearly scan for the minimum each time unless `V` is tiny. ``` What the heap is doing: @@ -656,12 +660,12 @@ repeat: push (d[v], v) ``` -**Edge cases** +Edge cases: -* Unreachable nodes remain at $\infty$. -* Stale heap entries occur; skip if already settled. +- Unreachable nodes remain at $\infty$. +- Stale heap entries occur; skip if already settled. -Unreachable nodes are not a failure, they’re a real output: “there is no path from s.” Stale heap entries happen because many implementations push a new `(betterDist, v)` instead of decreasing-key in place; later, the old worse entry resurfaces and must be ignored. The *do* is: keep a settled/visited check. The *don’t* is: assume every pop is “the one true current distance.” +Unreachable nodes are not a failure, they’re a real output: “there is no path from s.” Stale heap entries happen because many implementations push a new `(betterDist, v)` instead of decreasing-key in place; later, the old worse entry resurfaces and must be ignored. Keep a settled/visited check. Do not assume every pop is “the one true current distance.” ``` Stale entry example: @@ -674,13 +678,13 @@ Pop (6, v): settle v Later pop (10, v): v already settled -> skip ``` -#### Scheduling themes +### Scheduling themes -**Pattern**: Scheduling Greedy +Pattern: Scheduling Greedy Scheduling is greedy-friendly because “what you do early” shapes what’s possible later, and the proofs often boil down to a clean exchange: “finishing earlier leaves more room.” -That “leaves more room” line is the whole vibe of scheduling proofs. Time is a one-way resource: once you waste a slot early, you can’t buy it back later. So greedy strategies that **protect future flexibility** often win. The *do* is: look for a simple local choice that maximizes options later (earliest finish, earliest due date). The *don’t* is: get distracted by “what feels urgent” unless the objective matches that urgency. +Scheduling proofs often compare how much time remains after a choice. Earliest finish and earliest due date address different objectives, so state the objective and assumptions before selecting a rule. ``` Scheduling = packing into time @@ -691,15 +695,15 @@ Every choice consumes a chunk. Good greedy choices keep the remaining timeline usable. ``` -#### Interval scheduling (maximum number of non-overlapping intervals) +### Interval scheduling (maximum number of non-overlapping intervals) -This is the “fit as many meetings as possible” problem. The trick is realizing the objective is **count**, not total duration, not value, not “use time efficiently.” If you’re maximizing *how many*, then short/early-finishing intervals are gold because they leave space for more. The *do* is: optimize for free time after each pick. The *don’t* is: pick the earliest-starting interval (that’s a classic wrong greedy). +This is the “fit as many meetings as possible” problem. The trick is realizing the objective is count, not total duration, not value, not “use time efficiently.” If you’re maximizing how many, then the earliest-finishing compatible interval leaves the most time for subsequent choices. Choosing the shortest duration is a different rule and is not generally optimal. Optimize for free time after each pick. Do not pick the earliest-starting interval (that’s a classic wrong greedy). -**Baseline** +Baseline: Try all subsets → exponential. -The baseline is “check every combination of meetings and see which one works.” It’s correct but useless at scale, and it hides the structure: feasibility depends on *ordering* in time, not arbitrary set selection. The *do* is: use it to define correctness (“maximum number”). The *don’t* is: keep thinking in subsets once you see time order exists. +The baseline is “check every combination of meetings and see which one works.” It’s correct but useless at scale, and it hides the structure: feasibility depends on ordering in time, not arbitrary set selection. Use it to define correctness (“maximum number”). Do not keep thinking in subsets once you see time order exists. ``` Why subsets explode: @@ -710,11 +714,11 @@ Subsets: 2^n Even n=50 is astronomical. ``` -**Greedy rule** +Greedy rule: -Sort by finish time, take the next interval that starts after the last chosen ends. +Assume intervals have positive duration and use half-open ranges `[start, finish)`, so touching endpoints are compatible. Sort by finish time and accept an interval when `start >= last_finish`. Initialize `last_finish` below every possible start, rather than assuming times are non-negative. -This is the “always finish as early as possible” strategy. It’s not about being fast for its own sake, it’s about keeping the remaining timeline as wide as possible for future intervals. The *do* is: treat the finish time as the critical decision point. The *don’t* is: prioritize long intervals or early starts; those can block many later choices. +This is the “always finish as early as possible” strategy. It’s not about being fast for its own sake, it’s about keeping the remaining timeline as wide as possible for future intervals. Treat the finish time as the critical decision point. Do not prioritize long intervals or early starts; those can block many later choices. ``` Wrong greedy (earliest start) can fail: @@ -732,11 +736,11 @@ Earliest start picks A -> total 1 Earliest finish picks B,C,D,E,F,G,H -> total 7 ``` -**Proof sketch** +Proof sketch: -Exchange argument: if an optimal schedule picks an interval that finishes later than the earliest-finishing compatible interval, swap it out. You don’t reduce how many intervals fit afterward because you only *free up* time. +Exchange argument: if an optimal schedule picks an interval that finishes later than the earliest-finishing compatible interval, swap it out. You don’t reduce how many intervals fit afterward because you only free up time. -Here’s the “human” version of the exchange: suppose you and I both have a schedule, and at some step you picked a meeting that ends at 5pm, but there was another compatible meeting that ends at 3pm. If we swap yours for the 3pm one, everything after 5pm is still available, and now *more* time is available between 3pm and 5pm too. So swapping cannot make your future worse; it can only keep it the same or improve it. The *do* is: focus on the first decision where two schedules differ. The *don’t* is: try to prove it by complex induction before you see the simple “end earlier can’t hurt” fact. +Here’s the “human” version of the exchange: suppose you and I both have a schedule, and at some step you picked a meeting that ends at 5pm, but there was another compatible meeting that ends at 3pm. If we swap yours for the 3pm one, everything after 5pm is still available, and now more time is available between 3pm and 5pm too. So swapping cannot make your future worse; it can only keep it the same or improve it. Focus on the first decision where two schedules differ. Do not try to prove it by complex induction before you see the simple “end earlier can’t hurt” fact. ``` Exchange picture: @@ -751,22 +755,24 @@ Anything that starts after X end also starts after G end. So replacing X with G keeps all later options. ``` -**Complexity**: $O(n\log n)$ sort + $O(n)$ scan. +Complexity: $O(n\log n)$ sort + $O(n)$ scan. + +Most work is sorting once. The scan is just “take it if it fits.” implement as sort-by-end then one pass. Do not repeatedly search for the next interval in an inner loop. -Most work is sorting once. The scan is just “take it if it fits.” The *do* is: implement as sort-by-end then one pass. The *don’t* is: repeatedly search for the next interval in an inner loop. +Edge cases: -**Edge cases** +- If everything overlaps, you keep 1. +- Equal finish times can be broken arbitrarily. -* If everything overlaps, you keep 1. -* Equal finish times can be broken arbitrarily. +Equal finish times don’t change the “free up time” logic, finishing at the same time leaves the same future. Treat ties as harmless. Do not overfit tie-breaking rules unless you need reproducibility. -Equal finish times don’t change the “free up time” logic, finishing at the same time leaves the same future. The *do* is: treat ties as harmless. The *don’t* is: overfit tie-breaking rules unless you need reproducibility. +### Minimize the Maximum Lateness -#### Minimize the maximum lateness (unit-time jobs) +Assume one machine, all jobs available at time zero, and no precedence constraints. The example uses unit-time jobs; the exchange argument also extends to varying processing times as explained below. Define lateness for job $i$ as $L_i=C_i-d_i$ and objective $L_{\max}=\max_i L_i$. -This problem is subtle because it’s not “meet all deadlines” (sometimes you can’t). It’s “if someone is late, make the worst lateness as small as possible.” So fairness matters: you’re trying to prevent one job from being *catastrophically* late. The *do* is: think “minimize the worst-case pain.” The *don’t* is: optimize average lateness; that’s a different objective. +This problem is subtle because it’s not “meet all deadlines” (sometimes you can’t). It’s “if someone is late, make the worst lateness as small as possible.” So fairness matters: you’re trying to prevent one job from being catastrophically late. Think “minimize the worst-case pain.” Do not optimize average lateness; that’s a different objective. ``` Unit-time jobs = each job takes 1 slot: @@ -779,17 +785,17 @@ Lateness L_i = C_i - d_i Goal: minimize max lateness across all jobs. ``` -**Baseline** +Baseline: Try all $n!$ orders → impossible. -The baseline here is “try every permutation.” It screams that what matters is the order, and that we need a rule to choose an order without exploring all of them. The *do* is: acknowledge scheduling is about permutations. The *don’t* is: brute-force except for tiny `n`. +The baseline here is “try every permutation.” It screams that what matters is the order, and that we need a rule to choose an order without exploring all of them. Acknowledge scheduling is about permutations. Do not brute-force except for tiny `n`. -**Greedy rule (EDD)** +Greedy rule (EDD) Sort jobs by nondecreasing deadlines (Earliest Due Date first). -EDD is the “protect the earliest deadline from getting pushed back” strategy. If a job is due sooner, delaying it tends to create large lateness spikes. Putting it earlier is like paying the urgent bills first so your “overdue penalty” doesn’t explode. The *do* is: sort by `d_i`. The *don’t* is: sort by shortest job first here, processing times are all equal, so that idea is irrelevant. +EDD is the “protect the earliest deadline from getting pushed back” strategy. If a job is due sooner, delaying it tends to create large lateness spikes. Putting it earlier is like paying the urgent bills first so your “overdue penalty” doesn’t explode. Sort by `d_i`. Do not sort by shortest job first here, processing times are all equal, so that idea is irrelevant. ``` EDD visual: @@ -798,11 +804,11 @@ Deadlines: d=2 d=5 d=5 d=9 d=10 Order: earliest deadline jobs first ``` -**Proof sketch** +Proof sketch: Exchange argument on inversions: if two adjacent jobs are out of deadline order, swapping them cannot increase $L_{\max}$, and it improves (or preserves) the lateness of the earlier-deadline job. Repeatedly remove inversions → sorted order is optimal. -The “why” is very local: consider two adjacent jobs `A` then `B` where `d_A > d_B` (they’re inverted). Since both take one unit time, the pair occupies the same two slots either way, swapping only changes *which job gets the earlier slot*. Giving the earlier slot to the earlier deadline is never worse for the maximum lateness. So you bubble-sort away inversions without harming the objective, ending at EDD. The *do* is: zoom in on a two-job swap. The *don’t* is: attempt a global argument without this local swap lens. +The “why” is very local: consider two adjacent jobs `A` then `B` where `d_A > d_B` (they’re inverted). Since both take one unit time, the pair occupies the same two slots either way, swapping only changes which job gets the earlier slot. Giving the earlier slot to the earlier deadline is never worse for the maximum lateness. So you bubble-sort away inversions without harming the objective, ending at EDD. Zoom in on a two-job swap. Do not attempt a global argument without this local swap lens. ``` Inversion swap diagram (unit time) @@ -826,16 +832,16 @@ Key: the job with the tighter deadline stops being punished. Max lateness cannot increase by doing the swap. ``` -**Complexity**: $O(n\log n)$. +Complexity: $O(n\log n)$. -Again: sort once, then schedule in that order. The *do* is: implement as sort-by-deadline then compute completion times. The *don’t* is: simulate with complicated data structures when unit-time makes it simple. +Again: sort once, then schedule in that order. Implement as sort-by-deadline then compute completion times. Do not simulate with complicated data structures when unit-time makes it simple. -**Edge cases** +Edge cases: -* Same deadlines → any order ties. -* Different processing times $p_j$ requires a different model/algorithm. +- Same deadlines → any order ties. +- Varying processing times do not invalidate EDD for the classic single-machine maximum-lateness problem when all jobs are available at time zero. -EDD is optimal for **unit-time** jobs (or more generally, for minimizing maximum lateness on a single machine even with varying processing times under classic results, but the proof and model shift). Your note is a good guardrail: if job lengths differ, you must be precise about which theorem/problem variant you’re in. The *do* is: check assumptions (unit time? single machine? no release times?). The *don’t* is: apply EDD blindly when the model changes. +EDD remains optimal with varying positive processing times on one machine when all jobs are available at time zero and there are no precedence constraints. For adjacent inverted jobs A then B with $d_A>d_B$, swapping keeps their final completion time unchanged: B finishes earlier, while A's new lateness is less than B's old lateness. Thus the pair's maximum lateness cannot increase. Release times, multiple machines, and objectives such as the number of late jobs require separate reasoning. Lateness may be negative; tardiness is $\max(0,L_i)$. ``` Model checklist: @@ -845,14 +851,15 @@ Model checklist: - unit processing times? - objective is L_max (max lateness), not #on-time? -If any change: re-evaluate the greedy rule. +If machine count, release times, or the objective changes, re-evaluate the rule. +Varying positive processing times alone still permits EDD under the stated model. ``` -#### Huffman coding +### Huffman coding -**Pattern**: Merge-the-Two-Smallest Greedy +Pattern: Merge-the-Two-Smallest Greedy -Huffman coding is what happens when you take “be efficient” seriously: frequent symbols should be quick to write, rare symbols can be longer, and **prefix-free** means decoding is unambiguous (you never get stuck wondering where one codeword ends). The *do* is: think “short for common, long for rare.” The *don’t* is: chase clever-looking bitstrings by hand, this is a tree problem wearing a binary hat. +Huffman coding is what happens when you take “be efficient” seriously: frequent symbols should be quick to write, rare symbols can be longer, and prefix-free means decoding is unambiguous (you never get stuck wondering where one codeword ends). Think “short for common, long for rare.” Do not chase clever-looking bitstrings by hand, this is a tree problem wearing a binary hat. ``` Prefix-free means: no codeword is the prefix of another @@ -863,12 +870,17 @@ B: 10 C: 110 D: 111 -Bad (ambiguous): +Not prefix-free: +A: 0 +B: 01 <- "0" is a prefix of "01"; cannot decode 0 immediately + +Actually ambiguous: A: 0 -B: 01 <- "0" is a prefix of "01" +B: 01 +C: 1 <- 01 could mean B or A followed by C ``` -**Problem** +Problem: Given symbol frequencies $f_i>0$ with $\sum_i f_i=1$, find a prefix code minimizing average length: @@ -876,7 +888,7 @@ $$ \mathbb{E}[L]=\sum_i f_i L_i. $$ -The “why care” is baked into the formula: every extra bit on a high-frequency symbol costs you a lot, while extra bits on a rare symbol barely move the needle. Huffman is the clean, provably optimal way to trade length against frequency under the prefix constraint. The *do* is: keep seeing the objective as “weighted depth in a tree.” The *don’t* is: treat it like arithmetic on bitstrings; the tree is the real object. +The “why care” is baked into the formula: every extra bit on a high-frequency symbol costs you a lot, while extra bits on a rare symbol barely move the needle. Huffman is the clean, provably optimal way to trade length against frequency under the prefix constraint. Keep seeing the objective as “weighted depth in a tree.” Do not treat it like arithmetic on bitstrings; the tree is the real object. ``` Code tree view: @@ -888,7 +900,7 @@ Average length = sum(f_i * depth_i) So we want heavy leaves shallow, light leaves deep. ``` -**Baseline** +Baseline: Enumerate full binary trees → combinatorial explosion. Fixed-length codes (like all length $\lceil\log_2 k\rceil$) are easy but usually suboptimal. Greedy is the route to “optimal but still fast.” @@ -897,7 +909,7 @@ This baseline is important for two reasons: 1. It clarifies what “optimal” really means (best tree among all prefix trees). 2. It shows why we need structure: there are too many trees to brute-force. -Fixed-length coding is the “easy but wasteful” plan: it ignores the fact that some symbols happen way more than others. Huffman’s whole point is exploiting skew. The *do* is: compare “equal lengths” vs “frequency-aware lengths.” The *don’t* is: assume fixed-length is close to optimal unless frequencies are nearly uniform. +Fixed-length coding is the “easy but wasteful” plan: it ignores the fact that some symbols happen way more than others. Huffman’s whole point is exploiting skew. Compare “equal lengths” vs “frequency-aware lengths.” Do not assume fixed-length is close to optimal unless frequencies are nearly uniform. ``` Fixed-length code (k=8): @@ -907,11 +919,11 @@ But if one symbol is 50% of the data, making it length 1 can massively reduce average length. ``` -**Greedy rule** +Greedy rule: Repeatedly merge the two smallest weights $p$ and $q$ into $p+q$. -This is the heart of Huffman: if two symbols are the least frequent, you can “bury them deep” together with minimal pain. Merging is like saying: “in the final tree, these two will share a parent.” You then treat that parent as a combined pseudo-symbol with frequency `p+q` and repeat. The *do* is: see merging as “commit a deepest sibling pair.” The *don’t* is: merge something just because it looks convenient; it must be the **two smallest** for optimality. +This is the heart of Huffman: if two symbols are the least frequent, you can “bury them deep” together with minimal pain. Merging is like saying: “in the final tree, these two will share a parent.” You then treat that parent as a combined pseudo-symbol with frequency `p+q` and repeat. See merging as “commit a deepest sibling pair.” Do not merge something just because it looks convenient; it must be the two smallest for optimality. ``` Greedy visualization (weights only): @@ -926,7 +938,7 @@ Now weights: Repeat... ``` -**Why the cost increases by exactly $p+q$** +Why the cost increases by exactly $p+q$ When you merge two subtrees, every leaf under them increases depth by $1$, so the objective increases by the total frequency mass under them: @@ -934,7 +946,7 @@ $$ \Delta \mathbb{E}[L] = p+q. $$ -This is the “why this greedy is even measurable” moment: every merge has an exactly quantifiable cost. If you decide two subtrees will become siblings one level deeper, you’re charging every leaf underneath them one extra bit. The *do* is: remember “each merge adds its combined weight once.” The *don’t* is: think you need to build the entire codebook to compute the total cost, you can sum merge weights. +This is the “why this greedy is even measurable” moment: every merge has an exactly quantifiable cost. If you decide two subtrees will become siblings one level deeper, you’re charging every leaf underneath them one extra bit. Remember “each merge adds its combined weight once.” Do not think you need to build the entire codebook to compute the total cost, you can sum merge weights. ``` Depth bump diagram: @@ -960,19 +972,19 @@ Because each merge corresponds to "add 1 bit" to all leaves under the merged node. ``` -**Why safe** +Why safe Exchange argument: in an optimal prefix tree, the two least frequent symbols can be assumed to be deepest siblings. If a heavier symbol were deeper than a lighter one, swapping them would reduce cost. Collapsing the deepest sibling pair reduces the problem size and preserves optimality → induction. Here’s the flow that makes this feel human rather than mystical: 1. In any prefix tree, the deepest leaves pay the most bits. -2. So if *someone* must be deepest, it should be the least frequent symbols (cheapest to “punish” with extra length). +2. So if someone must be deepest, it should be the least frequent symbols (cheapest to “punish” with extra length). 3. Moreover, in a full binary prefix tree, the deepest leaves come in sibling pairs (they share a parent). 4. Therefore, you can assume the two smallest frequencies are deepest siblings in an optimal tree. 5. If you glue them into one pseudo-symbol of weight `p+q`, you’ve shrunk the problem while preserving optimal structure, then repeat. -The *do* is: connect “least frequent” → “deepest” → “siblings” → “merge.” The *don’t* is: accept “exchange argument” as a black box; it’s just “put heavier stuff shallower.” +Connect “least frequent” → “deepest” → “siblings” → “merge.” Do not accept “exchange argument” as a black box; it’s just “put heavier stuff shallower.” ``` Exchange intuition (swap reduces cost) @@ -1008,13 +1020,14 @@ Now solve smaller problem optimally. Expand back => optimal for original. ``` -**Implementation**: min-heap. +Implementation: min-heap. -The heap is how you make “always pick two smallest” fast. Each pop gives you the smallest remaining weight; push back the merged weight; repeat until one weight remains. The *do* is: treat it like repeatedly combining the cheapest two tasks. The *don’t* is: sort once and then do linear merges incorrectly, after every merge, the new `p+q` must re-enter the “smallest candidates” pool. +The heap is how you make “always pick two smallest” fast. Each pop gives you the smallest remaining weight; push back the merged weight; repeat until one weight remains. Treat it like repeatedly combining the cheapest two tasks. Do not sort once and then do linear merges incorrectly, after every merge, the new `p+q` must re-enter the “smallest candidates” pool. ``` Min-heap loop: +cost = 0 push all f_i while heap size > 1: p = pop_min() @@ -1023,16 +1036,17 @@ while heap size > 1: cost += p+q ``` -**Complexity**: $O(k\log k)$ for $k$ symbols. +Complexity: $O(k\log k)$ time and $O(k)$ tree/heap space for $k$ symbols. Frequencies may be positive counts instead of probabilities; scaling every weight equally does not change the optimal code. Keep child pointers at each merge to recover codewords; summing merge weights alone returns only the cost. -Interpretation: you do `k-1` merges; each merge does two pops and one push, each `log k`. The *do* is: remember it scales well even for large alphabets. The *don’t* is: assume it’s linear; the heap work matters (but it’s still very fast in practice). +Interpretation: you do `k-1` merges; each merge does two pops and one push, each `log k`. Remember it scales well even for large alphabets. Do not assume it’s linear; the heap work matters (but it’s still very fast in practice). -**Edge cases** +Edge cases: -* Ties → multiple optimal codebooks, same $\mathbb{E}[L]$. -* Two symbols → both length $1$. +- Ties → multiple optimal codebooks, same $\mathbb{E}[L]$. +- Two symbols → both length $1$. +- One symbol → an empty codeword has theoretical length zero if the decoder knows the message length; file formats often use one bit instead. An empty alphabet needs an explicit empty-input convention. -Ties are normal in real data (e.g., equal counts), and Huffman doesn’t care which tied pair you merge first, different trees, same optimal average length. With two symbols, there’s exactly one sensible prefix code: one gets `0`, the other gets `1`. The *do* is: expect non-uniqueness. The *don’t* is: treat different Huffman outputs as “wrong” if the cost matches. +Ties are normal in real data (e.g., equal counts), and Huffman doesn’t care which tied pair you merge first, different trees, same optimal average length. With two symbols, there’s exactly one sensible prefix code: one gets `0`, the other gets `1`. Expect non-uniqueness. Do not treat different Huffman outputs as “wrong” if the cost matches. ``` Two symbols: @@ -1043,11 +1057,11 @@ B: 1 Average length = 1 exactly. ``` -#### Maximum contiguous sum (Kadane) +### Maximum contiguous sum (Kadane) -**Pattern**: Discard-Negative-Prefix Greedy +Pattern: Discard-Negative-Prefix Greedy -This pattern is about emotional baggage: if the stuff you’ve accumulated so far is dragging you down, stop carrying it. Kadane’s algorithm works because *contiguous* means you’re forced to take whatever came immediately before, so the only real choice at each position is: **extend** the current block, or **start fresh** here. The *do* is: treat every index as a “restart checkpoint.” The *don’t* is: keep a running sum just because you started it earlier; sunk cost has no power here. +A contiguous subarray ending at the current position either extends a subarray ending at the previous position or starts at the current element. A negative preceding sum cannot improve an extension, which justifies starting again. ``` Contiguous block = one solid segment @@ -1060,11 +1074,11 @@ Either extend the segment ending at j-1 or start a new segment at j. ``` -**Problem** +Problem: -Given an array, pick one contiguous block maximizing its sum. +Given an array, pick one contiguous block maximizing its sum. Kadane's recurrence is also a dynamic-programming formulation; discarding a negative prefix explains the safe local choice. For the nonempty version, require a nonempty input and initialize both the ending-here sum and the best sum to the first element. -Why you should care: this shows up as “best streak,” “max profit over a period,” “strongest signal window,” “most energetic segment,” and a million other “find the hottest run” tasks. The *do* is: recognize it as “best window” not “best set.” The *don’t* is: mix it up with picking any subset (that’s a different problem). +Why you should care: this shows up as “best streak,” “max profit over a period,” “strongest signal window,” “most energetic segment,” and a million other “find the hottest run” tasks. Recognize it as “best window” not “best set.” Do not mix it up with picking any subset (that’s a different problem). ``` Subset vs contiguous (important!) @@ -1075,11 +1089,11 @@ subset: [x2, x7, x9] (can skip) Kadane is for contiguous. ``` -**Baseline** +Baseline: Try all $O(n^2)$ blocks (using prefix sums to evaluate quickly). -The baseline is the “I will brute-force my way to truth” approach: choose every possible start and end, compute sums, keep the best. It’s correct, and it teaches what “optimal” means, but it wastes time by recomputing overlapping windows. The *do* is: remember this baseline when you want to justify correctness. The *don’t* is: actually use it for large `n` unless you enjoy waiting. +The baseline is the “I will brute-force my way to truth” approach: choose every possible start and end, compute sums, keep the best. It’s correct, and it teaches what “optimal” means, but it wastes time by recomputing overlapping windows. Remember this baseline when you want to justify correctness. Do not actually use it for large `n` unless you enjoy waiting. ``` All blocks = all (L,R) pairs @@ -1091,13 +1105,13 @@ L=1: (1,1) (1,2) ... O(n^2) windows, tons of overlap. ``` -**Greedy rule** +Greedy rule: Never carry a negative-running prefix forward. If your running sum becomes negative, drop it, because adding a negative prefix to any future suffix only makes it worse. This is the “why” that makes the algorithm feel inevitable: if your current partial sum is negative, then for any future value `t`, `(negative sum) + t < t`. -So keeping that negative prefix can never help you build a best block that continues into the future. The *do* is: reset when the running sum goes below zero. The *don’t* is: reset when it merely “gets smaller”, smaller can still be useful if it stays non-negative. +So keeping that negative prefix can never help you build a best block that continues into the future. Reset when the running sum goes below zero. Do not reset when it merely “gets smaller”, smaller can still be useful if it stays non-negative. ``` Why negative prefixes are poison: @@ -1119,7 +1133,7 @@ $$ and track the best seen. -Here `E_j` is the best sum of a subarray that **must end at j**. That’s the key: “ending here” makes the choice local and greedy-friendly. Either you (1) start at `j` (take only `x[j]`), or (2) extend the best ending at `j-1`. There’s no third option if contiguity is mandatory. The *do* is: read the recurrence as “restart vs extend.” The *don’t* is: treat it like magic DP, this one is literally just the two possible ways to end at `j`. +Here `E_j` is the best sum among subarrays ending exactly at index `j`. Contiguity leaves two possibilities: take `x[j]` alone or add it to the best subarray ending at `j-1`. The recurrence chooses the better one. ``` Two options at j: @@ -1144,7 +1158,7 @@ At -6, the running sum becomes negative -> baggage dropped. Then the streak restarts at 3. ``` -**Proof sketch** +Proof sketch: A negative prefix cannot improve any future subarray sum, so discarding it is always safe. The loop maintains “best sum ending here” and “best overall.” @@ -1154,7 +1168,7 @@ The proof is basically a tight little logic loop: 2. If the best ending at `j-1` is negative, extending it only hurts, so starting fresh is optimal. 3. Therefore the recurrence computes the correct “best ending here,” and taking the maximum over all `j` gives the best overall. -The *do* is: anchor your reasoning on “must end at j.” The *don’t* is: try to prove it by enumerating all subarrays again, you’ll just reinvent the baseline. +Anchor your reasoning on “must end at j.” Do not try to prove it by enumerating all subarrays again, you’ll just reinvent the baseline. ``` Invariant view: @@ -1165,16 +1179,16 @@ BEST = max over all E_j seen so far Update keeps both truths correct each step. ``` -**Complexity**: $O(n)$ time, $O(1)$ space. +Complexity: $O(n)$ time, $O(1)$ space. -You scan once, keep two numbers (`E` and `BEST`). That’s the whole win: maximal information compression with minimal bookkeeping. The *do* is: implement it as a single pass. The *don’t* is: store all `E_j` unless you specifically need reconstruction. +You scan once, keep two numbers (`E` and `BEST`). That’s the whole win: maximal information compression with minimal bookkeeping. Implement it as a single pass. You do not need to store all `E_j` even to recover the chosen segment: track its candidate start and the best start/end indices whenever the best sum improves. -**Edge cases** +Edge cases: -* All negatives → answer is the least negative element (for non-empty requirement). -* If empty block allowed → initialize best at $0$. +- All negatives → answer is the least negative element (for non-empty requirement). +- If empty block allowed → initialize best at $0$. -These matter because Kadane’s “reset when negative” can otherwise trick you into returning 0 even when you’re required to pick *something*. The *do* is: decide up front whether “empty subarray” is allowed. The *don’t* is: mix conventions mid-solution. +These matter because Kadane’s “reset when negative” can otherwise trick you into returning 0 even when you’re required to pick something. Decide up front whether “empty subarray” is allowed. Do not mix conventions mid-solution. ``` All-negative example: [-5, -2, -8] @@ -1186,9 +1200,9 @@ Empty allowed: answer = 0 (pick empty block) ``` -### When Greedy Is Guaranteed: Matroids +## When Greedy Is Guaranteed: Matroids -There’s a clean world where greedy is always right for nonnegative weights: matroids. +A matroid consists of a finite ground set and a family of independent (allowed) subsets satisfying the properties below. For maximum-weight independence with non-negative weights, process elements in descending weight and keep each element that preserves independence. Informal rules: @@ -1196,20 +1210,20 @@ Informal rules: 2. Subsets of allowed sets are allowed. 3. Augmentation: if one allowed set is smaller than another, you can add something from the larger to the smaller and stay allowed. -Why you should care: matroids are a formal “certificate” that exchange arguments will succeed. They explain why “sort by weight and take what fits” works for MSTs (graphic matroid). +For a minimum-weight basis (a maximal independent set), process weights in ascending order until a basis is complete; weights may be negative. In the graphic matroid, independent sets are acyclic edge sets and bases are spanning trees of a connected graph, or spanning forests otherwise. This distinction matters: the empty set would minimize total positive weight if spanning were not required. Greedy can still work outside matroids (Dijkstra, Huffman), but then correctness depends on problem-specific structure rather than the general matroid guarantee. -### Common Failure Modes +## Common Failure Modes Greedy doesn’t always work. The failures are valuable because they teach you what a missing proof looks like. -#### Coin change with non-canonical coins +### Coin change with non-canonical coins Denominations ${1,3,4}$, target $6$. -* Greedy: $4+1+1$ → 3 coins. -* Optimal: $3+3$ → 2 coins. +- Greedy: $4+1+1$ → 3 coins. +- Optimal: $3+3$ → 2 coins. Lesson: “largest coin first” is not universally safe; it depends on special coin systems. @@ -1217,18 +1231,18 @@ Lesson: “largest coin first” is not universally safe; it depends on special Intervals $(1,10),(2,3),(4,5)$. -* Earliest start picks $(1,10)$ → blocks everything → 1 interval. -* Earliest finish picks $(2,3),(4,5)$ → 2 intervals. +- Earliest start picks $(1,10)$ → blocks everything → 1 interval. +- Earliest finish picks $(2,3),(4,5)$ → 2 intervals. Lesson: the key must match the exchange argument (“finishes earlier leaves more room”). -#### Fractional vs 0/1 knapsack +### Fractional vs 0/1 knapsack Greedy by value/weight ratio is optimal for fractional knapsack, but fails for 0/1. Example capacity $50$ with items $(10,60),(20,100),(30,120)$: -* Greedy ratio picks items 1 and 2 → value $160$. -* Optimal picks items 2 and 3 → value $220$. +- Greedy ratio picks items 1 and 2 → value $160$. +- Optimal picks items 2 and 3 → value $220$. Lesson: allowing fractions changes feasibility structure, and greedy safety can disappear. diff --git a/notes/math_set_relationship.md b/notes/math_set_relationship.md index 1bb62a2..f864ff8 100644 --- a/notes/math_set_relationship.md +++ b/notes/math_set_relationship.md @@ -1,10 +1,10 @@ -## Math Set Relationships +# Math Set Relationships -You begin with a small set of “building blocks,” and then you **systematically manufacture bigger collections** (pairs, sequences, subsets, orderings). The punchline is always the same: *what did you build, how many objects exist, and what does it cost to generate them?* If you care about algorithms, this is basically the bridge between “math objects” and “runtime explosions.” +Starting from a finite set, we can construct pairs, sequences, subsets, and orderings. For each construction, these notes explain what counts as a distinct result, how many results exist, and the time and storage required to generate them. -A do/don’t that will help as you read: **do** keep asking “what counts as a distinct object here?” (order matters? repetition allowed?), and **don’t** mix those rules mid-problem, most counting mistakes are just rule confusion. +Before counting, decide what makes two objects distinct: whether order matters and whether elements may be reused. Keep those rules fixed throughout the problem. -### Start with a set: the universe of building blocks +## Sets and Notation Let a finite set be: @@ -13,11 +13,13 @@ A = {a, b, c} |A| = n = 3 ``` -Almost everything below is about creating **new sets of objects** from `A`, then counting how many objects there are, then generating them efficiently. +Unless stated otherwise, $n$ and $k$ are non-negative integers, the elements of $A$ are distinct, and sampling without repetition requires $0 \le k \le n$. There is one empty subset and one empty sequence, so $0! = 1$ and choosing zero elements gives one result. Choosing more than $n$ elements without repetition gives no results. -This is the same move you do in programming all the time, turning a small input into a space of candidates. If the candidate space is small, brute force can work. If it’s huge (hello, $2^n$ and $n!$), you need smarter strategies. The math tells you *when “try everything” is doomed.* +Almost everything below is about creating new sets of objects from `A`, then counting how many objects there are, then generating them efficiently. -### Cartesian product: “pairing” choices (the product rule) +This is the same move you do in programming all the time, turning a small input into a space of candidates. If the candidate space is small, brute force can work. If it’s huge (hello, $2^n$ and $n!$), you need smarter strategies. The math tells you when “try everything” is doomed. + +## Cartesian Products When you see “choose one thing from here and one thing from there,” you’re in Cartesian-product land. It’s the formal version of nested loops: for each $x$ in $A$, loop over each $y$ in $B$. @@ -59,9 +61,9 @@ c | c,0 | c,1 | +-----+-----+ ``` -**Do** remember these are *ordered* pairs, $(a,0)\neq(0,a)$, and **don’t** treat it like “a set of two things.” Order is the entire point. +Do remember these are ordered pairs, $(a,0)\neq(0,a)$, and don’t treat it like “a set of two things.” Order is the entire point. -#### Counting +### Counting If $|A| = n$ and $|B| = m$ then: @@ -69,30 +71,30 @@ If $|A| = n$ and $|B| = m$ then: |A × B| = n·m ``` -#### Why it matters +### Why it matters This is the foundation of: -* counting multi-step choices (“choose x then choose y”) -* building tuples (records, coordinates) -* generating combinations/permutations via sequences (see below) +- counting multi-step choices (“choose x then choose y”) +- building tuples (records, coordinates) +- generating combinations/permutations via sequences (see below) This shows up constantly in real code: if you ever wrote two nested loops, joined two tables, or enumerated coordinate pairs on a grid, you were living in $A\times B$. The math just makes the “how many iterations is this?” question instantly answerable. -#### Algorithm + time +### Algorithm + time -To **enumerate** all pairs you must output $n\cdot m$ things: +To enumerate all pairs you must output $n\cdot m$ things: -* Time: **$\Theta(n\cdot m)$** -* Space: **$\Theta(1)$** extra (if streaming) or **$\Theta(n\cdot m)$** if storing +- Time: $\Theta(n\cdot m)$ +- Space: $\Theta(1)$ extra (if streaming) or $\Theta(n\cdot m)$ if storing -A practical do/don’t: **do** stream results when possible (generate and consume immediately), and **don’t** store the full product unless you truly need random access, memory can become your bottleneck before time does. +A practical do/don’t: do stream results when possible (generate and consume immediately), and don’t store the full product unless you truly need random access, memory can become your bottleneck before time does. -### Power set: “all subsets” (the mother of combinations) +## Power Sets -If Cartesian products feel like “two-loop land,” the power set is “every possible on/off choice.” This is where problems become *exponential* because you’re not picking one item, you’re deciding for each item whether it’s included. +If Cartesian products feel like “two-loop land,” the power set is “every possible on/off choice.” This is where problems become exponential because you’re not picking one item, you’re deciding for each item whether it’s included. -The power set $\mathcal{P}(A)$ is the set of **all subsets** of $A$. +The power set $\mathcal{P}(A)$ is the set of all subsets of $A$. For `A={a,b,c}`: @@ -120,7 +122,7 @@ Subset lattice (Boolean lattice): The power set is the search space behind “try all subsets” algorithms, feature selection, subset sum, knapsack-style brute force, and lots of graph subset problems. The lattice diagram is the map of that search space. -#### Counting +### Counting If $|A| = n$ then: @@ -144,20 +146,20 @@ a b c subset 1 1 1 {a,b,c} ``` -A do/don’t that saves headaches: **do** think “bitmask = subset” whenever you need to generate subsets efficiently, and **don’t** try to be “clever” by skipping the output cost, if the problem truly needs all subsets, $2^n$ is unavoidable. +A do/don’t that saves headaches: do think “bitmask = subset” whenever you need to generate subsets efficiently, and don’t try to be “clever” by skipping the output cost, if the problem truly needs all subsets, $2^n$ is unavoidable. -#### Algorithms + time +### Algorithms + time To enumerate all subsets, you must output $2^n$ subsets: -* Time: **$\Theta(n\cdot 2^n)$** if you build each subset explicitly (each subset may cost up to $n$ to construct) -* Space: **$\Theta(n)$** recursion/bitmask state (streaming) or **$\Theta(n\cdot 2^n)$** if storing all +- Time: $\Theta(n\cdot 2^n)$ if you build each subset explicitly (each subset may cost up to $n$ to construct) +- Space: $\Theta(n)$ recursion/bitmask state (streaming) or $\Theta(n\cdot 2^n)$ if storing all The big idea: output size dominates. If you’re generating all subsets, you’re not “being slow”, you’re paying the bill for the number of results you asked to print. -### Combinations: “choose k elements, order doesn’t matter” +## Combinations -Combinations are how you take the power set and say: “Okay, cool, but I only want subsets of a *specific size*.” This is a common “make the search space manageable” move: instead of *every* subset, you focus on one level of the lattice. +Combinations are how you take the power set and say: “Okay, cool, but I only want subsets of a specific size.” This is a common “make the search space manageable” move: instead of every subset, you focus on one level of the lattice. A $k$-combination is a subset of size $k$. @@ -175,13 +177,13 @@ If $|A| = n$: C(n,k) = "n choose k" = n! / (k!(n-k)!) ``` -In standard notation, this is +In standard notation, this is $$C(n,k)=\binom{n}{k}=\dfrac{n!}{k!(n-k)!}$$ -**Do** use $\binom{n}{k}$ when you write math (it’s clearer), and **don’t** forget the “order doesn’t matter” rule, if you start ordering the chosen elements, you’ve switched problems. +Do use $\binom{n}{k}$ when you write math (it’s clearer), and don’t forget the “order doesn’t matter” rule, if you start ordering the chosen elements, you’ve switched problems. -combinations are a “level” in the power-set lattice +Combinations form a level in the power-set lattice. For `A={a,b,c}`, `k=2`: @@ -189,7 +191,7 @@ For `A={a,b,c}`, `k=2`: level k=2: {a,b} {a,c} {b,c} ``` -**Power set is all combinations across all k:** +Power set is all combinations across all k: ``` 2^n = |𝒫(A)| = Σ_{k=0..n} C(n,k) @@ -199,35 +201,35 @@ Proper math: $$2^n = |\mathcal{P}(A)| = \sum_{k=0}^{n} \binom{n}{k}$$ -#### Algorithms + time +### Algorithms + time To enumerate all $k$-combinations you output $\binom{n}{k}$ objects: -* Time: **$\Theta(k \cdot C(n,k))$** (each output has $k$ items) -* Space: **$\Theta(k)$** recursion stack if streaming +- Time: $\Theta(k \cdot C(n,k))$ (each output has $k$ items) +- Space: $O(k)$ working storage for a generator that recurses over chosen positions. A literal take/skip recursion over all $n$ input positions can instead need $O(n)$ stack space. Common methods: -* recursive backtracking (“take / skip”) -* iterative lexicographic combination generation -* bitmask “next k-bit” tricks +- recursive backtracking (“take / skip”) +- iterative lexicographic combination generation +- bitmask “next k-bit” tricks -One practical tip: **do** pick your generation method based on what you need downstream (lexicographic order? streaming? constant extra memory?), and **don’t** assume “combinations are always small”, $\binom{n}{k}$ can still be huge around $k\approx n/2$. +One practical tip: do pick your generation method based on what you need downstream (lexicographic order? Streaming? Constant extra memory?), and don’t assume “combinations are always small”, $\binom{n}{k}$ can still be huge around $k\approx n/2$. -### Permutations: “arrangements, order matters” +## Permutations -Permutations are what happens when you stop treating a chosen set as a bag of items and start treating it like a *sequence*. In algorithm terms, you’ve moved from “which items?” to “in what order?” +Permutations are what happens when you stop treating a chosen set as a bag of items and start treating it like a sequence. In algorithm terms, you’ve moved from “which items?” to “in what order?” There are two closely related ideas: -#### (A) Permutations of all n elements +### (A) Permutations of all n elements All orderings of the entire set `A` (size $n$). Count: ``` -n! +n! ``` Example `A={a,b,c}`: @@ -236,7 +238,7 @@ Example `A={a,b,c}`: abc acb bac bca cab cba ``` -#### (B) k-permutations (arrangements of length k without repetition) +### (B) k-permutations (arrangements of length k without repetition) Ordered sequences of length $k$ drawn from $n$ distinct elements. @@ -246,13 +248,13 @@ Count: P(n,k) = n·(n-1)·...·(n-k+1) = n!/(n-k)! ``` -Proper math: +Proper math: $$P(n,k)=n(n-1)\cdots(n-k+1)=\dfrac{n!}{(n-k)!}$$ -A do/don’t that prevents classic mistakes: **do** decide early whether repetition is allowed; **don’t** mix “without repetition” formulas (like $P(n,k)$) with “with repetition” reasoning (like $n^k$). +A do/don’t that prevents classic mistakes: do decide early whether repetition is allowed; don’t mix “without repetition” formulas (like $P(n,k)$) with “with repetition” reasoning (like $n^k$). -#### permutations as “paths” of choices (multiplication rule) +### permutations as “paths” of choices (multiplication rule) For $n=3, k=2$: @@ -268,7 +270,7 @@ start This tree is the “nested loops in your head” picture: each level is a choice, and the number of leaves is the total count. -#### Relationship to combinations +### Relationship to combinations A $k$-combination becomes many $k$-permutations once you order it: @@ -276,35 +278,35 @@ A $k$-combination becomes many $k$-permutations once you order it: P(n,k) = C(n,k) · k! ``` -Proper math: +Proper math: $$P(n,k)=\binom{n}{k} k!$$ Because: -* choose the $k$ elements (order-free): $\binom{n}{k}$ -* order them: $k!$ +- choose the $k$ elements (order-free): $\binom{n}{k}$ +- order them: $k!$ -#### Algorithms + time +### Algorithms + time To enumerate all permutations of $n$ items: -* Time: **$\Theta(n \cdot n!)$** ($n$ work per permutation) -* Space: **$\Theta(n)$** recursion or in-place swaps +- Time: $\Theta(n \cdot n!)$ ($n$ work per permutation) +- Space: $\Theta(n)$ recursion or in-place swaps Classic algorithms: -* Heap’s algorithm (efficient swaps) -* next_permutation (lexicographic) -* backtracking swap recursion +- Heap’s algorithm (efficient swaps) +- next_permutation (lexicographic) +- backtracking swap recursion -**Do** use in-place swapping if you want low memory, and **don’t** generate permutations “just to test something” unless you’re sure $n$ is tiny, $n!$ grows so fast it’s basically a jump scare. +In-place swapping limits working memory, but explicitly producing every permutation still takes factorially many outputs. Estimate the output size before choosing enumeration, and use streaming when results can be processed one at a time. -### “With repetition” variants (multisets and strings) +## Selection with Repetition -So far we assumed you **don’t reuse** elements. If reuse is allowed, the counting flips in a really clean way: “no repetition” tends to create factorials and falling factorials; “repetition allowed” tends to create powers and stars-and-bars. +Subsets and permutations without repetition use each element at most once. Cartesian products and powers allow the same value in different positions. If reuse is allowed, the counting flips in a really clean way: “no repetition” tends to create factorials and falling factorials; “repetition allowed” tends to create powers and stars-and-bars. -#### Cartesian product becomes “strings of length k” +### Cartesian product becomes “strings of length k” If you choose $k$ positions and each position can be any of $n$ symbols, you get: @@ -313,7 +315,7 @@ A^k = A × A × ... × A (k times) |A^k| = n^k ``` -Proper Math: +Proper Math: $$A^k = \underbrace{A\times A\times\cdots\times A}_{k\text{ times}}$$ @@ -321,7 +323,7 @@ $$|A^k|=n^k$$ This is exactly “all $k$-length sequences over $A$” (like passwords). -#### Combinations with repetition (multisets) +### Combinations with repetition (multisets) Number of size-$k$ multisets from $n$ types: @@ -329,9 +331,9 @@ Number of size-$k$ multisets from $n$ types: C(n+k-1, k) ``` -That is $\binom{n+k-1}{k}$ (a “stars and bars” result). **Do** picture $k$ identical picks distributed among $n$ bins; **don’t** treat this like normal combinations, repetition changes the geometry. +For $n \ge 1$, that is $\binom{n+k-1}{k}$ (a “stars and bars” result). Place $k$ stars and $n-1$ separators in a row: the numbers of stars between separators specify the multiplicities of the $n$ types. If $n=0$, only the empty multiset ($k=0$) exists. Do picture $k$ identical picks distributed among $n$ bins; don’t treat this like normal combinations, repetition changes the geometry. -#### Permutations with repetition +### Permutations with repetition If you have $n$ symbols and length $k$, order matters and repetition allowed: @@ -339,11 +341,13 @@ If you have $n$ symbols and length $k$, order matters and repetition allowed: n^k ``` -### One unifying mental model: “objects = functions” +This use of “permutations with repetition” means unrestricted sequences. Arranging a fixed multiset is a different problem: if $N$ items have multiplicities $m_1,\ldots,m_r$ with sum $N$, the number of distinct orderings is $N!/(m_1!\cdots m_r!)$. For example, `aab` has three orderings: `aab`, `aba`, and `baa`. + +## Viewing the Constructions as Functions -This section is the “snap everything into place” moment. If you ever feel lost, switching to the function viewpoint usually makes the counting obvious: *how many ways can I assign labels to inputs?* +This section is the “snap everything into place” moment. If you ever feel lost, switching to the function viewpoint usually makes the counting obvious: how many ways can I assign labels to inputs? -#### Subset = function to {0,1} +### Subset = function to {0,1} A subset $S \subseteq A$ is equivalent to an indicator function: @@ -354,7 +358,7 @@ f(x)=1 if x∈S else 0 Number of such functions is $2^n$ → power set size. -#### k-length sequence = function from positions to A +### k-length sequence = function from positions to A A length-$k$ sequence is: @@ -364,7 +368,7 @@ g: {1..k} → A Number is $n^k$ → Cartesian power. -#### Permutation = bijection on {1..n} +### Permutation as a bijection from positions to elements A permutation is a bijection: @@ -376,15 +380,15 @@ Number is $n!$. So: -* **power set** = all ${0,1}$-labelings of $A$ -* **Cartesian powers** = all $k$-position labelings by $A$ -* **permutations** = all one-to-one labelings of positions by $A$ +- power set = all $\{0,1\}$-labelings of $A$ +- Cartesian powers = all $k$-position labelings by $A$ +- permutations = all one-to-one labelings of positions by $A$ -A do/don’t here: **do** use “functions” when you’re stuck; **don’t** overcomplicate it with new symbols, this is meant to be the simplest mental model. +A do/don’t here: do use “functions” when you’re stuck; don’t overcomplicate it with new symbols, this is meant to be the simplest mental model. -#### Complexity reality check: output size dominates +### Complexity reality check: output size dominates -If you *generate* these objects, the **minimum time** is at least the number of outputs. +If you generate these objects, the minimum time is at least the number of outputs. The bounds below assume nonempty outputs are copied explicitly and each element takes constant time to copy. For $n=0$ or $k=0$, emitting the single empty result costs constant time. Streaming saves storage for the accumulated results, but it does not remove the output time. So: @@ -399,21 +403,21 @@ So: This is why many problems that ask you to “try all subsets” are exponential: the search space literally has $2^n$ candidates. -One last do/don’t that matters in real projects: **do** treat these counts like early warning signs (a design review for your algorithm), and **don’t** wait until you’ve implemented everything to realize you built an $n!$ machine. +One last do/don’t that matters in real projects: do treat these counts like early warning signs (a design review for your algorithm), and don’t wait until you’ve implemented everything to realize you built an $n!$ machine. -### Quick “how to choose the right formula” cheat-sheet +## Choosing a Counting Formula Ask two questions: -#### Q1: Does order matter? +### Q1: Does order matter? -* **No** → combinations / subsets / multisets -* **Yes** → permutations / sequences / tuples +- No → combinations / subsets / multisets +- Yes → permutations / sequences / tuples -#### Q2: Can you reuse elements? +### Q2: Can you reuse elements? -* **No repetition** → factorial / falling factorial -* **Repetition allowed** → powers / stars-and-bars +- No repetition → factorial / falling factorial +- Repetition allowed → powers / stars-and-bars Decision table: @@ -423,15 +427,15 @@ Order matters P(n,k)=n!/(n-k)! n^k Order doesn't matter C(n,k)=n!/(k!(n-k)!) C(n+k-1,k) ``` -A quick “care factor”: this table is basically the fastest way to turn a word problem into a correct formula. The moment you decide “order?” and “reuse?”, you’ve already done the hard part. +Use the table by deciding first whether order matters and then whether elements can be reused. Those two assumptions distinguish the four counting problems. -### why these appear in algorithms +## Why these appear in algorithms -* **Power set / combinations** show up in: subset-sum, knapsack variants, feature selection, graph vertex subsets, brute-force optimization. -* **Permutations** show up in: traveling salesman, scheduling, ordering constraints, anagrams, backtracking search. -* **Cartesian products** show up in: nested loops, join operations in databases, grid/coordinate enumeration, state spaces. +- Power set / combinations show up in: subset-sum, knapsack variants, feature selection, graph vertex subsets, brute-force optimization. +- Permutations show up in: traveling salesman, scheduling, ordering constraints, anagrams, backtracking search. +- Cartesian products show up in: nested loops, join operations in databases, grid/coordinate enumeration, state spaces. -If you want a fun self-check: anytime you see an algorithm that “branches” at each element with a yes/no choice, your brain should whisper $2^n$. Anytime you see “try all orders,” it should whisper $n!$. Those whispers are your runtime instincts. +A yes/no branch for each of $n$ elements suggests $2^n$ candidate subsets. Trying all orders of $n$ distinct elements suggests $n!$ candidates. Constraints may prune the search, but these counts explain its initial size. A tiny “big picture” diagram: how they build on each other diff --git a/notes/matrices.md b/notes/matrices.md index 54ed559..6fb5676 100644 --- a/notes/matrices.md +++ b/notes/matrices.md @@ -1,12 +1,14 @@ -## Matrices and 2D Grids +# Matrices and 2D Grids Matrices represent images, game boards, and maps. Many classic problems reduce to transforming matrices, traversing them, or treating grids as graphs for search. -### Conventions +## Conventions -**Rows indexed $0..R-1$, columns $0..C-1$; cell $(r,c)$.** +Rows indexed $0..R-1$, columns $0..C-1$; cell $(r,c)$. -Rows increase **down**, columns increase **right**. Think “top-left is $(0,0)$”, not a Cartesian origin. +Assume a rectangular matrix: every row has the same number of columns. Handle empty inputs before reading the first row. In pseudocode, `a..b` includes both endpoints and is empty if its bounds oppose the stated step direction. Variables initialized to example dimensions should be replaced with the actual dimensions in a general implementation. + +Rows increase down, columns increase right. Think “top-left is $(0,0)$”, not a Cartesian origin. Visual index map (example $R=6$, $C=8$; each cell labeled $rc$): @@ -14,41 +16,39 @@ Visual index map (example $R=6$, $C=8$; each cell labeled $rc$): c → 0 1 2 3 4 5 6 7 r ↓ +----+----+----+----+----+----+----+----+ 0 | 00 | 01 | 02 | 03 | 04 | 05 | 06 | 07 | - +----+----+----+----+----+----+----+----+ + +----+----+----+----+----+----+----+----+ 1 | 10 | 11 | 12 | 13 | 14 | 15 | 16 | 17 | - +----+----+----+----+----+----+----+----+ + +----+----+----+----+----+----+----+----+ 2 | 20 | 21 | 22 | 23 | 24 | 25 | 26 | 27 | - +----+----+----+----+----+----+----+----+ + +----+----+----+----+----+----+----+----+ 3 | 30 | 31 | 32 | 33 | 34 | 35 | 36 | 37 | - +----+----+----+----+----+----+----+----+ + +----+----+----+----+----+----+----+----+ 4 | 40 | 41 | 42 | 43 | 44 | 45 | 46 | 47 | - +----+----+----+----+----+----+----+----+ + +----+----+----+----+----+----+----+----+ 5 | 50 | 51 | 52 | 53 | 54 | 55 | 56 | 57 | - +----+----+----+----+----+----+----+----+ + +----+----+----+----+----+----+----+----+ ``` -Handy conversions (for linearization / array-of-arrays): +For $C>0$, the following conversions describe a row-major flattened representation. An array of row objects need not itself occupy one contiguous block of memory: -* Linear index: $\text{id}=r\cdot C+c$. -* From id: $r=\lfloor \text{id}/C \rfloor$, $c=\text{id}\bmod C$. -* Row-major scan order (common in problems): for $r$ in $0..R-1$, for $c$ in $0..C-1$. +- Linear index: $\text{id}=r\cdot C+c$. +- From id: $r=\lfloor \text{id}/C \rfloor$, $c=\text{id}\bmod C$. +- Row-major scan order (common in problems): for $r$ in $0..R-1$, for $c$ in $0..C-1$. -**Row-major vs column-major arrows (same $3\times 6$ grid):** +Row-major and column-major visit order (same $3\times 6$ grid): ``` -Row-major (r, then c): Column-major (c, then r): -→ → → → → → ↓ ↓ ↓ - ↓ ↓ ↓ ↓ -← ← ← ← ← ← ↓ ↓ ↓ -↓ ↓ ↓ ↓ -→ → → → → → ↓ ↓ ↓ +Row-major visit numbers: Column-major visit numbers: + 0 1 2 3 4 5 0 3 6 9 12 15 + 6 7 8 9 10 11 1 4 7 10 13 16 +12 13 14 15 16 17 2 5 8 11 14 17 ``` -**Neighborhoods: $\mathbf{4}$-dir $\Delta={(-1,0),(1,0),(0,-1),(0,1)}$; $\mathbf{8}$-dir adds diagonals.** +Neighborhoods: $\mathbf{4}$-dir $\Delta={(-1,0),(1,0),(0,-1),(0,1)}$; $\mathbf{8}$-dir adds diagonals. The offsets $(\Delta r,\Delta c)$ are applied as $(r+\Delta r,\ c+\Delta c)$. -**4-neighborhood (“+”):** +4-neighborhood (“+”): ``` # @@ -59,7 +59,7 @@ The offsets $(\Delta r,\Delta c)$ are applied as $(r+\Delta r,\ c+\Delta c)$. (r+1,c) ``` -**8-neighborhood (“×” adds diagonals):** +8-neighborhood (“×” adds diagonals): ``` (r-1,c-1) (r-1,c) (r-1,c+1) @@ -83,13 +83,13 @@ dr8 = [-1,-1,-1, 0, 0, 1, 1, 1] dc8 = [-1, 0, 1,-1, 1,-1, 0, 1] ``` -**Boundary checks** (always guard neighbors): +Boundary checks (always guard neighbors): ``` 0 ≤ nr < R and 0 ≤ nc < C ``` -**Edge/inside intuition:** +Edge/inside intuition: ``` out of bounds @@ -104,15 +104,15 @@ dc8 = [-1, 0, 1,-1, 1,-1, 0, 1] └─────────────────┘ ``` -### Basic Operations (Building Blocks) +## Basic Operations (Building Blocks) -#### Transpose +### Transpose Swap across the main diagonal: $A_{r,c} \leftrightarrow A_{c,r}$ (square). For non-square, result shape is $C\times R$. -**Example inputs and outputs** +Example inputs and outputs: -*Example 1 (square)* +Example 1 (square) $$ A = \begin{bmatrix} @@ -129,13 +129,13 @@ A^{\mathsf{T}} = \end{bmatrix} $$ -**Mathematical formula (3×3)** +Mathematical formula (3×3) $$ (A^T)_{r,c}=A_{c,r},\quad 0\le r,c<3 $$ -**Pseudocode (square, in-place)** +Pseudocode (square, in-place) ``` n = 3 # for this example; generalize to n = size @@ -144,10 +144,10 @@ for r in 0..n-1: swap A[r][c], A[c][r] ``` -*Example 2 (rectangular)* +Example 2 (rectangular) $$ -\text{Input: } \quad +\text{Input: } \quad A = \begin{bmatrix} 1 & 2 & 3 \\ 4 & 5 & 6 @@ -156,7 +156,7 @@ A = \begin{bmatrix} $$ $$ -\text{Output: } \quad +\text{Output: } \quad A^{\mathsf{T}} = \begin{bmatrix} 1 & 4 \\ 2 & 5 \\ @@ -165,13 +165,13 @@ A^{\mathsf{T}} = \begin{bmatrix} \ (3 \times 2) $$ -**Mathematical formula (2×3 → 3×2)** +Mathematical formula (2×3 → 3×2) $$ (A^{\mathsf T})_{r,c}=A_{c,r},\quad 0\le r<3,\ 0\le c<2 $$ -**Pseudocode (rectangular, new matrix)** +Pseudocode (rectangular, new matrix) ``` R, C = 2, 3 @@ -181,20 +181,20 @@ for r in 0..R-1: B[c][r] = A[r][c] ``` -**How it works** +How it works: Iterate pairs once and swap. For square matrices, can be in-place by visiting only $c>r$. -* Time: $O(R\cdot C)$ -* Space: $O(1)$ in-place (square), else $O(R\cdot C)$ to allocate +- Time: $O(R\cdot C)$ +- Space: $O(1)$ in-place (square), else $O(R\cdot C)$ to allocate -#### Reverse Rows (Horizontal Flip) +### Reverse Rows (Horizontal Flip) Reverse each row left $\leftrightarrow$ right. -**Example inputs and outputs** +Example inputs and outputs: -*Example* +Example $$ \text{Input: } @@ -210,13 +210,13 @@ $$ \end{bmatrix} $$ -**Mathematical formula (2×3)** +Mathematical formula (2×3) $$ B_{r,c}=A_{r,\ C-1-c},\quad 0\le r<2,\ 0\le c<3 $$ -**Pseudocode (in-place)** +Pseudocode (in-place) ``` R, C = 2, 3 @@ -225,16 +225,16 @@ for r in 0..R-1: swap A[r][c], A[r][C-1-c] ``` -* Time: $O(R\cdot C)$ -* Space: $O(1)$ +- Time: $O(R\cdot C)$ +- Space: $O(1)$ -#### Reverse Columns (Vertical Flip) +### Reverse Columns (Vertical Flip) Reverse each column top $\leftrightarrow$ bottom. -**Example inputs and outputs** +Example inputs and outputs: -*Example* +Example $$ \text{Input: } @@ -252,13 +252,13 @@ $$ \end{bmatrix} $$ -**Mathematical formula (3×3)** +Mathematical formula (3×3) $$ B_{r,c}=A_{R-1-r,\ c},\quad 0\le r,c<3 $$ -**Pseudocode (in-place)** +Pseudocode (in-place) ``` R, C = 3, 3 @@ -267,30 +267,30 @@ for r in 0..(R//2 - 1): swap A[r][c], A[R-1-r][c] ``` -* Time: $O(R\cdot C)$ -* Space: $O(1)$ +- Time: $O(R\cdot C)$ +- Space: $O(1)$ -### Rotations (Composed from Basics) +## Rotations (Composed from Basics) -Use transpose + reversals for square in-place rotations; rectangular rotations produce new shape $(R\times C)\to(C\times R)$. +Use transpose and reversal for square in-place quarter turns. A 90° or 270° turn changes an $R\times C$ matrix into a $C\times R$ matrix; the straightforward rectangular algorithm allocates a new result. A 180° rotation preserves the shape and can be done in place for any rectangle. -#### 90° Clockwise (CW) +### 90° Clockwise (CW) Transpose, then reverse each row. -**Example inputs and outputs** +Example inputs and outputs: -*Example 1 (3×3)* +Example 1 (3×3) $$ -\text{Input: } +\text{Input: } \begin{bmatrix} 1 & 2 & 3 \\ 4 & 5 & 6 \\ 7 & 8 & 9 \end{bmatrix} \quad\Rightarrow\quad -\text{Output: } +\text{Output: } \begin{bmatrix} 7 & 4 & 1 \\ 8 & 5 & 2 \\ @@ -298,13 +298,13 @@ $$ \end{bmatrix} $$ -**Mathematical formula (n×n)** +Mathematical formula (n×n) $$ B_{r,c}=A_{n-1-c,\ r},\quad 0\le r,ct$): $(b,c)$ for $c=rgt-1,\ldots,left$ (decreasing). -* Left edge (if $rgt>left$): $(r,left)$ for $r=b-1,\ldots,t+1$ (decreasing). +- Top edge: $(t,c)$ for $c=left,\ldots,rgt$. +- Right edge: $(r,rgt)$ for $r=t+1,\ldots,b$. +- Bottom edge (if $b>t$): $(b,c)$ for $c=rgt-1,\ldots,left$ (decreasing). +- Left edge (if $rgt>left$): $(r,left)$ for $r=b-1,\ldots,t+1$ (decreasing). Concatenate these per layer until all elements are visited. -**Pseudocode (loops)** +Pseudocode (loops) ``` R, C = dims(A) @@ -630,20 +632,20 @@ while top <= bottom and left <= right: left += 1 ``` -**How it works** +How it works: Maintain top, bottom, left, right. Walk edges in order; after each edge, move the corresponding bound inward. -* Time: $O(R\cdot C)$ -* Space: $O(1)$ beyond output. +- Time: $O(R\cdot C)$ +- Space: $O(1)$ beyond output. -#### Diagonal Order (r+c layers) +### Diagonal Order (r+c layers) Visit cells grouped by $s=r+c$; alternate direction per diagonal to keep locality if desired. -**Example inputs and outputs** +Example inputs and outputs: -*Example* +Example $$ \text{Input: } @@ -655,7 +657,7 @@ d & e & f \text{One order: } a, b, d, e, c, f $$ -**Mathematical formulation (general $R\times C$)** +Mathematical formulation (general $R\times C$) Let $s=r+c$. For $s=0,1,\dots,R+C-2$, define @@ -680,7 +682,7 @@ $$ (This parity choice reproduces the example order $a,b,d,e,c,f$ for $R=2,C=3$.) -**Pseudocode (loops, alternating direction)** +Pseudocode (loops, alternating direction) ``` R, C = dims(A) @@ -691,29 +693,29 @@ for s in 0..(R + C - 2): r_hi = min(R - 1, s) if s % 2 == 0: - # even s: go downward-left (decreasing r) + # even s: go upward-right (decreasing r) for r in r_hi..r_lo step -1: c = s - r out.append(A[r][c]) else: - # odd s: go upward-right (increasing r) + # odd s: go downward-left (increasing r) for r in r_lo..r_hi: c = s - r out.append(A[r][c]) ``` -* Time: $O(R\cdot C)$ -* Space: $O(1)$ +- Time: $O(R\cdot C)$ +- Space: $O(1)$ beyond the $O(RC)$ output list. -### Grids as Graphs +## Grids as Graphs Each cell is a node; edges connect neighboring walkable cells. -**Grid-as-graph view (4-dir edges).** Each cell is a node; edges connect neighbors that are “passable”. Great for BFS shortest paths on unweighted grids. +Grid-as-graph view (4-dir edges). Each cell is a node; edges connect neighbors that are “passable”. Great for BFS shortest paths on unweighted grids. -**Example map (walls `#`, free `.`, start `S`, target `T`).** +Example map (walls `#`, free `.`, start `S`, target `T`). -Left: the map. Right: BFS distances (4-dir) from `S` until `T` is reached. +First the map, then BFS distances from `S` after traversing all reachable cells. Digits show distance modulo 10; `X` marks the target. ``` Original Map: @@ -735,15 +737,15 @@ BFS layers (distance mod 10): Legend: walls (#), goal reached (X) ``` -BFS explores in **expanding “rings”**; with 4-dir edges, each step increases Manhattan distance by 1 (unless blocked). Time $O(RC)$, space $O(RC)$ with a visited matrix/queue. +BFS layers increase shortest-path distance from the source by one. A particular grid move can increase or decrease Manhattan distance, and walls can force detours. For this map the target distance is 28. Time and auxiliary space are $O(RC)$ with a visited matrix and queue. -**Obstacles / costs / diagonals.** +Obstacles / costs / diagonals. -* Obstacles: skip neighbors that are `#` (or where cost is $\infty$). -* Weighted grids: Dijkstra / 0-1 BFS on the same neighbor structure. -* 8-dir with Euclidean costs: use $1$ for orthogonal moves and $\sqrt{2}$ for diagonals (A\* often pairs well here with an admissible heuristic). +- Obstacles: skip neighbors that are `#` (or where cost is $\infty$). +- Weighted grids: use Dijkstra for non-negative costs, or 0–1 BFS with a deque when every edge costs either zero or one. +- 8-dir with Euclidean costs: use $1$ for orthogonal moves and $\sqrt{2}$ for diagonals (A\* often pairs well here with an admissible heuristic). -**Common symbols:** +Common symbols: ``` . = free cell # = wall/obstacle @@ -751,13 +753,13 @@ S = start T = target/goal V = visited * = on current path / frontier ``` -#### BFS Shortest Path (Unweighted) +### BFS Shortest Path (Unweighted) Find the minimum steps from S to T. -**Example inputs and outputs** +Example inputs and outputs: -*Example* +Example $$ \text{Grid (0 = open, 1 = wall), } S = (0,0), T = (2,3) @@ -773,18 +775,18 @@ S & 0 & 1 & 0 \\ \text{Output: distance } = 5 $$ -**How it works** +How it works: Push S to a queue, expand in 4-dir layers, track distance/visited; stop when T is dequeued. -* Time: $O(R\cdot C)$ -* Space: $O(R\cdot C)$ +- Time: $O(R\cdot C)$ +- Space: $O(R\cdot C)$ -#### Connected Components (Islands) +### Connected Components (Islands) -Count regions of ‘1’s via DFS/BFS. +Count regions of ‘1’s via DFS/BFS using four-directional adjacency. Here `1` means land, unlike the preceding shortest-path example where `1` means wall. With eight-directional adjacency, the diagonal cells in this example form one island instead of two. -**Example inputs and outputs** +Example inputs and outputs: $$ \text{Input: } @@ -797,20 +799,20 @@ $$ \text{Output: } 2 \ \text{islands} $$ -**How it works** +How it works: Scan cells; when an unvisited ‘1’ is found, flood it (DFS/BFS) to mark the whole island. -* Time: $O(R\cdot C)$ -* Space: $O(R\cdot C)$ worst-case +- Time: $O(R\cdot C)$ +- Space: $O(R\cdot C)$ worst-case -### Backtracking on Grids +## Backtracking on Grids -#### Word Search (Single Word) +### Word Search (Single Word) Find a word by moving to adjacent cells (4-dir), using each cell once per path. -**Example inputs and outputs** +Example inputs and outputs: $$ \text{Board: } @@ -825,7 +827,7 @@ A & D & E & E \text{Output: true} $$ -**Mathematical formulation (general)** +Mathematical formulation (general) Let the word be $W=W_0W_1\cdots W_{L-1}$ and the grid be $G\in\Sigma^{R\times C}$. We seek a path $P=\big((r_0,c_0),\ldots,(r_{L-1},c_{L-1})\big)$ such that @@ -838,7 +840,7 @@ $$ \end{aligned} $$ -**Instantiation for the example (one valid path)** +Instantiation for the example (one valid path) $$ P=\big((0,0),(0,1),(0,2),(1,2),(2,2),(2,1)\big) @@ -846,17 +848,21 @@ $$ gives $A\to B\to C\to C\to E\to D = \text{"ABCCED"}$. -**Pseudocode (DFS with loops over starts and 4-neighbors)** +Pseudocode (DFS with loops over starts and 4-neighbors) ``` R, C = dims(board) L = len(word) +if L == 0: + return true +if R == 0 or C == 0 or L > R * C: + return false visited = array(R, C, fill=false) dr = [1, -1, 0, 0] dc = [0, 0, 1, -1] def dfs(r, c, i): - if r < 0 or r >= R or c < 0 or c >= C: + if r < 0 or r >= R or c < 0 or c >= C: return false if visited[r][c] or board[r][c] != word[i]: return false @@ -868,6 +874,7 @@ def dfs(r, c, i): nr = r + dr[k] nc = c + dc[k] if dfs(nr, nc, i + 1): + visited[r][c] = false return true visited[r][c] = false return false @@ -880,20 +887,20 @@ for r in 0..R-1: return false ``` -**How it works** +How it works: -From each starting match, DFS to next char; mark visited (temporarily), backtrack on failure. +Start from each cell matching the first character and search for the next character among adjacent cells. Mark cells only for the current path and restore those marks when returning. A failed attempt from one starting cell must not block another attempt. -* Time: up to $O(R\cdot C\cdot b^{L})$ (branching $b\in [3,4]$, word length $L$) -* Space: $O(L)$ +- Time: $O(RC\,4^L)$ is a simple upper bound. After the first move there are at most three forward choices because the previous cell cannot be reused, giving the tighter conventional bound $O(RC\,3^L)$. +- Auxiliary space: $O(RC+L)$ for the shown visited matrix and recursion stack. Marking cells temporarily in place can reduce this to $O(L)$, provided all original values are restored. Pruning: early letter mismatch; frequency precheck; prefix trie when searching many words. -#### Crossword-style Fill (Multiple Words) +### Crossword-style Fill (Multiple Words) Place words to slots with crossings; verify consistency at intersections. -**Mathematical formulation** +Mathematical formulation Let $S$ be the set of slots (across/down). Each slot $s\in S$ has a length $\ell(s)$ and ordered cell coordinates $\mathrm{cells}(s) = \big((r_0,c_0),\ldots,(r_{\ell(s)-1},c_{\ell(s)-1})\big)$. @@ -904,7 +911,7 @@ Find an assignment $f:S\to D$ such that, for all $s\in S$, $$ f(s)\in D_{\ell(s)}\quad\text{and}\quad -\forall i\ (P_s[i]\neq _ \Rightarrow f(s)[i]=P_s[i]), +\forall i\ (P_s[i]\neq \text{\_} \Rightarrow f(s)[i]=P_s[i]), $$ and for every intersection between slots $s$ at index $i$ and $t$ at index $j$, @@ -915,21 +922,19 @@ $$ (Optionally enforce all-different: $s\neq t \Rightarrow f(s)\neq f(t)$.) -**Pseudocode (backtracking with loops, MRV + trie filtering)** +Domains hold words of the correct length that match fixed letters. Forward checking removes words incompatible with a newly assigned crossing; every removal is recorded so it can be undone. Minimum remaining values (MRV) chooses the unassigned slot with the fewest currently available words, breaking ties by the largest number of unassigned neighbors. + +Pseudocode (backtracking with dynamic MRV and forward checking) ``` # Preprocess slots = extract_slots(grid) # with cells(s) and pattern P_s -trie = build_trie(dictionary) # for prefix/length checks # Build initial domains from patterns and lengths domains = dict() for s in slots: domains[s] = { w in dictionary | len(w) == len(s) and matches_pattern(w, P_s) } -# Order slots: Most-Restricted-Variable (smallest domain first) -slots.sort_by(|domains[s]| ascending, tiebreak by number_of_intersections) - used = set() # if words must be unique assignment = dict() @@ -941,8 +946,9 @@ def consistent(s, w): return true def forward_check_update_domains(s, w, removed): - # reduce neighbor domains by letter constraints from placing w at s for each intersection (s,i) with (t,j): + if t in assignment: + continue for each v in copy(domains[t]): if v[j] != w[i]: domains[t].remove(v); removed.append((t, v)) @@ -955,11 +961,12 @@ def backtrack(idx): if idx == len(slots): return true - s = slots[idx] + remaining = [slot for slot in slots if slot not in assignment] + s = min(remaining, key=(size(domains[slot] - used), -unassigned_neighbors(slot))) # iterate candidates; optionally skip ones already used for w in iterate(domains[s]): - if w in used: + if w in used: continue if not consistent(s, w): continue @@ -969,7 +976,7 @@ def backtrack(idx): removed = [] forward_check_update_domains(s, w, removed) - if backtrack(idx + 1): + if all(domains[slot] - used for slot in slots if slot not in assignment) and backtrack(idx + 1): return true undo_forward_check(removed) @@ -985,10 +992,9 @@ else: return failure ``` -**How it works** - -Backtrack over slot assignments; use a trie for prefix feasibility; order by most constrained slot first. - -* Time: exponential in slots; strong pruning and good heuristics are important. +How it works: +Backtrack over slot assignments, recomputing the most constrained slot after each placement. The shown version enforces distinct words through `used`; remove that restriction consistently if reuse is allowed. A trie is an optional way to generate candidates matching a pattern; it is not needed by this explicit-domain version. +- Time: exponential in the number of slots in the worst case. With $S$ slots and at most $D$ candidates per slot, there can be $D^S$ assignments before accounting for constraint-checking costs. +- Space: the initial domains, the current assignment, and the reversible domain-removal log, plus an $O(S)$ recursion stack. diff --git a/notes/searching.md b/notes/searching.md index eee025d..fd16e34 100644 --- a/notes/searching.md +++ b/notes/searching.md @@ -1,80 +1,82 @@ -## Searching +# Searching Searching is the task of finding whether a particular value exists in a collection and, if it does, where it lives (its index, pointer, node, or associated value). It shows up everywhere: checking if a username is taken, locating a record in a database, finding a file in an index, routing packets, or matching a word inside a document. The “shape” of your data, unsorted list, sorted array, hash table, tree, or text stream, determines what kinds of search strategies are possible and how fast they can be. Good search choices are really about trade-offs. A linear scan is universal and simple but grows slower as data grows; binary search is dramatically faster but requires sorted data; hash-based lookup is usually constant-time but depends on hashing quality and memory overhead; probabilistic filters trade perfect accuracy for huge space savings. Understanding these options helps you design systems that scale cleanly and keeps your code practical: fast where it needs to be, predictable under load, and correct for the guarantees your application requires. -### Which Search Should I Use? +All array examples use zero-based indices and ascending order unless stated otherwise. Bounds assume constant-time random access and comparisons. Hash-based bounds assume constant-time hashing and key comparison; variable-length keys add their own processing cost. + +## Which Search Should I Use? Choosing the right search algorithm depends on your data structure, constraints, and use case. Use this decision guide to quickly identify the best approach: -The fastest way to pick correctly is to start from your constraints, not from the algorithm names. Ask: **Is the data sorted?** **Do I need exact matches or “probably in set”?** **Am I searching text or keys?** If you answer those, most options eliminate themselves. +The fastest way to pick correctly is to start from your constraints, not from the algorithm names. Ask: Is the data sorted? Do I need exact matches or “probably in set”? Am I searching text or keys? If you answer those, most options eliminate themselves. -#### For Arrays and Lists +### For Arrays and Lists -**Unsorted data + need exact match?** +Unsorted data + need exact match? -* **Linear Search** , Simple scan left-to-right; $O(n)$ time, works on any list -* **Sentinel Linear Search** , Slight optimization removing bounds checks; still $O(n)$ +- Linear Search: Simple scan left-to-right; $O(n)$ time, works on any list +- Sentinel Linear Search: Slight optimization removing bounds checks; still $O(n)$ -**Sorted array + random access?** +Sorted array + random access? -* **Binary Search** (default choice) , Repeatedly halve the search space; $O(\log n)$ time +- Binary Search (default choice): Repeatedly halve the search space; $O(\log n)$ time - * *When target likely near start or unknown bounds?* → **Exponential Search** , Double index until range found, then binary search - * *Values uniformly distributed numeric?* → **Interpolation Search** , Estimate position based on value; expected $O(\log \log n)$ but can degrade to $O(n)$ - * *Want blocky scanning behavior?* → **Jump Search** , Jump by $\sqrt{n}$ blocks, then linear scan; $O(\sqrt{n})$ time; better cache locality than binary in some cases + - When target likely near start or unknown bounds? → Exponential Search: Double index until range found, then binary search + - Values uniformly distributed numeric? → Interpolation Search: Estimate position based on value; expected $O(\log \log n)$ but can degrade to $O(n)$ + - Want blocky scanning behavior? → Jump Search: Jump by $\sqrt{n}$ blocks, then linear scan; $O(\sqrt{n})$ time; better cache locality than binary in some cases -**Sorted array + finding first/last occurrence?** +Sorted array + finding first/last occurrence? -* **Binary Search variant** , Adjust binary search to continue after finding match +- Binary Search variant: Adjust binary search to continue after finding match A practical note: if you will search the same array many times, sorting up front can be worth it even if sorting costs $O(n\log n)$. If you only search once, sorting may be wasted work. -#### For Key-Based Lookup and Membership +### For Key-Based Lookup and Membership -**Key → value lookup or set membership?** +Key → value lookup or set membership? -* **Hash Tables** , Expected $O(1)$ lookup with good hash function and load factor +- Hash Tables: Expected $O(1)$ lookup with good hash function and load factor - * *Collision resolution:* Choose based on your needs: + - Collision resolution: Choose based on your needs: - * **Separate Chaining** , Lists/arrays per bucket; easiest deletions; steady performance at α ≈ 1 - * **Linear Probing** , Simple and cache-friendly; watch for primary clustering; keep α < 0.7 - * **Quadratic Probing** , Reduces primary clustering; use prime table sizes - * **Double Hashing** , Best probe distribution; minimizes clustering; slightly more complex - * **Cuckoo Hashing** , Guarantees $O(1)$ worst-case lookup (checks 2 positions); insertions may trigger rehashing + - Separate Chaining: Lists/arrays per bucket; easiest deletions; steady performance at α ≈ 1 + - Linear Probing: Simple and cache-friendly; watch for primary clustering; keep α < 0.7 + - Quadratic Probing: Reduces primary clustering; use prime table sizes + - Double Hashing: Key-dependent probe steps reduce clustering; slightly more complex + - Cuckoo Hashing: Guarantees $O(1)$ worst-case lookup (checks 2 positions); insertions may trigger rehashing Hash tables are the default when you want speed and don’t care about order. The “gotcha” is that hash tables only stay fast if you keep them from getting too full and you use a good hash function. -#### For Approximate Membership (Space-Efficient) +### For Approximate Membership (Space-Efficient) -**Need to test "is element in set?" with space priority?** +Need to test "is element in set?" with space priority? -* **Bloom Filter** , Probabilistic membership; answers "maybe present" or "definitely not"; no false negatives; no deletions -* **Counting Bloom Filter** , Like Bloom but supports deletions via counters; uses more space -* **Cuckoo Filter** , Space-efficient with deletions; high load factors (90%+); better performance than Bloom in many cases +- Bloom Filter: Probabilistic membership; answers "maybe present" or "definitely not"; no false negatives; no deletions +- Counting Bloom Filter: Like Bloom but supports deletions via counters; uses more space +- Cuckoo Filter: Space-efficient with deletions; high load factors (90%+); better performance than Bloom in many cases -Approximate membership structures are built for one move: quickly rejecting items that are *definitely not* present. They save time and memory by allowing **false positives** (“maybe present”), but never **false negatives** (“definitely not” is always correct). +Approximate membership structures are built for one move: quickly rejecting items that are definitely not present. They save time and memory by allowing false positives (“maybe present”), but never false negatives (“definitely not” is always correct). -#### For Substring/Pattern Search +### For Substring/Pattern Search -**Need to find pattern in text?** +Need to find pattern in text? -* **Naive String Search** , Simple sliding window; $O(n \cdot m)$ worst case; fine for short patterns -* **Knuth–Morris–Pratt (KMP)** , Guaranteed $O(n + m)$; uses failure function to avoid rechecks; best all-rounder -* **Boyer–Moore (BM)** , Often fastest in practice; scans right-to-left with smart skips; excellent for long patterns and large alphabets -* **Rabin–Karp (RK)** , Rolling hash comparison; $O(n + m)$ expected; great for multiple patterns or streaming data +- Naive String Search: Simple sliding window; $O(n \cdot m)$ worst case; fine for short patterns +- Knuth–Morris–Pratt (KMP): Guaranteed $O(n + m)$; uses failure function to avoid rechecks; best all-rounder +- Boyer–Moore (BM): Often fastest in practice; scans right-to-left with smart skips; excellent for long patterns and large alphabets +- Rabin–Karp (RK): Rolling hash comparison; $O(n + m)$ expected; great for multiple patterns or streaming data Text search is its own world because matching strings has structure you can exploit: repeated prefixes, character distributions, and the fact that mismatches can let you skip big chunks of work. -#### Quick Comparison Table +### Quick Comparison Table | Use Case | Data Type | Algorithm | Time Complexity | Notes | | ---------------------------- | ------------ | ------------------------------ | -------------------- | ------------------------------- | | Unsorted lookup | Array | Linear Search | $O(n)$ | Simple, works anywhere | | Sorted lookup | Array | Binary Search | $O(\log n)$ | Default for sorted arrays | -| Target near start | Sorted Array | Exponential Search | $O(\log p)$ | p = position found | +| Target near start | Sorted Array | Exponential Search | $O(\log(p+2))$ | p = match or insertion position | | Uniform distribution | Sorted Array | Interpolation Search | $O(\log \log n)$ avg | Can degrade to $O(n)$ | | Block-based scan | Sorted Array | Jump Search | $O(\sqrt{n})$ | Better cache locality | | Key-value pairs | Hash Table | Hash + Chaining/Probing | $O(1)$ expected | Depends on load factor | @@ -82,20 +84,20 @@ Text search is its own world because matching strings has structure you can expl | Set membership (space-first) | Bit Array | Bloom Filter | $O(k)$ | k = hash functions; FPR tunable | | Set with deletions | Counters | Counting Bloom / Cuckoo Filter | $O(k)$ | Supports removal | | Substring search | String | KMP | $O(n + m)$ | Guaranteed linear time | -| Substring (fast average) | String | Boyer–Moore | $O(n/m)$ avg | Best for long patterns | -| Multiple patterns | String | Rabin–Karp | $O(n + m)$ | Rolling hash enables batching | +| Substring search | String | Boyer–Moore | Variant/input dependent | Can skip many characters; include preprocessing | +| Multiple patterns | String | Rabin–Karp | Hashing plus verification/output | Convenient for equal-length patterns | -### Linear & Sequential Search +## Linear & Sequential Search Linear search is the baseline: it requires no assumptions and no preprocessing, which is exactly why it still matters. If your data is tiny, frequently changing, or you only search once, linear search is often “good enough” and the simplest correct solution. -#### Linear Search +### Linear Search Scan the list from left to right, comparing the target with each element until you either find a match (return its index) or finish the list (report “not found”). -**Example inputs and outputs** +Example inputs and outputs: -*Example 1* +Example 1 $$ \text{Input: } [7, 3, 5, 2, 9], \quad \text{target} = 5 @@ -105,7 +107,7 @@ $$ \text{Output: } \text{index} = 2 $$ -*Example 2* +Example 2 $$ \text{Input: } [4, 4, 4], \quad \text{target} = 4 @@ -115,7 +117,7 @@ $$ \text{Output: } \text{index} = 0 (\text{first match}) $$ -*Example 3* +Example 3 $$ \text{Input: } [10, 20, 30], \quad \text{target} = 25 @@ -125,11 +127,11 @@ $$ \text{Output: } \text{not found} $$ -**How Linear Search Works** +How Linear Search Works -We start at index `0`, compare the value with the target, and keep moving right until we either **find it** or reach the **end**. +We start at index `0`, compare the value with the target, and keep moving right until we either find it or reach the end. -Target **5** in `[7, 3, 5, 2, 9]` +Target 5 in `[7, 3, 5, 2, 9]` ``` Indexes: 0 1 2 3 4 @@ -137,7 +139,7 @@ List: [7] [3] [5] [2] [9] Target: 5 ``` -*Step 1:* pointer at index 0 +Step 1: pointer at index 0 ``` | @@ -147,7 +149,7 @@ v → compare 7 vs 5 → no ``` -*Step 2:* pointer moves to index 1 +Step 2: pointer moves to index 1 ``` | @@ -157,19 +159,19 @@ v → compare 3 vs 5 → no ``` -*Step 3:* pointer moves to index 2 +Step 3: pointer moves to index 2 ``` | v 7 3 5 2 9 -→ compare 5 vs 5 → YES ✅ → return index 2 +→ compare 5 vs 5 → YES → return index 2 ``` -**Worst Case (Not Found)** +Worst Case (Not Found) -Target **9** in `[1, 2, 3]` +Target 9 in `[1, 2, 3]` ``` Indexes: 0 1 2 @@ -184,26 +186,26 @@ Checks: → 2 ≠ 9 → 3 ≠ 9 → end -→ not found ❌ +→ not found ``` -* Works on any list; no sorting or structure required. -* Returns the first index containing the target; if absent, reports “not found.” -* Time: $O(n)$ comparisons on average and in the worst case; best case $O(1)$ if the first element matches. -* Space: $O(1)$ extra memory. -* Naturally finds the earliest occurrence when duplicates exist. -* Simple and dependable for short or unsorted data. -* Assumes 0-based indexing in these notes. +- Works on any list; no sorting or structure required. +- Returns the first index containing the target; if absent, reports “not found.” +- Time: $O(n)$ comparisons on average and in the worst case; best case $O(1)$ if the first element matches. +- Space: $O(1)$ extra memory. +- Naturally finds the earliest occurrence when duplicates exist. +- Simple and dependable for short or unsorted data. +- Assumes 0-based indexing in these notes. -#### Sentinel Linear Search +### Sentinel Linear Search Place one copy of the target at the very end as a “sentinel” so the scan can run without checking bounds each step; afterward, decide whether the match was inside the original list or only at the sentinel position. Sentinel search is the same algorithm with a small engineering twist: remove a bounds check from the loop. That doesn’t change big-O, but in performance-sensitive code it can matter, especially in low-level languages where bounds checks are not free. -**Example inputs and outputs** +Example inputs and outputs: -*Example 1* +Example 1 $$ \text{Input: } [12, 8, 6, 15], \quad \text{target} = 6 @@ -213,7 +215,7 @@ $$ \text{Output: } \text{index} = 2 $$ -*Example 2* +Example 2 $$ \text{Input: } [2, 4, 6, 8], \quad \text{target} = 5 @@ -223,11 +225,11 @@ $$ \text{Output: } \text{not found } (\text{only the sentinel matched}) $$ -**How it works** +How it works: Put the target at one extra slot at the end so the loop is guaranteed to stop on a match; afterward, check whether the match was inside the original range. -Target **11** not in the list +Target 11 not in the list ``` Original list (n=5): @@ -253,9 +255,9 @@ Scan step by step: 11 = 11 → pointer at 5 (sentinel) ``` -Therefore, **not found** in original list. +Therefore, not found in original list. -Target **6** inside the list +Target 6 inside the list ``` Original list (n=4): @@ -275,28 +277,28 @@ Scan: ``` 12 ≠ 6 → index 0 8 ≠ 6 → index 1 - 6 = 6 → index 2 ✅ + 6 = 6 → index 2 ``` -* Removes the per-iteration “have we reached the end?” check; the sentinel guarantees termination. -* Same $O(n)$ time in big-O terms, but slightly fewer comparisons in tight loops. -* Space: needs one extra slot; if you cannot append, you can temporarily overwrite the last element (store it, write the target, then restore it). -* After scanning, decide by index: if the first match index < original length, it’s a real match; otherwise, it’s only the sentinel. -* Use when micro-optimizing linear scans over arrays where bounds checks are costly. -* Behavior with duplicates: still returns the first occurrence within the original range. -* Be careful to restore any overwritten last element if you used the in-place variant. +- Removes the per-iteration “have we reached the end?” check; the sentinel guarantees termination. +- Same $O(n)$ time in big-O terms, but slightly fewer comparisons in tight loops. +- Space: needs one extra slot; if you cannot append, you can temporarily overwrite the last element (store it, write the target, then restore it). +- For an appended sentinel, a match before the original length is real; a match at the original length is not. If overwriting the last element, save and restore it, then accept a match only if it occurred before the last index or the saved last element equals the target. Handle an empty array separately. +- Use when micro-optimizing linear scans over arrays where bounds checks are costly. +- Behavior with duplicates: still returns the first occurrence within the original range. +- Be careful to restore any overwritten last element if you used the in-place variant. -### Divide & Conquer Search +## Divide & Conquer Search Divide-and-conquer search algorithms assume structure, almost always sorted order, and then exploit it to throw away most of the search space quickly. The “why you should care” is simple: when $n$ gets big, cutting the problem in half repeatedly is the difference between instant and unbearable. -#### Binary Search +### Binary Search On a sorted array, repeatedly halve the search interval by comparing the target to the middle element until found or the interval is empty. -**Example inputs and outputs** +Example inputs and outputs: -*Example 1* +Example 1 $$ \text{Input: } A = [2, 5, 8, 12, 16, 23, 38], \quad \text{target} = 16 @@ -306,7 +308,7 @@ $$ \text{Output: } \text{index} = 4 $$ -*Example 2* +Example 2 $$ \text{Input: } A = [1, 3, 3, 3, 9], \quad \text{target} = 3 @@ -316,7 +318,7 @@ $$ \text{Output: } \text{index} = 2 \quad (\text{any valid match; first/last needs a variant}) $$ -*Example 3* +Example 3 $$ \text{Input: } A = [10, 20, 30, 40], \quad \text{target} = 35 @@ -326,18 +328,18 @@ $$ \text{Output: } \text{not found} $$ -**How it works** +How it works: -We repeatedly check the **middle** element, and then discard half the list based on comparison. +We repeatedly check the middle element, and then discard half the list based on comparison. -Find **16** in: +Find 16 in: ``` A = [ 2 ][ 5 ][ 8 ][ 12 ][ 16 ][ 23 ][ 38 ] i = 0 1 2 3 4 5 6 ``` -*Step 1* +Step 1 ``` low = 0, high = 6 @@ -351,7 +353,7 @@ i = 0 1 2 3 4 5 6 Active range: indices 0..6 ``` -*Step 2* +Step 2 ``` low = 4, high = 6 @@ -365,12 +367,12 @@ i = 0 1 2 3 4 5 6 Active range: indices 4..6 ``` -*Step 3* +Step 3 ``` low = 4, high = 4 mid = 4 -A[4] = 16 == target ✅ +A[4] = 16 == target A = [ 2 ][ 5 ][ 8 ][ 12 ][ 16 ][ 23 ][ 38 ] i = 0 1 2 3 4 5 6 @@ -381,21 +383,23 @@ Active range: indices 4..4 FOUND at index 4 -* Requires a sorted array (assume ascending here). -* Time: $O(log n)$; Space: $O(1)$ iterative. -* Returns any one matching index by default; “first/last occurrence” is a small, common refinement. -* Robust, cache-friendly, and a building block for many higher-level searches. -* Beware of off-by-one errors when shrinking bounds. +- Requires a sorted array (assume ascending here). +- Time: $O(log n)$; Space: $O(1)$ iterative. +- Returns any one matching index by default; “first/last occurrence” is a small, common refinement. +- Robust, cache-friendly, and a building block for many higher-level searches. +- Beware of off-by-one errors when shrinking bounds. + +For inclusive bounds, compute `mid = low + (high - low) // 2` to avoid adding two large indices in fixed-width arithmetic. Every unsuccessful comparison must exclude `mid` from the next interval. An empty input starts with `high = -1` and returns “not found” without accessing the array. A common refinement in real work is “find the first/last occurrence.” The trick is not to stop when you see the target, keep going left (or right) while preserving correctness. Same structure, slightly different termination condition. -#### Ternary Search +### Ternary Search -Like binary, but splits the current interval into three parts using two midpoints; used mainly for unimodal functions or very specific array cases. +For sorted-array lookup, ternary search splits the interval using two midpoints. A related but distinct optimization technique finds an extremum of a unimodal function by comparing function values; it does not use the target-comparison rules shown below. -**Example inputs and outputs** +Example inputs and outputs: -*Example 1* +Example 1 $$ \text{Input: } A = [1, 4, 7, 9, 12, 15], \quad \text{target} = 9 @@ -405,7 +409,7 @@ $$ \text{Output: } \text{index} = 3 $$ -*Example 2* +Example 2 $$ \text{Input: } A = [2, 6, 10, 14], \quad \text{target} = 5 @@ -415,13 +419,15 @@ $$ \text{Output: } \text{not found} $$ -**How it works** +How it works: -We divide the array into **three parts** using two midpoints `m1` and `m2`. +We divide the array into three parts using two midpoints `m1` and `m2`. -* If `target < A[m1]` → search $[low .. m1-1]$ -* Else if `target > A[m2]` → search $[m2+1 .. high]$ -* Else → search $[m1+1 .. m2-1]$ +First return a match if the target equals either midpoint value. Otherwise: + +- If `target < A[m1]` → search $[low .. m1-1]$ +- Else if `target > A[m2]` → search $[m2+1 .. High]$ +- Else → search $[m1+1 .. m2-1]$ ``` A = [ 1 ][ 4 ][ 7 ][ 9 ][ 12 ][ 15 ] @@ -429,37 +435,37 @@ i = 0 1 2 3 4 5 Target: 9 ``` -*Step 1* +Step 1 ``` low = 0, high = 5 m1 = low + (high - low)//3 = 0 + (5)//3 = 1 -m2 = high - (high - low)//3 = 5 - (5)//3 = 3 +m2 = high - (high - low)//3 = 5 - (5)//3 = 4 A[m1] = 4 -A[m2] = 9 +A[m2] = 12 A = [ 1 ][ 4 ][ 7 ][ 9 ][ 12 ][ 15 ] i = 0 1 2 3 4 5 - ↑L ↑m1 ↑m2 ↑H - 0 1 3 5 + ↑L ↑m1 ↑m2 ↑H + 0 1 4 5 ``` -FOUND at index 3 +Since `4 < 9 < 12`, continue in indices `2..3`. The next midpoints are 2 and 3; `A[3] = 9`, so the target is found at index 3. -* Also assumes a sorted array. -* For discrete sorted arrays, it does **not** beat binary search asymptotically; it performs more comparisons per step. -* Most valuable for searching the extremum of a **unimodal function** on a continuous domain; for arrays, prefer binary search. -* Complexity: $O(log n)$ steps but with larger constant factors than binary search. +- Also assumes a sorted array. +- For discrete sorted arrays, it does not beat binary search asymptotically; it performs more comparisons per step. +- Most valuable for searching the extremum of a unimodal function on a continuous domain; for arrays, prefer binary search. +- Complexity: $O(log n)$ steps but with larger constant factors than binary search. -#### Jump Search +### Jump Search On a sorted array, jump ahead in fixed block sizes to find the block that may contain the target, then do a linear scan inside that block. -**Example inputs and outputs** +Example inputs and outputs: -*Example 1* +Example 1 $$ \text{Input: } A = [1, 4, 9, 16, 25, 36, 49], \quad \text{target} = 25, \quad \text{jump} = \lfloor \sqrt{7} \rfloor = 2 @@ -469,7 +475,7 @@ $$ \text{Output: } \text{index} = 4 $$ -*Example 2* +Example 2 $$ \text{Input: } A = [3, 8, 15, 20, 22, 27], \quad \text{target} = 21, \quad \text{jump} = 2 @@ -479,23 +485,23 @@ $$ \text{Output: } \text{not found} $$ -**How it works** +How it works: -Here’s a clear trace of **jump search** in action: +Here’s a clear trace of jump search in action: -We’re applying **jump search** to find $25$ in +We’re applying jump search to find $25$ in $$ A = [1, 4, 9, 16, 25, 36, 49] $$ -with $n=7$, block size $\approx \sqrt{7} \approx 2$, so **jump=2**. +with $n=7$, block size $\approx \sqrt{7} \approx 2$, so jump=2. We probe every 2nd index: -* probe = 0 → $A[0] = 1 < 25$ → jump to 2 -* probe = 2 → $A[2] = 9 < 25$ → jump to 4 -* probe = 4 → $A[4] = 25 \geq 25$ → stop +- probe = 0 → $A[0] = 1 < 25$ → jump to 2 +- probe = 2 → $A[2] = 9 < 25$ → jump to 4 +- probe = 4 → $A[4] = 25 \geq 25$ → stop So target is in block $(2..4]$ @@ -507,8 +513,8 @@ So target is in block $(2..4]$ Linear Scan in block (indexes 3..4) -* i = 3 → $A[3] = 16 < 25$ -* i = 4 → $A[4] = 25 = 25$ ✅ FOUND +- i = 3 → $A[3] = 16 < 25$ +- i = 4 → $A[4] = 25 = 25$ FOUND ``` Block [16 ][25 ] @@ -516,20 +522,20 @@ Block [16 ][25 ] i=3 i=4 (found!) ``` -The element $25$ is found at **index 4**. +The element $25$ is found at index 4. -* Works on sorted arrays; pick jump ≈ √n for good balance. -* Time: $O(√n)$ comparisons on average; Space: $O(1)$. -* Useful when binary search isn’t desirable: binary search has poor locality (random jumps); jump search scans a block sequentially, which is more cache-friendly and has better branch prediction. -* Degrades gracefully to “scan block then stop.” +- Works on sorted arrays; pick jump ≈ √n for good balance. +- Time: $O(√n)$ comparisons on average; Space: $O(1)$. +- Useful when binary search isn’t desirable: binary search has poor locality (random jumps); jump search scans a block sequentially, which is more cache-friendly and has better branch prediction. +- Degrades gracefully to “scan block then stop.” -#### Exponential Search +### Exponential Search On a sorted array, grow the right boundary exponentially (1, 2, 4, 8, …) to find a containing range, then finish with binary search in that range. -**Example inputs and outputs** +Example inputs and outputs: -*Example 1* +Example 1 $$ \text{Input: } A = [2, 3, 5, 7, 11, 13, 17, 19, 23], \quad \text{target} = 19 @@ -539,7 +545,7 @@ $$ \text{Output: } \text{index} = 7 $$ -*Example 2* +Example 2 $$ \text{Input: } A = [10, 20, 30, 40, 50], \quad \text{target} = 12 @@ -549,7 +555,7 @@ $$ \text{Output: } \text{not found} $$ -**How it works** +How it works: ``` A = [ 2 ][ 3 ][ 5 ][ 7 ][ 11 ][ 13 ][ 17 ][ 19 ][ 23 ] @@ -559,30 +565,30 @@ Target = 19 Find range by exponential jumps -*Start* at `i=1`, double each step until `A[i] ≥ target` (or end). +Start at `i=1`, double each step until `A[i] ≥ target` (or end). -*Jump 1:* `i=1` +Jump 1: `i=1` ``` A[i]=3 ≤ 19 → continue ... ``` -*Jump 2:* `i=2` +Jump 2: `i=2` ``` A[i]=5 ≤ 19 → continue ... ``` -*Jump 3:* `i=4` +Jump 3: `i=4` ``` A[i]=11 ≤ 19 → continue ... ``` -*Jump 4:* `i=8` +Jump 4: `i=8` ``` A[i]=23 > 19 → stop @@ -590,7 +596,7 @@ Range is (previous power of two .. i] = (4 .. 8] → search indices 5..8 ... ``` -*Range for binary search:* `low=5, high=8`. +Range for binary search: `low=5, high=8`. Binary search on $A[5..8]$ @@ -599,7 +605,7 @@ Subarray: [ 13 ][ 17 ][ 19 ][ 23 ] Indices : 5 6 7 8 ``` -*Step 1* +Step 1 ``` low=5, high=8 → mid=(5+8)//2=6 @@ -607,28 +613,28 @@ A[6]=17 < 19 → move right → low=7 ... ``` -*Step 2* +Step 2 ``` low=7, high=8 → mid=(7+8)//2=7 -A[7]=19 == target ✅ → FOUND +A[7]=19 == target → FOUND ... ``` -Found at **index 7**. +Found at index 7. -* Great when the target is likely to be near the beginning or when the array is **unbounded**/**stream-like** but sorted (you can probe indices safely). -* Time: $O(log p)$ to find the range where p is the final bound, plus $O(log p)$ for binary search → overall $O(log p)$. -* Space: $O(1)$. -* Often paired with data sources where you can test “is index i valid?” while doubling i. +- Useful when the target is near the beginning or the length of a sorted random-access source is unknown. A sequential-only stream cannot support the indexed probes without additional storage or traversal. +- Check index zero first. Finding and searching a range ending near position $p$ takes $O(\log(p+2))$ time, including constant-time cases at the beginning. For a finite array, cap the upper bound at `n-1` and handle an empty array before probing. +- Space: $O(1)$. +- Often paired with data sources where you can test “is index i valid?” while doubling i. -#### Interpolation Search +### Interpolation Search On a sorted (roughly uniformly distributed) array, estimate the likely position using the values themselves and probe there; repeat on the narrowed side. -**Example inputs and outputs** +Example inputs and outputs: -*Example 1* +Example 1 $$ \text{Input: } A = [10, 20, 30, 40, 50, 60, 70], \quad \text{target} = 55 @@ -638,7 +644,7 @@ $$ \text{Output: } \text{not found } (\text{probes near indices } 4\text{–}5) $$ -*Example 2* +Example 2 $$ \text{Input: } A = [5, 15, 25, 35, 45, 55, 65], \quad \text{target} = 45 @@ -648,7 +654,7 @@ $$ \text{Output: } \text{index} = 4 $$ -*Example 3* +Example 3 $$ \text{Input: } A = [1, 1000, 1001, 1002], \quad \text{target} = 2 @@ -658,12 +664,12 @@ $$ \text{Output: } \text{not found } (\text{bad distribution for interpolation}) $$ -**How it works** +How it works: -* Guard against division by zero: if `A[high] == A[low]`, stop (or binary-search fallback). -* Clamp the computed `pos` to `[low, high]` before probing. -* Works best when values are **uniformly distributed**; otherwise it can degrade toward linear time. -* Assumes `A` is sorted and values are uniform. +- Before probing, check that the interval is nonempty and the target lies between its endpoint values. If `A[high] == A[low]`, return a match if that value equals the target; otherwise report absence without dividing. +- Clamp the computed `pos` to `[low, high]` before probing. +- Works best when values are uniformly distributed; otherwise it can degrade toward linear time. +- Sorted numeric data is required for correctness. Approximately uniform values are a performance assumption, not a correctness requirement. Probe formula: @@ -679,7 +685,7 @@ i = 0 1 2 3 4 5 6 target = 45 ``` -*Step 1 , initial probe* +Step 1: initial probe ``` low=0 (A[0]=10), high=6 (A[6]=70) @@ -690,40 +696,40 @@ pos ≈ 0 + (6-0) * (45-10)/(70-10) ... ``` -Probe **index 3**: `A[3]=40 < 45` → set `low = 3 + 1 = 4` +Probe index 3: `A[3]=40 < 45` → set `low = 3 + 1 = 4` -*Step 2 , after moving low* +Step 2: after moving low ``` A = [ 10 ][ 20 ][ 30 ][ 40 ][ 50 ][ 60 ][ 70 ] ... ``` -At this point, an **early-stop check** already tells us `target (45) < A[low] (50)` → cannot exist in `A[4..6]` → **not found**. +At this point, an early-stop check already tells us `target (45) < A[low] (50)` → cannot exist in `A[4..6]` → not found. -* Best on **uniformly distributed** sorted data; expected time $O(log log n)$. -* Worst case can degrade to $O(n)$, especially on skewed or clustered values. -* Space: $O(1)$. -* Very fast when value-to-index mapping is close to linear (e.g., near-uniform numeric keys). -* Requires careful handling when A[high] = A[low] (avoid division by zero); also sensitive to integer rounding in discrete arrays. +- Best on uniformly distributed sorted data; expected time $O(log log n)$. +- Worst case can degrade to $O(n)$, especially on skewed or clustered values. +- Space: $O(1)$. +- Very fast when value-to-index mapping is close to linear (e.g., near-uniform numeric keys). +- Requires careful handling when A[high] = A[low] (avoid division by zero); also sensitive to integer rounding in discrete arrays. -### Hash-based Search +## Hash-based Search Hash-based search is for when you want direct access by key. Instead of narrowing down ranges, you compute a location in (roughly) constant time. The cost is setup complexity: you need a hash function, a collision strategy, and resizing rules to keep performance stable. -* **Separate chaining:** Easiest deletions, steady $O(1)$ with α≈1; good when memory fragmentation isn’t a concern. -* **Open addressing (double hashing):** Best probe quality among OA variants; great cache locality; keep α < 0.8. -* **Open addressing (linear/quadratic):** Simple and fast at low α; watch clustering and tombstones. -* **Cuckoo hashing:** Tiny and predictable lookup cost; inserts costlier and may rehash; great for read-heavy workloads. -* In all cases: pick strong hash functions and resize early to keep α healthy. +- Separate chaining: Easiest deletions, steady $O(1)$ with α≈1; good when memory fragmentation isn’t a concern. +- Open addressing (double hashing): key-dependent probe steps reduce clustering, but scattered probes usually have poorer locality than linear probing. Choose a load threshold appropriate to the implementation. +- Open addressing (linear/quadratic): Simple and fast at low α; watch clustering and tombstones. +- Cuckoo hashing: Tiny and predictable lookup cost; inserts costlier and may rehash; great for read-heavy workloads. +- In all cases: pick strong hash functions and resize early to keep α healthy. -#### Hash Table Search +### Hash Table Search Map a key to an array index with a hash function; look at that bucket to find the key, giving expected $O(1)$ lookups under a good hash and healthy load factor. -**Example inputs and outputs** +Example inputs and outputs: -*Example 1* +Example 1 $$ \text{Table size: } m = 7, @@ -735,7 +741,7 @@ $$ \text{Output: } \text{found (bucket 3)} $$ -*Example 2* +Example 2 $$ \text{Table size: } m = 7, @@ -747,7 +753,7 @@ $$ \text{Output: } \text{not found} $$ -**How it works** +How it works: ``` +-----+ hash +-----------------+ search/compare +--------+ @@ -755,9 +761,9 @@ $$ +-----+ +-----------------+ +--------+ ``` -* With chaining, the “collision path” is the **list inside one bucket**. -* With linear probing, the “collision path” is the **probe sequence** across buckets (3 → 4 → 5 → …). -* Both keep your original flow: hash → inspect bucket (and collision path) → match? +- With chaining, the “collision path” is the list inside one bucket. +- With linear probing, the “collision path” is the probe sequence across buckets (3 → 4 → 5 → …). +- Both keep your original flow: hash → inspect bucket (and collision path) → match? ``` Array (buckets/indexes 0..6): @@ -768,9 +774,9 @@ Idx: 0 1 2 3 4 5 6 +---+-----+-----+-----+-----+-----+-----+ ``` -**Example mapping with** `h(k) = k mod 7`, **stored keys** `{10, 24, 31}` all hash to index `3`. +Example mapping with `h(k) = k mod 7`, stored keys `{10, 24, 31}` all hash to index `3`. -*Strategy A , Separate Chaining (linked list per bucket)* +Strategy A: Separate Chaining (linked list per bucket) Insertions @@ -787,7 +793,7 @@ Idx: 0 1 2 3 4 5 6 bucket[3] chain: [10] → [24] → [31] → ∅ ``` -*Search(24)* +Search(24) ``` 1) Compute index = h(24) = 3 @@ -797,7 +803,7 @@ bucket[3] chain: [10] → [24] → [31] → ∅ 3) Return FOUND (bucket 3) ``` -*Strategy B , Open Addressing (Linear Probing)* +Strategy B: Open Addressing (Linear Probing) Insertions @@ -812,7 +818,7 @@ Idx: 0 1 2 3 4 5 6 +---+-----+-----+-----+-----+-----+-----+ ``` -*Search(24)* +Search(24) ``` 1) Compute index = h(24) = 3 @@ -822,18 +828,18 @@ Idx: 0 1 2 3 4 5 6 (If not found, continue probing until an empty slot or wrap limit.) ``` -* Quality hash + low load factor (α = n/m) ⇒ expected $O(1)$ search/insert/delete. -* Collisions are inevitable; the collision strategy (open addressing vs. chaining vs. cuckoo) dictates actual steps. -* Rehashing (growing and re-inserting) is used to keep α under control. -* Uniform hashing assumption underpins the $O(1)$ expectation; adversarial keys or poor hashes can degrade performance. +- Quality hash + low load factor (α = n/m) ⇒ expected $O(1)$ search/insert/delete. +- Collisions are inevitable; the collision strategy (open addressing vs. Chaining vs. Cuckoo) dictates actual steps. +- Rehashing (growing and re-inserting) is used to keep α under control. +- Uniform hashing assumption underpins the $O(1)$ expectation; adversarial keys or poor hashes can degrade performance. -#### Open Addressing , Linear Probing +### Open Addressing: Linear Probing Keep everything in one array; on collision, probe alternative positions in a deterministic sequence until an empty slot or the key is found. -**Example inputs and outputs** +Example inputs and outputs: -*Example 1* +Example 1 $$ m = 10, @@ -845,7 +851,7 @@ $$ \text{Output: } \text{found (index 3)} $$ -*Example 2* +Example 2 $$ m = 10, @@ -857,20 +863,20 @@ $$ \text{Output: } \text{not found} $$ -**How it works** +How it works: -*Hash function:* +Hash function: ``` h(k) = k mod 10 Probe sequence: i, i+1, i+2, ... (wrap around) ``` -*Insertions* +Insertions -* Insert 12 → `h(12)=2` → place at index 2 -* Insert 22 → `h(22)=2` occupied → probe 3 → place at 3 -* Insert 32 → `h(32)=2` occupied → probe 3 (occupied) → probe 4 → place at 4 +- Insert 12 → `h(12)=2` → place at index 2 +- Insert 22 → `h(22)=2` occupied → probe 3 → place at 3 +- Insert 32 → `h(32)=2` occupied → probe 3 (occupied) → probe 4 → place at 4 Resulting table (indexes 0..9): @@ -881,11 +887,11 @@ Value: | | | 12 | 22 | 32 | | | | | | +---+---+----+----+----+---+---+---+---+---+ ``` -*Search(22)* +Search(22) -* Start at `h(22)=2` -* index 2 → 12 ≠ 22 → probe → -* index 3 → 22 ✅ FOUND +- Start at `h(22)=2` +- index 2 → 12 ≠ 22 → probe → +- index 3 → 22 FOUND Path followed: @@ -893,13 +899,13 @@ Path followed: 2 → 3 ``` -*Search(42)* +Search(42) -* Start at `h(42)=2` -* index 2 → 12 ≠ 42 → probe → -* index 3 → 22 ≠ 42 → probe → -* index 4 → 32 ≠ 42 → probe → -* index 5 → empty slot → stop → ❌ NOT FOUND +- Start at `h(42)=2` +- index 2 → 12 ≠ 42 → probe → +- index 3 → 22 ≠ 42 → probe → +- index 4 → 32 ≠ 42 → probe → +- index 5 → empty slot → stop → NOT FOUND Path followed: @@ -907,16 +913,16 @@ Path followed: 2 → 3 → 4 → 5 (∅) ``` -* Simple and cache-friendly; clusters form (“primary clustering”) which can slow probes. -* Deletion uses **tombstones** to keep probe chains intact. -* Performance depends sharply on load factor; keep α well below 1 (e.g., α ≤ 0.7). -* Expected search~ $O(1)$ at low α; degrades as clusters grow. +- Simple and cache-friendly; clusters form (“primary clustering”) which can slow probes. +- Deletion uses tombstones to keep probe chains intact. +- Performance depends sharply on load factor; keep α well below 1 (e.g., α ≤ 0.7). +- Expected search~ $O(1)$ at low α; degrades as clusters grow. -#### Open Addressing , Quadratic Probing +### Open Addressing: Quadratic Probing -**Example inputs and outputs** +Example inputs and outputs: -*Example 1* +Example 1 $$ m = 11 (\text{prime}), @@ -932,7 +938,7 @@ $$ \text{Output: found (index 1)} $$ -*Example 2* +Example 2 $$ m = 11 (\text{prime}), @@ -948,15 +954,15 @@ $$ \text{Output: not found} $$ -**How it works** +How it works: -*Hash function:* +Hash function: ``` h(k) = k mod 11 ``` -*Probe sequence (relative offsets):* +Probe sequence (relative offsets): ``` +1², +2², +3², +4², +5², ... mod 11 @@ -969,11 +975,11 @@ So from `h(k)`, we try slots in this order: h, h+1, h+4, h+9, h+5, h+3, ... (all mod 11) ``` -*Insertions* +Insertions -* Insert **22** → `h(22)=0` → place at index 0 -* Insert **33** → `h(33)=0` occupied → try `0+1²=1` → index 1 free → place at 1 -* Insert **44** → `h(44)=0` occupied → probe 1 (occupied) → probe `0+4=4` → place at 4 +- Insert 22 → `h(22)=0` → place at index 0 +- Insert 33 → `h(33)=0` occupied → try `0+1²=1` → index 1 free → place at 1 +- Insert 44 → `h(44)=0` occupied → probe 1 (occupied) → probe `0+4=4` → place at 4 Resulting table: @@ -984,10 +990,10 @@ Val: | 22 | 33 | | | 44 | | | | | | | +----+----+---+---+----+---+---+---+---+---+---+ ``` -*Search(33)* +Search(33) -* Start `h(33)=0` → slot 0 = 22 ≠ 33 -* Probe `0+1²=1` → slot 1 = 33 ✅ FOUND +- Start `h(33)=0` → slot 0 = 22 ≠ 33 +- Probe `0+1²=1` → slot 1 = 33 FOUND Path: @@ -995,12 +1001,12 @@ Path: 0 → 1 ``` -*Search(55)* +Search(55) -* Start `h(55)=0` → slot 0 = 22 ≠ 55 -* Probe `0+1²=1` → slot 1 = 33 ≠ 55 -* Probe `0+2²=4` → slot 4 = 44 ≠ 55 -* Probe `0+3²=9` → slot 9 = empty → stop → ❌ NOT FOUND +- Start `h(55)=0` → slot 0 = 22 ≠ 55 +- Probe `0+1²=1` → slot 1 = 33 ≠ 55 +- Probe `0+2²=4` → slot 4 = 44 ≠ 55 +- Probe `0+3²=9` → slot 9 = empty → stop → NOT FOUND Path: @@ -1008,16 +1014,16 @@ Path: 0 → 1 → 4 → 9 (∅) ``` -* Reduces primary clustering but can exhibit **secondary clustering** (keys with same h(k) follow same probe squares). -* Table size choice matters (often prime); ensure the probe sequence can reach many slots. -* Keep α modest; deletion still needs tombstones. -* Expected $O(1)$ at healthy α; simpler than double hashing. +- Reduces primary clustering but can exhibit secondary clustering (keys with same h(k) follow same probe squares). +- For the specific rule `h + i²` modulo an odd prime, only about half the slots are visited. Keeping the table less than half full guarantees an available slot during insertion. Other quadratic rules have different coverage guarantees; prime size alone does not guarantee a full cycle. +- Keep α modest; deletion still needs tombstones. +- Expected $O(1)$ at healthy α; simpler than double hashing. -#### Open Addressing , Double Hashing +### Open Addressing: Double Hashing -**Example inputs and outputs** +Example inputs and outputs: -*Hash functions* +Hash functions $$ h_{1}(k) = k \bmod 11, @@ -1030,7 +1036,7 @@ $$ h(k,i) = \big(h_{1}(k) + i \cdot h_{2}(k)\big) \bmod 11 $$ -*Example 1* +Example 1 $$ m = 11, @@ -1059,7 +1065,7 @@ $$ \text{Output: found (index 4)} $$ -*Example 2* +Example 2 $$ m = 11, @@ -1086,16 +1092,16 @@ $$ \text{Output: not found} $$ -**How it works** +How it works: -We use **two hash functions**: +We use two hash functions: ``` h₁(k) = k mod m h₂(k) = 1 + (k mod 10) ``` -*Probe sequence:* +Probe sequence: ``` i, i + h₂, i + 2·h₂, i + 3·h₂, ... (all mod m) @@ -1103,25 +1109,25 @@ i, i + h₂, i + 2·h₂, i + 3·h₂, ... (all mod m) This ensures fewer clustering issues compared to linear or quadratic probing. -*Insertions (m = 11)* +Insertions (m = 11) -Insert **22** +Insert 22 -* `h₁(22)=0` → place at index 0 +- `h₁(22)=0` → place at index 0 -Insert **33** +Insert 33 -* `h₁(33)=0` (occupied) -* `h₂(33)=1+(33 mod 10)=4` -* Probe sequence: 0, 4 → place at index 4 +- `h₁(33)=0` (occupied) +- `h₂(33)=1+(33 mod 10)=4` +- Probe sequence: 0, 4 → place at index 4 -Insert **44** +Insert 44 -* `h₁(44)=0` (occupied) -* `h₂(44)=1+(44 mod 10)=5` -* Probe sequence: 0, 5 → place at index 5 +- `h₁(44)=0` (occupied) +- `h₂(44)=1+(44 mod 10)=5` +- Probe sequence: 0, 5 → place at index 5 -*Table State* +Table State ``` Idx: 0 1 2 3 4 5 6 7 8 9 10 @@ -1130,10 +1136,10 @@ Val: |22 | | | |33 |44 | | | | | | +---+---+---+---+---+---+---+---+---+---+---+ ``` -*Search(33)* +Search(33) -* Start at `h₁(33)=0` → slot 0 = 22 ≠ 33 -* Next: `0+1·h₂(33)=0+4=4` → slot 4 = 33 ✅ FOUND +- Start at `h₁(33)=0` → slot 0 = 22 ≠ 33 +- Next: `0+1·h₂(33)=0+4=4` → slot 4 = 33 FOUND Path: @@ -1141,11 +1147,11 @@ Path: 0 → 4 ``` -*Search(55)* +Search(55) -* `h₁(55)=0`, `h₂(55)=1+(55 mod 10)=6` -* slot 0 = 22 ≠ 55 -* slot 6 = empty → stop → ❌ NOT FOUND +- `h₁(55)=0`, `h₂(55)=1+(55 mod 10)=6` +- slot 0 = 22 ≠ 55 +- slot 6 = empty → stop → NOT FOUND Path: @@ -1153,18 +1159,18 @@ Path: 0 → 6 (∅) ``` -* Minimizes clustering; probe steps depend on the key. -* Choose h₂ so it’s **non-zero** and relatively prime to m, ensuring a full cycle. -* Excellent performance at higher α than linear/quadratic, but still sensitive if α → 1. -* Deletion needs tombstones; implementation slightly more complex. +- Minimizes clustering; probe steps depend on the key. +- Choose h₂ so it’s non-zero and relatively prime to m, ensuring a full cycle. +- Excellent performance at higher α than linear/quadratic, but still sensitive if α → 1. +- Deletion needs tombstones; implementation slightly more complex. -#### Separate Chaining +### Separate Chaining Each array cell holds a small container (e.g., a linked list); colliding keys live together in that bucket. -**Example inputs and outputs** +Example inputs and outputs: -*Setup* +Setup $$ m = 5, \quad h(k) = k \bmod 5, \quad \text{buckets hold linked lists} @@ -1198,7 +1204,7 @@ $$ h(14) = 14 \bmod 5 = 4 \;\Rightarrow\; \text{bucket 4: } [14] $$ -*Example 1* +Example 1 $$ \text{Target: } 22 @@ -1208,13 +1214,13 @@ $$ h(22) = 2 \Rightarrow \text{bucket 2} = [12, 22, 7] $$ -Found at **position 2** in the list. +Found at zero-based index 1 in the bucket list (the second element). $$ -\text{Output: found (bucket 2, position 2)} +\text{Output: found (bucket 2, list index 1)} $$ -*Example 2* +Example 2 $$ \text{Target: } 9 @@ -1230,7 +1236,7 @@ $$ \text{Output: not found} $$ -**How it works** +How it works: ``` h(k) = k mod 5 @@ -1248,19 +1254,19 @@ Search(9): - Bucket 4: [14] → 9 not present → NOT FOUND ``` -* Simple deletes (remove from a bucket) and no tombstones. -* Expected $O(1 + α)$ time; with good hashing and α kept near/below 1, bucket lengths stay tiny. -* Memory overhead for bucket nodes; cache locality worse than open addressing. -* Buckets can use **ordered lists** or **small vectors** to accelerate scans. -* Rehashing still needed as n grows; α = n/m controls performance. +- Simple deletes (remove from a bucket) and no tombstones. +- Expected $O(1 + α)$ time; with good hashing and α kept near/below 1, bucket lengths stay tiny. +- Memory overhead for bucket nodes; cache locality worse than open addressing. +- Buckets can use ordered lists or small vectors to accelerate scans. +- Rehashing still needed as n grows; α = n/m controls performance. -#### Cuckoo Hashing +### Cuckoo Hashing Keep two (or more) hash positions per key; insert by “kicking out” occupants to their alternate home so lookups check only a couple of places. -**Example inputs and outputs** +Example inputs and outputs: -*Setup* +Setup Two hash tables $T_{1}$ and $T_{2}$, each of size $$ @@ -1275,11 +1281,11 @@ $$ Cuckoo hashing invariant: -* Each key is stored either in $T_{1}[h_{1}(k)]$ or $T_{2}[h_{2}(k)]$. -* On insertion, if a spot is occupied, the existing key is **kicked out** and reinserted into the other table. -* If relocations form a cycle, the table is **rebuilt (rehash)** with new hash functions. +- Each key is stored either in $T_{1}[h_{1}(k)]$ or $T_{2}[h_{2}(k)]$. +- On insertion, if a spot is occupied, the existing key is kicked out and reinserted into the other table. +- If relocations form a cycle, the table is rebuilt (rehash) with new hash functions. -*Example 1* +Example 1 $$ \text{Target: } 15 @@ -1300,7 +1306,7 @@ $$ \text{Output: found (T₂, index 4)} $$ -*Example 2* +Example 2 If insertion causes repeated displacements and eventually loops: @@ -1312,9 +1318,9 @@ $$ \text{Output: rebuild / rehash required} $$ -**How it works** +How it works: -We keep **two hash tables (T₁, T₂)**, each with its own hash function. Every key can live in **exactly one of two possible slots**: +We keep two hash tables (T₁, T₂), each with its own hash function. Every key can live in exactly one of two possible slots: Hash functions: @@ -1322,85 +1328,87 @@ $$ h_1(k) = k \bmod 5, \quad h_2(k) = 1 + (k \bmod 4) $$ -Every key can live in **exactly one of two slots**: $T_1[h_1(k)]$ or $T_2[h_2(k)]$. -If a slot is occupied, we **evict** the old occupant and reinsert it at its alternate location. +These toy functions make the arithmetic easy to follow; they are not independent random hashes, and the second function never uses index zero. They illustrate relocation, not a recommended production hash family. + +Every key can live in exactly one of two slots: $T_1[h_1(k)]$ or $T_2[h_2(k)]$. +If a slot is occupied, we evict the old occupant and reinsert it at its alternate location. -*Start empty:* +Start empty: ``` T₁: [ ][ ][ ][ ][ ] T₂: [ ][ ][ ][ ][ ] ``` -*Insert 10* → goes to $T_1[h_1(10)=0]$: +Insert 10 → goes to $T_1[h_1(10)=0]$: ``` T₁: [10 ][ ][ ][ ][ ] T₂: [ ][ ][ ][ ][ ] ``` -*Insert 15* +Insert 15 -* $T_1[0]$ already has 10 → evict 10 -* Place 15 at $T_1[0]$ -* Reinsert evicted 10 at $T_2[h_2(10)=3]$: +- $T_1[0]$ already has 10 → evict 10 +- Place 15 at $T_1[0]$ +- Reinsert evicted 10 at $T_2[h_2(10)=3]$: ``` T₁: [15 ][ ][ ][ ][ ] T₂: [ ][ ][ ][10 ][ ] ``` -*Insert 20* +Insert 20 -* $T_1[0]$ has 15 → evict 15 -* Place 20 at $T_1[0]$ -* Reinsert 15 at $T_2[h_2(15)=4]$: +- $T_1[0]$ has 15 → evict 15 +- Place 20 at $T_1[0]$ +- Reinsert 15 at $T_2[h_2(15)=4]$: ``` T₁: [20 ][ ][ ][ ][ ] T₂: [ ][ ][ ][10 ][15 ] ``` -*Insert 25* +Insert 25 -* $T_1[0]$ has 20 → evict 20 -* Place 25 at $T_1[0]$ -* Reinsert 20 at $T_2[h_2(20)=1]$: +- $T_1[0]$ has 20 → evict 20 +- Place 25 at $T_1[0]$ +- Reinsert 20 at $T_2[h_2(20)=1]$: ``` T₁: [25 ][ ][ ][ ][ ] T₂: [ ][20 ][ ][10 ][15 ] ``` -🔎 *Search(15)* + Search(15) -* $T_1[h_1(15)=0] \to 25 \neq 15$ -* $T_2[h_2(15)=4] \to 15$ ✅ FOUND +- $T_1[h_1(15)=0] \to 25 \neq 15$ +- $T_2[h_2(15)=4] \to 15$ FOUND -**FOUND in T₂ at index 4** +FOUND in T₂ at index 4 -* Lookups probe at **most two places** (with two hashes) → excellent constant factors. -* Inserts may trigger a chain of evictions; detect cycles and **rehash** with new functions. -* High load factors achievable (e.g.,~0.5–0.9 depending on variant and number of hashes/tables). -* Deletions are easy (remove key); no tombstones, but ensure invariants remain. -* Sensitive to hash quality; poor hashes increase cycle risk. +- Lookups probe at most two places (with two hashes) → excellent constant factors. +- Inserts may trigger a chain of evictions; detect cycles and rehash with new functions. +- The basic two-table, one-slot scheme typically requires load below about one half of the combined capacity. Higher occupancy requires variants such as multiple slots per bucket; do not apply their load thresholds to this basic scheme. +- Deletions are easy (remove key); no tombstones, but ensure invariants remain. +- Sensitive to hash quality; poor hashes increase cycle risk. -### Probabilistic Membership Filters +## Probabilistic Membership Filters -Probabilistic membership filters answer a very specific question: **“Is this item in the set?”** but with a twist, they allow uncertainty in one direction. They can say **“definitely not”** with full confidence, or **“maybe”** when the filter thinks the item could be present. This trade lets them be extremely space-efficient and extremely fast. +Probabilistic membership filters answer a very specific question: “Is this item in the set?” but with a twist, they allow uncertainty in one direction. They can say “definitely not” with full confidence, or “maybe” when the filter thinks the item could be present. This trade lets them be extremely space-efficient and extremely fast. The main rule to remember is: -* **No false negatives**: if the filter says “definitely not,” the item is not in the set. -* **Possible false positives**: the filter might say “maybe present” even when the item isn’t actually there. +- No false negatives: if the filter says “definitely not,” the item is not in the set. +- Possible false positives: the filter might say “maybe present” even when the item isn’t actually there. That’s often exactly what you want when a “maybe” result can be followed by a real lookup in a slower structure (like a hash table or database). -#### Bloom Filter +### Bloom Filter Overview A Bloom filter uses a bit array of length `m` and `k` independent hash functions. To insert an element, hash it `k` times and set those `k` bit positions to `1`. To query an element, hash it again and check the same `k` positions: if any bit is `0`, the element is definitely not present; if all are `1`, it might be present. -**How it works (tiny example)** +How it works (tiny example) Bit array of length 10 (all zeros initially): @@ -1421,98 +1429,61 @@ Insert `"dog"` hashes to 1, 5, 9 bits : [0 1 1 0 0 1 0 0 1 1] ``` -Query `"cat"` → check 2,5,8 → all 1 → **maybe present** ✅ -Query `"cow"` → hashes to 3,5,7 → bit 3 is 0 → **definitely not** ❌ +Query `"cat"` → check 2,5,8 → all 1 → maybe present +Query `"cow"` → hashes to 3,5,7 → bit 3 is 0 → definitely not -* Insert: $O(k)$ bit sets -* Query: $O(k)$ bit checks -* Space: $O(m)$ bits total -* No deletions (basic Bloom): clearing bits could break other entries. -* False positive rate depends on `m`, `k`, and number of inserted items `n`. +- Insert: $O(k)$ bit sets +- Query: $O(k)$ bit checks +- Space: $O(m)$ bits total +- No deletions (basic Bloom): clearing bits could break other entries. +- False positive rate depends on `m`, `k`, and number of inserted items `n`. -#### Counting Bloom Filter +### Counting Bloom Filter Overview A Counting Bloom filter replaces the bit array with small counters (e.g., 4-bit or 8-bit). Insertion increments counters; deletion decrements counters. Query checks whether all counters are non-zero. This is the “Bloom filter that supports removals,” but you pay for it in memory: counters cost more than bits, and counter overflow must be handled carefully. -* Insert/query/delete: $O(k)$ -* More space than Bloom filter -* Supports deletions safely (within counter limits) +- Insert/query/delete: $O(k)$ +- More space than Bloom filter +- Delete only a key known to have a corresponding live insertion. Keep insertion and deletion counts balanced, and handle counter overflow without losing multiplicities. -#### Cuckoo Filter +### Cuckoo Filter Overview A Cuckoo filter stores small fingerprints (short hashes) in a table and uses cuckoo-style relocation (similar spirit to cuckoo hashing). Queries check a small number of candidate buckets. Compared to Bloom filters, cuckoo filters often: -* support deletions naturally, -* achieve high load factors, -* and can be faster at query time for similar false positive rates. +- support deletions naturally, +- achieve high load factors, +- and can be faster at query time for similar false positive rates. - Insert/query/delete: expected $O(1)$, with occasional relocations - No false negatives (for stored fingerprints, assuming correct implementation) - False positives possible (fingerprints can collide) -### String Matching (Substring/Pattern Search) - -Searching in strings is about finding a pattern of length `m` inside a text of length `n`. The naive method works by checking every alignment, but smarter algorithms exploit structure, especially the fact that mismatches can tell you how far you can skip ahead. - -#### Naive String Search - -Slide the pattern across the text one position at a time; at each position, compare characters until mismatch or full match. - -* Worst case: $O(n \cdot m)$ -* Great when patterns are short, texts are small, or you just want the simplest correct solution. - -#### Knuth–Morris–Pratt (KMP) - -KMP avoids re-checking characters by precomputing a failure function (often called the LPS array: longest proper prefix which is also a suffix). When a mismatch occurs, KMP uses the failure function to shift the pattern without moving backward in the text. - -* Preprocessing: $O(m)$ -* Search: $O(n)$ -* Total: $O(n+m)$ guaranteed - -Why it matters: KMP gives you a **hard guarantee**. No “bad day” inputs that suddenly become quadratic. - -#### Boyer–Moore (BM) - -Boyer–Moore compares from right to left and uses skip heuristics (bad-character and good-suffix rules) to jump the pattern forward by more than one step when mismatches happen. - -* Often fastest in practice for long patterns and large alphabets. -* Worst-case can be higher than linear depending on variant, but real-world performance is excellent. - -Why it matters: BM is what you use when you care about speed on typical text, not just asymptotic guarantees. - -#### Rabin–Karp (RK) - -Rabin–Karp uses a rolling hash. Hash the pattern and each length-`m` window of the text; if hashes match, verify by direct comparison to avoid hash collision errors. - -* Expected: $O(n + m)$ -* Great for multiple patterns (hash them all) or streaming windows. +## Membership Filter Walkthroughs -Why it matters: the rolling hash lets you update the window hash in constant time, which is powerful when you’re scanning huge text or many patterns. - -### Probabilistic & Approximate Search +The earlier overview introduced the guarantees. The following larger examples trace collisions, false positives, and deletion step by step. -Probabilistic search structures are built for one very specific job: **fast membership tests** when you don’t want (or can’t afford) exact storage. Instead of storing full keys, they store compact “evidence” that a key might be present. That trade buys you speed and memory savings, at the cost of occasionally saying *“maybe”* when the real answer is *“no.”* +Probabilistic search structures are built for one very specific job: fast membership tests when you don’t want (or can’t afford) exact storage. Instead of storing full keys, they store compact “evidence” that a key might be present. That trade buys you speed and memory savings, at the cost of occasionally saying “maybe” when the real answer is “no.” The core promise is usually this: -* **“Definitely not present” is trustworthy.** -* **“Maybe present” means “check the real structure if you need certainty.”** +- “Definitely not present” is trustworthy. +- “Maybe present” means “check the real structure if you need certainty.” In practice, these filters sit in front of slower systems (databases, disk caches, network calls). If the filter says “definitely not,” you skip the expensive work. If it says “maybe,” you pay the cost, but only for a smaller set of candidates. -#### Bloom Filter +### Bloom Filter -Space-efficient structure for fast membership tests; answers **“maybe present”** or **“definitely not present”** with a tunable false-positive rate and no false negatives (if built correctly, without deletions). +Space-efficient structure for fast membership tests; answers “maybe present” or “definitely not present” with a tunable false-positive rate and no false negatives (if built correctly, without deletions). -A Bloom filter is basically a **bit array + a few hash functions**. Insertion flips a handful of bits to 1. Lookup checks the same handful of bits. The reason this works is monotonicity: bits only ever go from 0 → 1, so a missing 1-bit is a solid proof the item was never inserted. The flip side is that different items can collide onto the same bits, causing false positives. +A Bloom filter is basically a bit array + a few hash functions. Insertion flips a handful of bits to 1. Lookup checks the same handful of bits. The reason this works is monotonicity: bits only ever go from 0 → 1, so a missing 1-bit is a solid proof the item was never inserted. The flip side is that different items can collide onto the same bits, causing false positives. -**Example inputs and outputs** +Example inputs and outputs: -*Setup* +Setup $$ m = 16 \text{bits}, @@ -1525,7 +1496,7 @@ $$ {"cat", "dog"} $$ -*Example 1* +Example 1 $$ \text{Query: contains("cat")} @@ -1537,7 +1508,7 @@ $$ \text{Output: maybe present (true positive)} $$ -*Example 2* +Example 2 $$ \text{Query: contains("cow")} @@ -1549,7 +1520,7 @@ $$ \text{Output: definitely not present} $$ -*Example 3* +Example 3 $$ \text{Query: contains("eel")} @@ -1561,9 +1532,9 @@ $$ \text{Output: maybe present (false positive)} $$ -**How it works** +How it works: -*Initial state* (all zeros): +Initial state (all zeros): ``` Idx: 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 @@ -1599,14 +1570,14 @@ Query `"cow"` ``` h1(cow) = 1 → bit[1] = 1 h2(cow) = 3 → bit[3] = 1 -h3(cow) = 6 → bit[6] = 0 ❌ +h3(cow) = 6 → bit[6] = 0 Idx: 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 A = [0 1 0 1 0 0 0 1 0 1 0 0 1 0 0 0] ✓ ✓ ✗ ``` -At least one zero → **DEFINITELY NOT PRESENT** +At least one zero → DEFINITELY NOT PRESENT Query `"eel"` @@ -1620,26 +1591,26 @@ A = [0 1 0 1 0 0 0 1 0 1 0 0 1 0 0 0] ✓ ✓ ✓ ``` -All ones → **MAYBE PRESENT** (could be a **false positive**) +All ones → MAYBE PRESENT (could be a false positive) -* Answers: **maybe present** / **definitely not present**; never false negatives (without deletions). -* False-positive rate is tunable via bit-array size **m**, number of hashes **k**, and items **n**; more space & good **k** → lower FPR. -* Time: $O(k)$ per insert/lookup; Space:~m bits. -* No deletions in the basic form; duplicates are harmless (idempotent sets). -* Union = bitwise OR; intersection = bitwise AND (for same m,k,hashes). -* Choose independent, well-mixed hash functions to avoid correlated bits. +- Answers: maybe present / definitely not present; never false negatives (without deletions). +- False-positive rate is tunable via bit-array size m, number of hashes k, and items n; more space & good k → lower FPR. +- Time: $O(k)$ per insert/lookup; Space:~m bits. +- No deletions in the basic form; duplicates are harmless (idempotent sets). +- For equal array sizes and identical hash functions, bitwise OR represents the union. Bitwise AND gives a filter with no false negatives for the intersection, but generally differs from rebuilding a Bloom filter from the exact intersection and may retain extra false positives. +- Choose independent, well-mixed hash functions to avoid correlated bits. A helpful intuition: a Bloom filter is like a nightclub bouncer with a stamp list. If you don’t have the stamp pattern, you’re definitely not in. If you do, you might still be an imposter who happens to match the pattern, but it’s usually rare enough to be worth the speed. -#### Counting Bloom Filter +### Counting Bloom Filter -Bloom filter variant that keeps a small counter per bit so you can **delete** by decrementing; still probabilistic and may have false positives. +Bloom filter variant that keeps a small counter per bit so you can delete by decrementing; still probabilistic and may have false positives. -Counting Bloom filters exist because normal Bloom filters only ever turn bits on. That’s great for proving “definitely not,” but it makes deletions unsafe: turning a bit off might accidentally remove evidence needed for other elements. Counters fix that by tracking **how many inserts contributed** to each position. +Counting Bloom filters exist because normal Bloom filters only ever turn bits on. That’s great for proving “definitely not,” but it makes deletions unsafe: turning a bit off might accidentally remove evidence needed for other elements. Counters fix that by tracking how many inserts contributed to each position. -**Example inputs and outputs** +Example inputs and outputs: -*Setup* +Setup $$ m = 12 \text{counters (each 2–4 bits)}, @@ -1654,7 +1625,7 @@ $$ Then delete `"alpha"`. -*Example 1* +Example 1 $$ \text{Query: contains("alpha")} @@ -1666,7 +1637,7 @@ $$ \text{Output: definitely not present} $$ -*Example 2* +Example 2 $$ \text{Query: contains("beta")} @@ -1678,7 +1649,7 @@ $$ \text{Output: maybe present} $$ -*Example 3* +Example 3 $$ \text{Query: contains("gamma")} @@ -1690,10 +1661,10 @@ $$ \text{Output: definitely not present} $$ -**How it works** +How it works: -Each cell is a **small counter** (e.g. 4-bits, range 0..15). -This allows **deletions**: increment on insert, decrement on delete. +Each cell is a small counter (e.g. 4-bits, range 0..15). +This allows deletions: increment on insert, decrement on delete. Initial state @@ -1763,23 +1734,23 @@ A = [0 0 0 1 0 1 0 0 0 0 0 1] ✗ ✓ ✗ ``` -* Supports **deletion** by decrementing counters; insertion increments. -* Still probabilistic: may return false positives; avoids false negatives **if counters never underflow** and hashes are consistent. -* Space: more than Bloom (a few bits per counter instead of 1). -* Watch for counter **saturation** (caps at max value) and **underflow** (don’t decrement below 0). -* Good for dynamic sets with frequent inserts and deletes. +- Supports deletion by decrementing counters; insertion increments. +- Still probabilistic: false positives remain possible. Avoid false negatives by deleting only known live insertions, preserving exact counter contributions without overflow or underflow, and using consistent hashes. +- Space: more than Bloom (a few bits per counter instead of 1). +- Simply saturating a counter loses multiplicity information. Later decrements can then reach zero while another key still needs that counter. Use sufficient counter capacity or an explicit overflow scheme; checking for underflow alone is insufficient. +- Good for dynamic sets with frequent inserts and deletes. If Bloom filters are stamps, counting Bloom filters are a sign-in sheet: you can erase a name only when you’re sure nobody else relied on that same mark. The counters track that “how many” safely. -#### Cuckoo Filter +### Cuckoo Filter -Hash-table–style filter that stores short **fingerprints** in two possible buckets; supports **insert, lookup, delete** with low false-positive rates and high load factors. +Hash-table–style filter that stores short fingerprints in two possible buckets; supports insert, lookup, delete with low false-positive rates and high load factors. A cuckoo filter feels more like a compact hash table than a Bloom filter. Instead of setting bits, it stores small fingerprints in buckets. The reason you should care is practical: cuckoo filters support deletions naturally (you remove a fingerprint), and they often achieve excellent performance at high occupancy. -**Example inputs and outputs** +Example inputs and outputs: -*Setup* +Setup $$ b = 8 \text{buckets}, @@ -1795,7 +1766,7 @@ $$ Each element is stored as a short fingerprint in one of two candidate buckets. -*Example 1* +Example 1 $$ \text{Query: contains("cat")} @@ -1807,7 +1778,7 @@ $$ \text{Output: maybe present (true positive)} $$ -*Example 2* +Example 2 $$ \text{Query: contains("fox")} @@ -1819,7 +1790,7 @@ $$ \text{Output: definitely not present} $$ -*Example 3 (Deletion)* +Example 3 (Deletion) $$ \text{Operation: remove("dog")} @@ -1831,14 +1802,16 @@ $$ \text{Result: deletion supported directly by removing the fingerprint} $$ -**How it works** +How it works: -Each key `x` → short **fingerprint** `f = FP(x)` +Each key `x` → short fingerprint `f = FP(x)` Two candidate buckets: -* `i1 = H(x) mod b` -* `i2 = (i1 XOR H(f)) mod b` - (`f` can be stored in either bucket; moving between buckets preserves the invariant.) +- `i1 = H(x) mod b` +- `i2 = (i1 XOR H(f)) mod b` + (`f` can be stored in either bucket.) + +For this XOR construction, the bucket count must be a power of two, as it is here (`b = 8`), and the hash contribution is restricted to the bucket-index bits. Applying the same XOR again recovers the other bucket. Arbitrary modulo reduction with a non-power-of-two bucket count does not preserve that property. Start (empty) @@ -1905,11 +1878,50 @@ Check buckets 0 and 7 → fingerprint not found Result: DEFINITELY NOT PRESENT -* Stores **fingerprints**, not full keys; answers **maybe present** / **definitely not present**. -* Supports **deletion** by removing a matching fingerprint from either bucket. -* Very high load factors (often 90%+ with small buckets) and excellent cache locality. -* False-positive rate controlled by fingerprint length (more bits → lower FPR). -* Insertions can trigger **eviction chains**; worst case requires a **rehash/resize**. -* Two buckets per item (or more in variants); lookups check a tiny, fixed set of places. +- Stores fingerprints, not full keys; answers maybe present / definitely not present. +- Delete a matching fingerprint only for a key known to be present. A positive filter query is not enough to establish that a deletion is safe: a false positive can identify another key's fingerprint. See the [original cuckoo-filter paper, deletion section](https://www.cs.cmu.edu/~binfan/papers/conext14_cuckoofilter.pdf). +- Very high load factors (often 90%+ with small buckets) and excellent cache locality. +- False-positive rate controlled by fingerprint length (more bits → lower FPR). +- Insertion can trigger eviction chains or fail. Preserve all previously stored fingerprints on failure and report it to the caller. Rebuilding with new hashes generally requires access to the original keys, which the filter itself does not store. +- Two buckets per item (or more in variants); lookups check a tiny, fixed set of places. A good mental model: cuckoo filters are like a coat check where every coat has exactly two allowed hooks. If both hooks are full, you temporarily move coats around (“cuckooing”) until everyone has a legal spot, or you expand the room. + +## String Matching (Substring/Pattern Search) + +Searching in strings is about finding a pattern of length `m` inside a text of length `n`. The naive method works by checking every alignment, but smarter algorithms exploit structure, especially the fact that mismatches can tell you how far you can skip ahead. + +### Naive String Search + +Slide the pattern across the text one position at a time; at each position, compare characters until mismatch or full match. + +- Worst case: $O(n \cdot m)$ +- Great when patterns are short, texts are small, or you just want the simplest correct solution. + +### Knuth–Morris–Pratt (KMP) + +KMP avoids re-checking characters by precomputing a failure function (often called the LPS array: longest proper prefix which is also a suffix). When a mismatch occurs, KMP uses the failure function to shift the pattern without moving backward in the text. + +- Preprocessing: $O(m)$ +- Search: $O(n)$ +- Total: $O(n+m)$ guaranteed + +Why it matters: KMP gives you a hard guarantee. No “bad day” inputs that suddenly become quadratic. + +### Boyer–Moore (BM) + +Boyer–Moore compares from right to left and uses skip heuristics (bad-character and good-suffix rules) to jump the pattern forward by more than one step when mismatches happen. + +- Often fastest in practice for long patterns and large alphabets. +- Worst-case can be higher than linear depending on variant, but real-world performance is excellent. + +Why it matters: BM is what you use when you care about speed on typical text, not just asymptotic guarantees. + +### Rabin–Karp (RK) + +Rabin–Karp uses a rolling hash. Hash the pattern and each length-`m` window of the text; if hashes match, verify by direct comparison to avoid hash collision errors. + +- For finding the first match, expected time is $O(n+m)$ under suitable hashing assumptions; collisions can make worst-case time $O(nm)$. When reporting all matches and verifying each one, include $O(zm)$ for $z$ verified matches. +- Great for multiple patterns (hash them all) or streaming windows. + +Why it matters: the rolling hash lets you update the window hash in constant time, which is powerful when you’re scanning huge text or many patterns. diff --git a/notes/sorting.md b/notes/sorting.md index b1f9c2f..a3b5853 100644 --- a/notes/sorting.md +++ b/notes/sorting.md @@ -1,24 +1,24 @@ -## Sorting +# Sorting -In the realm of computer science, 'sorting' refers to the process of arranging a collection of items in a specific, predetermined order. This order is based on certain criteria that are defined beforehand. +Sorting arranges a collection of items according to a defined ordering, such as ascending numeric value or a record's timestamp. The comparison rule must be consistent; contradictory comparisons make a sorted result ill-defined. For instance: -* **Numbers** can be sorted according to their numerical values, either in ascending or descending order. -* **Strings** are typically sorted in alphabetical order, similar to how words are arranged in dictionaries. +- Numbers can be sorted according to their numerical values, either in ascending or descending order. +- Strings are typically sorted in alphabetical order, similar to how words are arranged in dictionaries. Sorting isn't limited to just numbers and strings. Virtually any type of object can be sorted, provided there is a clear definition of order for them. What makes sorting worth caring about is that it quietly powers “everyday” features: search bars, leaderboards, timelines, deduplicating data, grouping logs, ranking recommendations, and making reports readable. If you can sort efficiently (and correctly), you can often make a system feel instantly faster and more polished, because so many tasks become simpler once data is ordered. -### Stability in Sorting +## Stability in Sorting Stability, in the context of sorting, refers to preserving the relative order of equal items in the sorted output. In simple terms, when two items have equal keys: -* In an **unstable** sorting algorithm, their order might be reversed in the sorted output. -* In a **stable** sorting algorithm, their relative order remains unchanged. +- In an unstable sorting algorithm, their order might be reversed in the sorted output. +- In a stable sorting algorithm, their relative order remains unchanged. -Stability sounds like a “detail,” but it becomes a superpower the moment you sort objects by multiple keys. A common pattern is: sort by the *secondary* key first, then sort by the *primary* key. That second pass only works as intended if the second sort is stable, because stability preserves the secondary ordering within groups that tie on the primary key. +Stability is useful when sorting records by multiple keys. Sort by the secondary key first, then stably by the primary key. The second pass preserves the secondary ordering among records whose primary keys are equal. Let’s take an analogous list of medieval knights (each with its original 0-based index): @@ -26,33 +26,33 @@ Let’s take an analogous list of medieval knights (each with its original 0-bas [Lancelot0] [Gawain1] [Lancelot2] [Percival3] [Gawain4] ``` -We’ll “sort” them by name, bringing all **Lancelot**s to the front, then **Gawain**, then **Percival**. +Use the custom name order Lancelot, Gawain, Percival for this example. This is not alphabetical order; it keeps the illustration focused on the relative order of equal names. -#### Stable Sort +### Stable Sort -A **stable** sort preserves the left-to-right order of equal items. +A stable sort preserves the left-to-right order of equal items. -I. **Bring every “Lancelot” to the front**, in the order they appeared (index 0 before 2): +I. Bring every “Lancelot” to the front, in the order they appeared (index 0 before 2): ``` [Lancelot0] [Lancelot2] [Gawain1] [Percival3] [Gawain4] ``` -II. **Next, move the two “Gawain”s** ahead of “Percival”, again preserving 1 before 4: +II. Next, move the two “Gawain”s ahead of “Percival”, again preserving 1 before 4: ``` [Lancelot0] [Lancelot2] [Gawain1] [Gawain4] [Percival3] ``` -So the **stable-sorted** sequence is: +So the stable-sorted sequence is: ``` [Lancelot0] [Lancelot2] [Gawain1] [Gawain4] [Percival3] ``` -#### Unstable Sort +### Unstable Sort -An **unstable** sort may reorder equal items arbitrarily. +An unstable sort may reorder equal items arbitrarily. I. When collecting the “Lancelot”s, it might pick index 2 before 0: @@ -66,46 +66,46 @@ II. Later, the two “Gawain”s might swap (4 before 1): [Lancelot2] [Lancelot0] [Gawain4] [Gawain1] [Percival3] ``` -So one possible **unstable-sorted** sequence is: +So one possible unstable-sorted sequence is: ``` [Lancelot2] [Lancelot0] [Gawain4] [Gawain1] [Percival3] ``` -If you then did a second pass (say, sorting by rank or battle-honors) you’d only want to reorder knights of different names, trusting that ties (same-name knights) are still in their intended original order, something only a stable sort guarantees. +If you next sort stably by rank, rank becomes the primary key. Among equal-rank knights, the previously established name order remains intact. A stable sort preserves prior order only among records equal under the current comparison key. A quick rule of thumb: -* If you’re sorting **primitive values** (just numbers), stability often doesn’t matter. -* If you’re sorting **records** (like `(name, age, score)`), stability matters a lot more than people expect. +- If you’re sorting primitive values (just numbers), stability often doesn’t matter. +- If you’re sorting records (like `(name, age, score)`), stability matters a lot more than people expect. And because many real systems sort records, stability is one of those “small” choices that saves you from weird bugs later. -### Bubble Sort +## Bubble Sort -Bubble sort is one of the simplest sorting algorithms. It is often used as an **introductory algorithm** because it is easy to understand, even though it is not efficient for large datasets. +Bubble sort is one of the simplest sorting algorithms. It is often used as an introductory algorithm because it is easy to understand, even though it is not efficient for large datasets. -The name comes from the way **larger elements "bubble up"** to the top (end of the list), just as bubbles rise in water. +The name comes from the way larger elements "bubble up" to the top (end of the list), just as bubbles rise in water. The basic idea: -* Compare **adjacent elements**. -* Swap them if they are in the wrong order. -* Repeat until no swaps are needed. +- Compare adjacent elements. +- Swap them if they are in the wrong order. +- Repeat until no swaps are needed. -Bubble sort is valuable mostly as a teaching tool: it helps you build intuition for comparison-based sorting, swapping, and passes. In real programs, it’s rarely the best choice, but understanding why it’s slow teaches you what *good* sorting algorithms avoid (repeated full scans with tiny progress each pass). +Bubble sort is valuable mostly as a teaching tool: it helps you build intuition for comparison-based sorting, swapping, and passes. In real programs, it’s rarely the best choice, but understanding why it’s slow teaches you what good sorting algorithms avoid (repeated full scans with tiny progress each pass). -**Step-by-Step Walkthrough** +Step-by-Step Walkthrough -1. Start from the **first element**. -2. Compare it with its **neighbor to the right**. -3. If the left is greater, **swap** them. +1. Start from the first element. +2. Compare it with its neighbor to the right. +3. If the left is greater, swap them. 4. Move to the next pair and repeat until the end of the list. -5. After the **first pass**, the largest element is at the end. +5. After the first pass, the largest element is at the end. 6. On each new pass, ignore the elements already in their correct place. 7. Continue until the list is sorted. -**Example Run** +Example Run: We will sort the array: @@ -113,7 +113,7 @@ We will sort the array: [ 5 ][ 1 ][ 4 ][ 2 ][ 8 ] ``` -**Pass 1** +Pass 1 Compare adjacent pairs and push the largest to the end. @@ -133,9 +133,9 @@ Compare 5 and 8 → no swap [ 1 ][ 4 ][ 2 ][ 5 ][ 8 ] ``` -✔ Largest element **8** has bubbled to the end. +Largest element 8 has bubbled to the end. -**Pass 2** +Pass 2 Now we only need to check the first 4 elements. @@ -152,9 +152,9 @@ Compare 4 and 5 → no swap [ 1 ][ 2 ][ 4 ][ 5 ] [8] ``` -✔ Second largest element **5** is now in place. +Second largest element 5 is now in place. -**Pass 3** +Pass 3 Check only the first 3 elements. @@ -168,17 +168,17 @@ Compare 2 and 4 → no swap [ 1 ][ 2 ][ 4 ] [5][8] ``` -✔ Sorted order is now reached. +Sorted order is now reached. -**Final Result** +Final Result: ``` [ 1 ][ 2 ][ 4 ][ 5 ][ 8 ] ``` -**Visual Illustration of Bubble Effect** +Visual Illustration of Bubble Effect -Here’s how the **largest values "bubble up"** to the right after each pass: +Here’s how the largest values "bubble up" to the right after each pass: ``` Pass 1: [ 5 1 4 2 8 ] → [ 1 4 2 5 8 ] @@ -186,58 +186,58 @@ Pass 2: [ 1 4 2 5 ] → [ 1 2 4 5 ] [8] Pass 3: [ 1 2 4 ] → [ 1 2 4 ] [5 8] ``` -Sorted! ✅ +Sorted! -**Optimizations** +Optimizations: -* By keeping track of whether any swaps were made during a pass, Bubble Sort can terminate early if the array is already sorted. This optimization makes Bubble Sort’s **best case** much faster ($O(n)$). +- By keeping track of whether any swaps were made during a pass, Bubble Sort can terminate early if the array is already sorted. This optimization makes Bubble Sort’s best case much faster ($O(n)$). -This “early exit” is bubble sort’s one redeeming trick: when data is already nearly sorted, it doesn’t keep doing pointless work. That idea, detecting when you can stop early, shows up in far better algorithms too. +Early exit avoids further passes after a pass makes no swaps. Nearly sorted input is not always fast: a small element near the end may still need many passes to move left. That idea, detecting when you can stop early, shows up in far better algorithms too. -**Stability** +Stability: -Bubble sort is **stable**. +Bubble sort is stable when it swaps only strictly out-of-order neighbors, leaving equal keys untouched. -* If two elements have the same value, they remain in the same order relative to each other after sorting. -* This is important when sorting complex records where a secondary key matters. +- If two elements have the same value, they remain in the same order relative to each other after sorting. +- This is important when sorting complex records where a secondary key matters. -**Complexity** +Complexity: | Case | Time Complexity | Notes | | ---------------- | --------------- | ---------------------------------------- | -| **Worst Case** | $O(n^2)$ | Array in reverse order | -| **Average Case** | $O(n^2)$ | Typically quadratic comparisons | -| **Best Case** | $O(n)$ | Already sorted + early exit optimization | -| **Space** | $O(1)$ | In-place, requires no extra memory | +| Worst Case | $O(n^2)$ | Array in reverse order | +| Average Case | $O(n^2)$ | Typically quadratic comparisons | +| Best Case | $O(n)$ | Already sorted + early exit optimization | +| Space | $O(1)$ | In-place, requires no extra memory | -**Implementation** +Implementation: -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/sorting/cpp/bubble_sort/src/bubble_sort.cpp) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/sorting/python/bubble_sort/src/bubble_sort.py) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/sorting/cpp/bubble_sort/src/bubble_sort.cpp) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/sorting/python/bubble_sort/src/bubble_sort.py) -### Selection Sort +## Selection Sort Selection sort is another simple sorting algorithm, often introduced right after bubble sort because it is equally easy to understand. -Instead of repeatedly "bubbling" elements, **selection sort works by repeatedly selecting the smallest (or largest) element** from the unsorted portion of the array and placing it into its correct position. +Instead of repeatedly "bubbling" elements, selection sort works by repeatedly selecting the smallest (or largest) element from the unsorted portion of the array and placing it into its correct position. Think of it like arranging books: -* Look through all the books, find the smallest one, and put it first. -* Then, look through the rest, find the next smallest, and put it second. -* Repeat until the shelf is sorted. +- Look through all the books, find the smallest one, and put it first. +- Then, look through the rest, find the next smallest, and put it second. +- Repeat until the shelf is sorted. -Selection sort is a nice contrast to bubble sort: bubble sort makes lots of small swaps, while selection sort makes **very few swaps** but still pays for a full scan each pass. That’s a useful engineering lesson: minimizing swaps doesn’t automatically mean you’re minimizing total work. +Selection sort is a nice contrast to bubble sort: bubble sort makes lots of small swaps, while selection sort makes very few swaps but still pays for a full scan each pass. That’s a useful engineering lesson: minimizing swaps doesn’t automatically mean you’re minimizing total work. -**Step-by-Step Walkthrough** +Step-by-Step Walkthrough -1. Start at the **first position**. -2. Search the **entire unsorted region** to find the smallest element. +1. Start at the first position. +2. Search the entire unsorted region to find the smallest element. 3. Swap it with the element in the current position. 4. Move the boundary of the sorted region one step forward. 5. Repeat until all elements are sorted. -**Example Run** +Example Run: We will sort the array: @@ -245,75 +245,75 @@ We will sort the array: [ 64 ][ 25 ][ 12 ][ 22 ][ 11 ] ``` -**Pass 1** +Pass 1 Find the smallest element in the entire array and put it in the first position. ``` Initial: [ 64 ][ 25 ][ 12 ][ 22 ][ 11 ] -Smallest = 11 -Swap 64 ↔ 11 +Smallest = 11 +Swap 64 ↔ 11 Result: [ 11 ][ 25 ][ 12 ][ 22 ][ 64 ] ``` -✔ The first element is now in its correct place. +The first element is now in its correct place. -**Pass 2** +Pass 2 Find the smallest element in the remaining unsorted region. ``` Start: [ 11 ][ 25 ][ 12 ][ 22 ][ 64 ] -Smallest in [25,12,22,64] = 12 -Swap 25 ↔ 12 +Smallest in [25,12,22,64] = 12 +Swap 25 ↔ 12 Result: [ 11 ][ 12 ][ 25 ][ 22 ][ 64 ] ``` -✔ The second element is now in place. +The second element is now in place. -**Pass 3** +Pass 3 Repeat for the next unsorted region. ``` Start: [ 11 ][ 12 ][ 25 ][ 22 ][ 64 ] -Smallest in [25,22,64] = 22 -Swap 25 ↔ 22 +Smallest in [25,22,64] = 22 +Swap 25 ↔ 22 Result: [ 11 ][ 12 ][ 22 ][ 25 ][ 64 ] ``` -✔ The third element is now in place. +The third element is now in place. -**Pass 4** +Pass 4 Finally, sort the last two. ``` Start: [ 11 ][ 12 ][ 22 ][ 25 ][ 64 ] -Smallest in [25,64] = 25 -Already in correct place → no swap +Smallest in [25,64] = 25 +Already in correct place → no swap Result: [ 11 ][ 12 ][ 22 ][ 25 ][ 64 ] ``` -✔ Array fully sorted. +Array fully sorted. -**Final Result** +Final Result: ``` [ 11 ][ 12 ][ 22 ][ 25 ][ 64 ] ``` -**Visual Illustration of Selection** +Visual Illustration of Selection -Here’s how the **sorted region expands** from left to right: +Here’s how the sorted region expands from left to right: ``` Pass 1: [ 64 25 12 22 11 ] → [ 11 ] [ 25 12 22 64 ] @@ -324,53 +324,53 @@ Pass 4: [ 11 12 22 ][ 25 64 ] → [ 11 12 22 25 ] [ 64 ] At each step: -* The **left region is sorted** ✅ -* The **right region is unsorted** 🔄 +- The left region is sorted +- The right region is unsorted -**Optimizations** +Optimizations: -* Unlike bubble sort, **early exit is not possible** because selection sort always scans the entire unsorted region to find the minimum. -* But it does fewer swaps: **at most (n-1) swaps**, compared to potentially many in bubble sort. +- Unlike bubble sort, early exit is not possible because selection sort always scans the entire unsorted region to find the minimum. +- But it does fewer swaps: at most (n-1) swaps, compared to potentially many in bubble sort. This is why selection sort sometimes appears in environments where writes are expensive (think: limited-write memory). Even there, it’s a trade: fewer writes, but still lots of comparisons. -**Stability** +Stability: -* **Selection sort is NOT stable** in its classic form. -* If two elements are equal, their order may change due to swapping. -* Stability can be achieved by inserting instead of swapping, but this makes the algorithm more complex. +- Selection sort is NOT stable in its classic form. +- If two elements are equal, their order may change due to swapping. +- Stability can be achieved by inserting instead of swapping, but this makes the algorithm more complex. -**Complexity** +Complexity: | Case | Time Complexity | Notes | | ---------------- | --------------- | ------------------------------------------ | -| **Worst Case** | $O(n^2)$ | Scanning full unsorted region every pass | -| **Average Case** | $O(n^2)$ | Quadratic comparisons | -| **Best Case** | $O(n^2)$ | No improvement, still must scan every pass | -| **Space** | $O(1)$ | In-place sorting | +| Worst Case | $O(n^2)$ | Scanning full unsorted region every pass | +| Average Case | $O(n^2)$ | Quadratic comparisons | +| Best Case | $O(n^2)$ | No improvement, still must scan every pass | +| Space | $O(1)$ | In-place sorting | -**Implementation** +Implementation: -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/sorting/cpp/selection_sort/src/selection_sort.cpp) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/sorting/python/selection_sort/src/selection_sort.py) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/sorting/cpp/selection_sort/src/selection_sort.cpp) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/sorting/python/selection_sort/src/selection_sort.py) -### Insertion Sort +## Insertion Sort Insertion sort is a simple, intuitive sorting algorithm that works the way people often sort playing cards in their hands. -It builds the **sorted portion one element at a time**, by repeatedly taking the next element from the unsorted portion and inserting it into its correct position among the already sorted elements. +It builds the sorted portion one element at a time, by repeatedly taking the next element from the unsorted portion and inserting it into its correct position among the already sorted elements. -Insertion sort is the first “small” sort that feels genuinely useful in practice. Why? Because many real datasets are already *almost* sorted, and insertion sort loves that. It’s also a common helper inside faster algorithms: when the subproblems become tiny, insertion sort finishes them efficiently with low overhead. +Insertion sort is the first “small” sort that feels genuinely useful in practice. Why? Because many real datasets are already almost sorted, and insertion sort loves that. It’s also a common helper inside faster algorithms: when the subproblems become tiny, insertion sort finishes them efficiently with low overhead. The basic idea: -1. Start with the **second element** (the first element by itself is trivially sorted). -2. Compare it with elements to its **left**. +1. Start with the second element (the first element by itself is trivially sorted). +2. Compare it with elements to its left. 3. Shift larger elements one position to the right. 4. Insert the element into the correct spot. 5. Repeat until all elements are processed. -**Example Run** +Example Run: We will sort the array: @@ -378,7 +378,7 @@ We will sort the array: [ 12 ][ 11 ][ 13 ][ 5 ][ 6 ] ``` -**Pass 1: Insert 11** +Pass 1: Insert 11 Compare 11 with 12 → shift 12 right → insert 11 before it. @@ -388,9 +388,9 @@ Action: Insert 11 before 12 After: [ 11 ][ 12 ][ 13 ][ 5 ][ 6 ] ``` -✔ Sorted portion: $[11, 12]$ +Sorted portion: $[11, 12]$ -**Pass 2: Insert 13** +Pass 2: Insert 13 Compare 13 with 12 → already greater → stays in place. @@ -399,9 +399,9 @@ Before: [ 11 ][ 12 ][ 13 ][ 5 ][ 6 ] After: [ 11 ][ 12 ][ 13 ][ 5 ][ 6 ] ``` -✔ Sorted portion: [11, 12, 13] +Sorted portion: [11, 12, 13] -**Pass 3: Insert 5** +Pass 3: Insert 5 Compare 5 with 13 → shift 13 Compare 5 with 12 → shift 12 @@ -414,9 +414,9 @@ Action: Move 13 → Move 12 → Move 11 → Insert 5 After: [ 5 ][ 11 ][ 12 ][ 13 ][ 6 ] ``` -✔ Sorted portion: [5, 11, 12, 13] +Sorted portion: [5, 11, 12, 13] -**Pass 4: Insert 6** +Pass 4: Insert 6 Compare 6 with 13 → shift 13 Compare 6 with 12 → shift 12 @@ -429,15 +429,15 @@ Action: Move 13 → Move 12 → Move 11 → Insert 6 After: [ 5 ][ 6 ][ 11 ][ 12 ][ 13 ] ``` -✔ Sorted! +Sorted! -**Final Result** +Final Result: ``` [ 5 ][ 6 ][ 11 ][ 12 ][ 13 ] ``` -**Visual Growth of Sorted Region** +Visual Growth of Sorted Region ``` Start: [ 12 | 11 13 5 6 ] @@ -447,54 +447,86 @@ Pass 3: [ 5 11 12 13 | 6 ] Pass 4: [ 5 6 11 12 13 ] ``` -✔ The **bar ( | )** shows the boundary between **sorted** and **unsorted**. +The bar ( | ) shows the boundary between sorted and unsorted. -**Optimizations** +Optimizations: -* Efficient for **small arrays**. -* Useful as a **helper inside more complex sorts** (e.g., Quick Sort or Merge Sort) for small subarrays. -* Can be optimized with **binary search** to find insertion positions faster (but shifting still takes linear time). +- Efficient for small arrays. +- Useful as a helper inside more complex sorts (e.g., Quick Sort or Merge Sort) for small subarrays. +- Can be optimized with binary search to find insertion positions faster (but shifting still takes linear time). That last bullet is a subtle but important point: binary search reduces comparisons, not movement. If moving elements is the expensive part (and it often is), the big win comes from having fewer shifts, which happens naturally when the array is already nearly sorted. -**Stability** +Stability: -Insertion sort is **stable** (equal elements keep their relative order). +Insertion sort is stable when it shifts only strictly larger keys and inserts the new item after existing equal keys. -**Complexity** +Complexity: | Case | Time Complexity | Notes | | ---------------- | --------------- | -------------------------------------------------- | -| **Worst Case** | $O(n^2)$ | Reverse-sorted input | -| **Average Case** | $O(n^2)$ | | -| **Best Case** | $O(n)$ | Already sorted input , only comparisons, no shifts | -| **Space** | $O(1)$ | In-place | +| Worst Case | $O(n^2)$ | Reverse-sorted input | +| Average Case | $O(n^2)$ | | +| Best Case | $O(n)$ | Already sorted input: only comparisons, no shifts | +| Space | $O(1)$ | In-place | -**Implementation** +Implementation: -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/sorting/cpp/insertion_sort/src/insertion_sort.cpp) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/sorting/python/insertion_sort/src/insertion_sort.py) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/sorting/cpp/insertion_sort/src/insertion_sort.cpp) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/sorting/python/insertion_sort/src/insertion_sort.py) -### Quick Sort +## Merge Sort -Quick Sort is a **divide-and-conquer** algorithm. Unlike bubble sort or selection sort, which work by repeatedly scanning the whole array, Quick Sort works by **partitioning** the array into smaller sections around a "pivot" element and then sorting those sections independently. +Merge sort divides the input into two halves, sorts each half recursively, and merges the two sorted results. Unlike quicksort, its split does not depend on the values, so the recursion stays balanced. -It is one of the **fastest sorting algorithms in practice**, widely used in libraries and systems. +To merge, compare the first unconsumed item from each half and copy the smaller one to the output. When one half is exhausted, append the remainder of the other half. Each item is processed once during a merge. -Quick sort’s magic is that it does a lot of work *once* up front (partitioning), and that one operation pays off by shrinking the problem dramatically. The danger is also in that same choice: if your pivots are consistently bad, you stop shrinking the problem effectively and performance collapses. +Example: + +```text +Input: [38, 27, 43, 3] +Split: [38, 27] [43, 3] +Split again: [38] [27] [43] [3] +Merge pairs: [27, 38] [3, 43] +Final merge: + 3 < 27 -> [3] + 27 < 43 -> [3, 27] + 38 < 43 -> [3, 27, 38] + append 43 -> [3, 27, 38, 43] +``` + +A subarray of length zero or one is already sorted. For larger inputs, the recurrence is $T(n)=2T(n/2)+\Theta(n)$, giving $\Theta(n\log n)$ time for the standard version in all input-order cases. + +- Stability: take from the left half first when keys compare equal. +- Auxiliary space: $O(n)$ for array merging, plus $O(\log n)$ recursive stack space. +- Optimization: reuse one merge buffer rather than allocate a new buffer at every call. A bottom-up version merges runs of length 1, 2, 4, and so on without recursion. +- Applications: stable record sorting and external sorting, where sequentially merging sorted runs is useful. + +Implementations: + +- [C++](../src/sorting/cpp/merge_sort/) +- [Python](../src/sorting/python/merge_sort/) + +## Quick Sort + +Quick Sort is a divide-and-conquer algorithm. Unlike bubble sort or selection sort, which work by repeatedly scanning the whole array, Quick Sort works by partitioning the array into smaller sections around a "pivot" element and then sorting those sections independently. + +It is one of the fastest sorting algorithms in practice, widely used in libraries and systems. + +Partitioning puts a pivot in place and separates the remaining items into smaller subproblems. Balanced partitions reduce the problem quickly; consistently extreme pivots leave one large subproblem and can produce quadratic total work. The basic idea: -1. Choose a **pivot element** (commonly the last, first, middle, or random element). +1. Choose a pivot element (commonly the last, first, middle, or random element). 2. Rearrange (partition) the array so that: -* All elements **smaller than the pivot** come before it. -* All elements **larger than the pivot** come after it. + - All elements smaller than the pivot come before it. + - All elements larger than the pivot come after it. -3. The pivot is now in its **final sorted position**. -4. Recursively apply Quick Sort to the **left subarray** and **right subarray**. +3. In the pivot-placement scheme illustrated here, the pivot is now in its final position. Other schemes, such as Hoare partitioning, return a split boundary rather than a finalized pivot. Equal keys must be handled consistently so recursion always shrinks. +4. Recursively apply Quick Sort to the left subarray and right subarray. -**Example Run** +Example Run: We will sort the array: @@ -502,143 +534,143 @@ We will sort the array: [ 10 ][ 80 ][ 30 ][ 90 ][ 40 ][ 50 ][ 70 ] ``` -**Step 1: Choose Pivot (last element = 70)** +Step 1: Choose Pivot (last element = 70) Partition around 70. ``` Initial: [ 10 ][ 80 ][ 30 ][ 90 ][ 40 ][ 50 ][ 70 ] -→ Elements < 70: [ 10, 30, 40, 50 ] +→ Elements < 70: [ 10, 30, 40, 50 ] → Pivot (70) goes here ↓ Sorted split: [ 10 ][ 30 ][ 40 ][ 50 ][ 70 ][ 90 ][ 80 ] ``` -*(ordering of right side may vary during partition; only pivot’s position is guaranteed)* +(ordering of right side may vary during partition; only pivot’s position is guaranteed) -✔ Pivot (70) is in correct place. +Pivot (70) is in correct place. -**Step 2: Left Subarray [10, 30, 40, 50]** +Step 2: Left Subarray [10, 30, 40, 50] Choose pivot = 50. ``` [ 10 ][ 30 ][ 40 ][ 50 ] → pivot = 50 -→ Elements < 50: [10, 30, 40] -→ Pivot at correct place +→ Elements < 50: [10, 30, 40] +→ Pivot at correct place Result: [ 10 ][ 30 ][ 40 ][ 50 ] ``` -✔ Pivot (50) fixed. +Pivot (50) fixed. -**Step 3: Left Subarray of Left [10, 30, 40]** +Step 3: Left Subarray of Left [10, 30, 40] Choose pivot = 40. ``` [ 10 ][ 30 ][ 40 ] → pivot = 40 -→ Elements < 40: [10, 30] -→ Pivot at correct place +→ Elements < 40: [10, 30] +→ Pivot at correct place Result: [ 10 ][ 30 ][ 40 ] ``` -✔ Pivot (40) fixed. +Pivot (40) fixed. -**Step 4: [10, 30]** +Step 4: [10, 30] Choose pivot = 30. ``` [ 10 ][ 30 ] → pivot = 30 -→ Elements < 30: [10] +→ Elements < 30: [10] Result: [ 10 ][ 30 ] ``` -✔ Sorted. +The left side is sorted. On the right, partition `[90, 80]` around pivot `80` to obtain `[80, 90]`. Both recursive branches are now complete. -**Final Result** +Final Result: ``` [ 10 ][ 30 ][ 40 ][ 50 ][ 70 ][ 80 ][ 90 ] ``` -**Visual Partition Illustration** +Visual Partition Illustration Here’s how the array gets partitioned step by step: ``` -Pass 1: [ 10 80 30 90 40 50 | 70 ] - ↓ pivot = 70 +Pass 1: [ 10 80 30 90 40 50 | 70 ] + ↓ pivot = 70 [ 10 30 40 50 | 70 | 90 80 ] -Pass 2: [ 10 30 40 | 50 ] [70] [90 80] - ↓ pivot = 50 +Pass 2: [ 10 30 40 | 50 ] [70] [90 80] + ↓ pivot = 50 [ 10 30 40 | 50 ] [70] [90 80] -Pass 3: [ 10 30 | 40 ] [50] [70] [90 80] - ↓ pivot = 40 +Pass 3: [ 10 30 | 40 ] [50] [70] [90 80] + ↓ pivot = 40 [ 10 30 | 40 ] [50] [70] [90 80] -Pass 4: [ 10 | 30 ] [40] [50] [70] [90 80] - ↓ pivot = 30 +Pass 4: [ 10 | 30 ] [40] [50] [70] [90 80] + ↓ pivot = 30 [ 10 | 30 ] [40] [50] [70] [90 80] ``` -✔ Each pivot splits the problem smaller and smaller until fully sorted. +Each pivot splits the problem smaller and smaller until fully sorted. -**Optimizations** +Optimizations: -* **Pivot Choice:** Choosing a good pivot (e.g., median or random) improves performance. -* **Small Subarrays:** For very small partitions, switch to Insertion Sort for efficiency. -* **Tail Recursion:** Can optimize recursion depth. +- Pivot Choice: Choosing a good pivot (e.g., median or random) improves performance. +- Small Subarrays: For very small partitions, switch to Insertion Sort for efficiency. +- Recursion depth: recurse on the smaller partition and loop over the larger partition to guarantee $O(\log n)$ stack space even when partitions are uneven. These are not “nice-to-haves”, they’re what separates textbook quick sort from industrial quick sort. Real implementations work hard to avoid worst-case behavior, because worst-case quick sort can be painfully slow on already-sorted or adversarial inputs. -**Stability** +Stability: -* Quick Sort is **not stable** by default (equal elements may be reordered). -* Stable versions exist, but require modifications. +- Quick Sort is not stable by default (equal elements may be reordered). +- Stable versions exist, but require modifications. -**Complexity** +Complexity: | Case | Time Complexity | Notes | | ---------------- | --------------- | ------------------------------------------------------------------ | -| **Worst Case** | $O(n^2)$ | Poor pivot choices (e.g., always smallest/largest in sorted array) | -| **Average Case** | $O(n \log n)$ | Expected performance, very fast in practice | -| **Best Case** | $O(n \log n)$ | Balanced partitions | -| **Space** | $O(\log n)$ | Due to recursion stack | +| Worst Case | $O(n^2)$ | Poor pivot choices (e.g., always smallest/largest in sorted array) | +| Average Case | $O(n \log n)$ | Expected performance, very fast in practice | +| Best Case | $O(n \log n)$ | Balanced partitions | +| Space | $O(\log n)$ expected, $O(n)$ worst | Ordinary recursive implementation | -**Implementation** +Implementation: -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/sorting/cpp/quick_sort/src/quick_sort.cpp) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/sorting/python/quick_sort/src/quick_sort.py) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/sorting/cpp/quick_sort/src/quick_sort.cpp) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/sorting/python/quick_sort/src/quick_sort.py) -### Heap sort +## Heap Sort -Heap Sort is a **comparison-based sorting algorithm** that uses a special data structure called a **binary heap**. -It is efficient, with guaranteed $O(n \log n)$ performance, and sorts **in-place** (no extra array needed). +Heap Sort is a comparison-based sorting algorithm that uses a special data structure called a binary heap. +It is efficient, with guaranteed $O(n \log n)$ performance, and sorts in-place (no extra array needed). -Heap sort is what you reach for when you want predictability. It doesn’t have quick sort’s “usually blazing fast, occasionally terrible” personality. It’s consistently good, and that reliability matters in systems where worst-case latency is a big deal. +Heap sort is useful when predictable worst-case time and low auxiliary memory matter. Its $O(n\log n)$ worst-case bound avoids quicksort's dependence on pivot quality. The basic idea: -1. **Build a max heap** from the input array. +1. Build a max heap from the input array. -* In a max heap, every parent is greater than its children. -* This ensures the **largest element is at the root** (first index). + - In a max heap, every parent is greater than or equal to its children. + - This ensures the largest element is at the root (first index). -2. Swap the **root (largest element)** with the **last element** of the heap. +2. Swap the root (largest element) with the last element of the heap. 3. Reduce the heap size by 1 (ignore the last element, which is now in place). -4. **Heapify** (restore heap property). +4. Heapify (restore heap property). 5. Repeat until all elements are sorted. -**Example Run** +Example Run: We will sort the array: @@ -646,7 +678,7 @@ We will sort the array: [ 4 ][ 10 ][ 3 ][ 5 ][ 1 ] ``` -**Step 1: Build Max Heap** +Step 1: Build Max Heap Binary tree view: @@ -672,7 +704,7 @@ Heap: Array: [ 10 ][ 5 ][ 3 ][ 4 ][ 1 ] ``` -**Step 2: Swap Root with Last** +Step 2: Swap Root with Last Swap 10 ↔ 1 → largest (10) moves to correct final place. @@ -693,7 +725,7 @@ Heap: Array: [ 5 ][ 4 ][ 3 ][ 1 ][ 10 ] ``` -**Step 3: Swap Root with Last** +Step 3: Swap Root with Last Swap 5 ↔ 1. @@ -712,7 +744,7 @@ Heap: Array: [ 4 ][ 1 ][ 3 ][ 5 ][ 10 ] ``` -**Step 4: Swap Root with Last** +Step 4: Swap Root with Last Swap 4 ↔ 3. @@ -731,7 +763,7 @@ Heap: Array: [ 3 ][ 1 ][ 4 ][ 5 ][ 10 ] ``` -**Step 5: Swap Root with Last** +Step 5: Swap Root with Last Swap 3 ↔ 1. @@ -739,15 +771,15 @@ Swap 3 ↔ 1. [ 1 ][ 3 ][ 4 ][ 5 ][ 10 ] ``` -✔ Sorted array achieved. +Sorted array achieved. -**Final Result** +Final Result: ``` [ 1 ][ 3 ][ 4 ][ 5 ][ 10 ] ``` -**Visual Progress** +Visual Progress: ``` Initial: [ 4 10 3 5 1 ] @@ -759,48 +791,48 @@ Step 4: [ 1 | 3 4 5 10 ] Sorted: [ 1 3 4 5 10 ] ``` -✔ Each step places the largest element into its correct final position. +Each step places the largest element into its correct final position. -**Optimizations** +Optimizations: -* Building the heap can be done in **O(n)** time using bottom-up heapify. -* After building, each extract-max + heapify takes **O(log n)**. +- Building the heap can be done in O(n) time using bottom-up heapify. +- After building, each extract-max + heapify takes O(log n). -**Stability** +Stability: -Heap sort is **not stable**. Equal elements may not preserve their original order because of swaps. +Heap sort is not stable. Equal elements may not preserve their original order because of swaps. -**Complexity** +Complexity: | Case | Time Complexity | Notes | | ---------------- | --------------- | ---------------------- | -| **Worst Case** | $O(n \log n)$ | | -| **Average Case** | $O(n \log n)$ | | -| **Best Case** | $O(n \log n)$ | No early exit possible | -| **Space** | $O(1)$ | In-place | +| Worst Case | $O(n \log n)$ | | +| Average Case | $O(n \log n)$ | | +| Best Case | $O(n \log n)$ | Upper bound; equal keys can allow linear time with early-stopping sift-down | +| Space | $O(1)$ | In-place | -**Implementation** +Implementation: -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/sorting/cpp/heap_sort/src/heap_sort.cpp) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/sorting/python/heap_sort/src/heap_sort.py) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/sorting/cpp/heap_sort/src/heap_sort.cpp) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/sorting/python/heap_sort/src/heap_sort.py) -### Radix Sort +## Radix Sort -Radix Sort is a **non-comparison-based sorting algorithm**. -Instead of comparing elements directly, it processes numbers digit by digit, from either the **least significant digit (LSD)** or the **most significant digit (MSD)**, using a stable intermediate sorting algorithm (commonly **Counting Sort**). +Radix Sort is a non-comparison-based sorting algorithm. +Instead of comparing elements directly, it processes numbers digit by digit, from either the least significant digit (LSD) or the most significant digit (MSD), using digit-based distribution. LSD radix sort requires a stable intermediate sort, commonly counting sort; MSD radix sort recursively sorts digit buckets and may use an unstable in-place partition. -Radix sort is the “cheat code” you unlock when your values have structure (digits, bytes, fixed-length keys). Comparisons are expensive because they only tell you “less than or greater than.” Radix sort uses richer information: digits let you bucket items directly. That’s how it reaches near-linear performance for fixed-width integers. +Radix sort uses the structure of keys, such as digits or bytes, to distribute items directly into buckets. It can avoid the comparison-sorting lower bound because it performs more informative operations than pairwise comparisons. -Because it avoids comparisons, Radix Sort can achieve **linear time complexity** in many cases. +Because it avoids comparisons, Radix Sort can achieve linear time complexity in many cases. The basic idea: -1. Pick a **digit position** (units, tens, hundreds, etc.). -2. Sort the array by that digit using a **stable sorting algorithm**. +1. Pick a digit position (units, tens, hundreds, etc.). +2. Sort the array by that digit using a stable sorting algorithm. 3. Move to the next digit. 4. Repeat until all digits are processed. -**Example Run (LSD Radix Sort)** +Example Run (LSD Radix Sort) We will sort the array: @@ -808,7 +840,7 @@ We will sort the array: [ 170 ][ 45 ][ 75 ][ 90 ][ 802 ][ 24 ][ 2 ][ 66 ] ``` -**Step 1: Sort by 1s place (units digit)** +Step 1: Sort by 1s place (units digit) ``` Original: [170, 45, 75, 90, 802, 24, 2, 66] @@ -823,7 +855,7 @@ By 1s digit: Result: [170][90][802][2][24][45][75][66] ``` -**Step 2: Sort by 10s place** +Step 2: Sort by 10s place ``` [170][90][802][2][24][45][75][66] @@ -839,7 +871,7 @@ By 10s digit: Result: [802][2][24][45][66][170][75][90] ``` -**Step 3: Sort by 100s place** +Step 3: Sort by 100s place ``` [802][2][24][45][66][170][75][90] @@ -852,13 +884,13 @@ By 100s digit: Result: [2][24][45][66][75][90][170][802] ``` -**Final Result** +Final Result: ``` [ 2 ][ 24 ][ 45 ][ 66 ][ 75 ][ 90 ][ 170 ][ 802 ] ``` -**Visual Process** +Visual Process: ``` Step 1 (1s): [170 90 802 2 24 45 75 66] @@ -866,56 +898,55 @@ Step 2 (10s): [802 2 24 45 66 170 75 90] Step 3 (100s): [2 24 45 66 75 90 170 802] ``` -✔ Each pass groups by digit → final sorted order. - -**LSD vs MSD** +Each pass groups by digit → final sorted order. -* **LSD (Least Significant Digit first):** Process digits from right (units) to left (hundreds). Most common, simpler. -* **MSD (Most Significant Digit first):** Process from left to right, useful for variable-length data like strings. +LSD vs MSD -**Stability** +- LSD (Least Significant Digit first): Process digits from right (units) to left (hundreds). Most common, simpler. +- MSD (Most Significant Digit first): Process from left to right, useful for variable-length data like strings. -* Radix Sort **is stable**, because it relies on a stable intermediate sort (like Counting Sort). -* Equal elements remain in the same order across passes. +Stability: -**Complexity** +- The LSD variant shown is stable because each digit pass is stable. Stability of MSD variants depends on their implementation. +- Equal elements remain in the same order across passes. -* **Time Complexity:** $O(n \cdot k)$ +Complexity: - * $n$ = number of elements - * $k$ = number of digits (or max digit length) +Let $n$ be the number of keys, $d$ the number of digit positions, and $b$ the radix (number of possible digit values). -* **Space Complexity:** $O(n + k)$ (depends on the stable sorting method used, e.g., Counting Sort). +- Time: $O(d(n+b))$ with counting sort for each digit. +- Auxiliary space: $O(n+b)$, reusing the output and count arrays between passes. +- For fixed $d$ and $b$, time is $O(n)$. Digit extraction is assumed constant time. -* For integers with fixed number of digits, Radix Sort can be considered **linear time**. +The example uses non-negative base-10 integers, treating missing leading digits as zero. Signed integers require an order-preserving key transformation or separate handling of signs; strings need a defined alphabet and end-of-string ordering. -**Implementation** +Implementation: -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/sorting/cpp/radix_sort/src/radix_sort.cpp) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/sorting/python/radix_sort/src/radix_sort.py) +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/sorting/cpp/radix_sort/src/radix_sort.cpp) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/sorting/python/radix_sort/src/radix_sort.py) -### Counting Sort +## Counting Sort -Counting Sort is a **non-comparison-based sorting algorithm** that works by **counting occurrences** of each distinct element and then calculating their positions in the output array. +Counting Sort is a non-comparison-based sorting algorithm that works by counting occurrences of each distinct element and then calculating their positions in the output array. It is especially efficient when: -* The input values are integers. -* The **range of values (k)** is not significantly larger than the number of elements (n). +- The input values are integers. +- The range of values (k) is not significantly larger than the number of elements (n). Counting sort is the engine that often makes radix sort practical. It’s incredibly fast when the value range is reasonable, but it’s not a general-purpose sort: it relies on keys being integers in a known range. When it fits, it feels like you’re sorting by flipping a few switches rather than doing a lot of comparisons. The basic idea: -1. Find the **range** of the input (min to max). -2. Create a **count array** to store the frequency of each number. -3. Modify the count array to store **prefix sums** (cumulative counts). +1. Find the range of the input (min to max). +2. Create a count array to store the frequency of each number. +3. Modify the count array to store prefix sums (cumulative counts). -* This gives the final position of each element. + - The cumulative count identifies the end of each key's output region. 4. Place elements into the output array in order, using the count array. -**Example Run** +Example Run: We will sort the array: @@ -923,23 +954,23 @@ We will sort the array: [ 4 ][ 2 ][ 2 ][ 8 ][ 3 ][ 3 ][ 1 ] ``` -**Step 1: Count Frequencies** +Step 1: Count Frequencies ``` Elements: 1 2 3 4 5 6 7 8 Counts: 1 2 2 1 0 0 0 1 ``` -**Step 2: Prefix Sums** +Step 2: Prefix Sums ``` Elements: 1 2 3 4 5 6 7 8 Counts: 1 3 5 6 6 6 6 7 ``` -✔ Now each number tells us the **last index position** where that value should go. +Each prefix count gives the number of elements less than or equal to that value. Its last available zero-based output position is `count[value] - 1`. After placing an element there, decrement its count. -**Step 3: Place Elements** +Step 3: Place Elements Process input from right → left (for stability). @@ -955,13 +986,13 @@ Place 2 → index 1 Place 4 → index 5 ``` -**Final Result** +Final Result: ``` [ 1 ][ 2 ][ 2 ][ 3 ][ 3 ][ 4 ][ 8 ] ``` -**Visual Process** +Visual Process: ``` Step 1 Count: [0,1,2,2,1,0,0,0,1] @@ -969,38 +1000,41 @@ Step 2 Prefix: [0,1,3,5,6,6,6,6,7] Step 3 Output: [1,2,2,3,3,4,8] ``` -✔ Linear-time sorting by counting positions. +Linear-time sorting by counting positions. -**Stability** +Stability: -Counting Sort is **stable** if we place elements **from right to left** into the output array. +With cumulative end positions, counting sort is stable when it scans the input right to left and decrements the count after each placement. Negative keys can be handled by indexing the count array with `key - minimum`; then $k = maximum - minimum + 1$. Return immediately for an empty input before computing its minimum or maximum. -**Complexity** +Complexity: | Case | Time Complexity | Notes | | ----------- | --------------- | ------------------------------------------- | -| **Overall** | $O(n + k)$ | $n$ = number of elements, $k$ = value range | -| **Space** | $O(n + k)$ | Extra array for counts + output | +| Overall | $O(n + k)$ | $n$ = number of elements, $k$ = value range | +| Space | $O(n + k)$ | Extra array for counts + output | + +Implementation: -**Implementation** +- [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/sorting/cpp/counting_sort/src/counting_sort.cpp) +- [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/sorting/python/counting_sort/src/counting_sort.py) -* [C++](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/sorting/cpp/counting_sort/src/counting_sort.cpp) -* [Python](https://github.com/djeada/Algorithms-And-Data-Structures/blob/master/src/sorting/python/counting_sort/src/counting_sort.py) +## Comparison Table -### Comparison Table +Space means auxiliary storage, including the recursion stack. Average-case bounds require input or randomization assumptions; the table does not predict elapsed time. -Below is a consolidated **side-by-side comparison** of all the sorts we’ve covered so far: +Below is a consolidated side-by-side comparison of all the sorts we’ve covered so far: -Before using any table like this, remember what it can and can’t tell you. Big-O is a great compass, but not a GPS: constant factors, memory behavior, stability requirements, and input shape (random vs nearly sorted vs adversarial) all affect real-world speed. The most “correct” choice is the one that matches your constraints, not the one with the prettiest asymptotic line. +Use complexity bounds together with stability, memory requirements, and input characteristics. Constant factors and cache behavior also affect performance, so an asymptotically good bound alone does not identify the best implementation for every workload. | Algorithm | Best Case | Average | Worst Case | Space | Stable? | Notes | | ------------------ | ------------- | ------------- | ------------- | ----------- | ------- | ---------------------- | -| **Bubble Sort** | $O(n)$ | $O(n^2)$ | $O(n^2)$ | $O(1)$ | Yes | Simple, slow | -| **Selection Sort** | $O(n^2)$ | $O(n^2)$ | $O(n^2)$ | $O(1)$ | No | Few swaps | -| **Insertion Sort** | $O(n)$ | $O(n^2)$ | $O(n^2)$ | $O(1)$ | Yes | Good for small inputs | -| **Quick Sort** | $O(n \log n)$ | $O(n \log n)$ | $O(n^2)$ | $O(\log n)$ | No | Very fast in practice | -| **Heap Sort** | $O(n \log n)$ | $O(n \log n)$ | $O(n \log n)$ | $O(1)$ | No | Guaranteed performance | -| **Counting Sort** | $O(n + k)$ | $O(n + k)$ | $O(n + k)$ | $O(n + k)$ | Yes | Integers only | -| **Radix Sort** | $O(nk)$ | $O(nk)$ | $O(nk)$ | $O(n + k)$ | Yes | Uses Counting Sort | - -**Use insertion sort for tiny or nearly-sorted arrays, quick sort for general high speed (with good pivot strategy), heap sort for worst-case guarantees, and counting/radix when your keys make it possible.** +| Bubble Sort | $O(n)$ | $O(n^2)$ | $O(n^2)$ | $O(1)$ | Yes | Simple, slow | +| Selection Sort | $O(n^2)$ | $O(n^2)$ | $O(n^2)$ | $O(1)$ | No | Few swaps | +| Insertion Sort | $O(n)$ | $O(n^2)$ | $O(n^2)$ | $O(1)$ | Yes | Good for small inputs | +| Quick Sort | $O(n\log n)$ | $O(n\log n)$ | $O(n^2)$ | $O(n)$ worst | No | Smaller-side recursion reduces stack to $O(\log n)$ | +| Merge Sort | $O(n\log n)$ | $O(n\log n)$ | $O(n\log n)$ | $O(n)$ | Yes | Stable merging | +| Heap Sort | $O(n \log n)$ | $O(n \log n)$ | $O(n \log n)$ | $O(1)$ | No | Guaranteed performance | +| Counting Sort | $O(n + k)$ | $O(n + k)$ | $O(n + k)$ | $O(n + k)$ | Yes | Integers only | +| LSD Radix Sort | $O(d(n+b))$ | $O(d(n+b))$ | $O(d(n+b))$ | $O(n+b)$ | Yes | $d$ digits, radix $b$ | + +Use insertion sort for tiny or nearly-sorted arrays, quick sort for general high speed (with good pivot strategy), heap sort for worst-case guarantees, and counting/radix when your keys make it possible.