Skip to content

Commit b0ea47d

Browse files
committed
Complete Chapter 14 (Sorting & QuickSelect): merge sort, quickselect, custom comparators, LIS-by-sort
1 parent c7b2d77 commit b0ea47d

10 files changed

Lines changed: 1742 additions & 0 deletions

‎CodingInterviewFightClub/src/SUMMARY.md‎

Lines changed: 10 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -154,3 +154,13 @@
154154
- [13.5 Search Suggestion System](ch13-tries/search-suggestion-system.md)
155155
- [13.6 Word Squares](ch13-tries/word-squares.md)
156156
- [13.7 Design Auto Complete System](ch13-tries/design-autocomplete-system.md)
157+
158+
- [14. Sorting & QuickSelect](ch14-sorting/index.md)
159+
- [14.0 Pattern Primer: Sort as Preprocessing](ch14-sorting/pattern-primer.md)
160+
- [14.1 Merge Sort](ch14-sorting/merge-sort.md)
161+
- [14.2 Kth Largest Element](ch14-sorting/kth-largest-element.md)
162+
- [14.3 K Closest Points To Origin](ch14-sorting/k-closest-points-to-origin.md)
163+
- [14.4 Largest Number](ch14-sorting/largest-number.md)
164+
- [14.5 H-Index](ch14-sorting/h-index.md)
165+
- [14.6 Russian Doll Envelopes](ch14-sorting/russian-doll-envelopes.md)
166+
- [14.7 Top K Frequent Elements (QuickSelect)](ch14-sorting/top-k-frequent-elements-quickselect.md)
Lines changed: 162 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,162 @@
1+
# 14.5 H-Index
2+
3+
> **Source:** [`src/main/kotlin/sorting/HIndex.kt`](https://github.com/arpanpathak/AdvancedAlgorithmPatterns/blob/main/src/main/kotlin/sorting/HIndex.kt)
4+
> **Pattern:** sort + scan · **Core page**
5+
6+
## The Problem
7+
8+
Given an array `citations` where `citations[i]` is the citation count of paper `i`, return the **h-index**: the largest `h` such that **at least `h` papers have at least `h` citations**.
9+
10+
- Constraints: $1 \le n \le 5000$; $0 \le citations[i] \le 1000$.
11+
12+
## Examples
13+
14+
```
15+
Input: citations = [3,0,6,1,5] -> Output: 3 (papers with >= 3 citations: 3,6,5)
16+
Input: citations = [1,3,1] -> Output: 1
17+
```
18+
19+
## Intuition — sort descending, then "rank vs citations" is the check
20+
21+
Sort descending. Now `citations[i]` is the citation count of the `i+1`-th most-cited paper, and the h-index definition becomes a single scan:
22+
23+
> `h` is the largest value where `citations[i] >= i + 1` still holds.
24+
25+
Walk the sorted array; the first position where `citations[i] < i + 1` breaks the run — and the answer is `i` (that many papers met the bar). If no break, every paper clears the bar, and the answer is `n`.
26+
27+
**Why does the sorted scan capture the definition?** "At least h papers with ≥ h citations" — in descending order, that's "the first h papers all have ≥ h citations". The scan finds the largest such h by checking the boundary position where the requirement fails: papers `0..i-1` have ≥ i citations, paper `i` doesn't.
28+
29+
**Why not test all h?** You *could* binary search h or count frequencies (the counting variant, $O(n + \text{max citation})$). The sort-then-scan is the simplest correct shape; the counting version is the "no sort needed" optimization when citations are bounded (≤ 1000 here, so `O(n + 1000)` counting beats `O(n log n)`).
30+
31+
## Approach 1 — Count frequencies (O(n + maxC))
32+
33+
`count[c]` = papers with exactly c citations; walk from max down accumulating papers ≥ h: $O(n + \text{maxC})$, no sort. The "values are bounded" optimization worth mentioning.
34+
35+
## Approach 2 — Sort descending + scan (the repo's version, optimal)
36+
37+
```kotlin
38+
class HIndex {
39+
/**
40+
* @param citations citations[i] = citation count of paper i
41+
* @return the h-index
42+
*/
43+
fun hIndex(citations: IntArray): Int {
44+
// Step 1: Sort the citations in descending order
45+
citations.sortDescending()
46+
47+
// Step 2: Find the h-index
48+
for (i in citations.indices) {
49+
// The current index represents the number of papers.
50+
// Check if the current citation count is >= index + 1.
51+
if (citations[i] < i + 1) {
52+
return i // papers 0..i-1 met the bar
53+
}
54+
}
55+
return citations.size // every paper met the bar
56+
}
57+
}
58+
```
59+
60+
```java
61+
import java.util.*;
62+
63+
public class HIndex {
64+
/**
65+
* @param citations citations[i] = citation count of paper i
66+
* @return the h-index
67+
*/
68+
public int hIndex(int[] citations) {
69+
Integer[] sorted = Arrays.stream(citations).boxed()
70+
.sorted(Collections.reverseOrder()).toArray(Integer[]::new); // descending
71+
72+
for (int i = 0; i < sorted.length; i++) {
73+
if (sorted[i] < i + 1) return i; // papers 0..i-1 met the bar
74+
}
75+
return sorted.length; // every paper met the bar
76+
}
77+
}
78+
```
79+
80+
```cpp
81+
#include <algorithm>
82+
#include <vector>
83+
84+
class HIndex {
85+
public:
86+
/**
87+
* @param citations citations[i] = citation count of paper i
88+
* @return the h-index
89+
*/
90+
int hIndex(std::vector<int>& citations) {
91+
std::sort(citations.begin(), citations.end(), std::greater<int>()); // descending
92+
93+
for (int i = 0; i < (int)citations.size(); i++) {
94+
if (citations[i] < i + 1) return i; // papers 0..i-1 met the bar
95+
}
96+
return citations.size(); // every paper met the bar
97+
}
98+
};
99+
```
100+
101+
```python
102+
def h_index(citations: list[int]) -> int:
103+
"""
104+
@param citations: citations[i] = citation count of paper i
105+
@return: the h-index
106+
"""
107+
citations.sort(reverse=True) # descending
108+
109+
for i, c in enumerate(citations):
110+
if c < i + 1:
111+
return i # papers 0..i-1 met the bar
112+
return len(citations) # every paper met the bar
113+
```
114+
115+
```rust
116+
impl Solution {
117+
/// @param citations citations[i] = citation count of paper i
118+
/// @return the h-index
119+
pub fn h_index(citations: Vec<i32>) -> i32 {
120+
let mut citations = citations;
121+
citations.sort_unstable_by(|a, b| b.cmp(a)); // descending
122+
123+
for (i, &c) in citations.iter().enumerate() {
124+
if c < (i + 1) as i32 {
125+
return i as i32; // papers 0..i-1 met the bar
126+
}
127+
}
128+
citations.len() as i32 // every paper met the bar
129+
}
130+
}
131+
```
132+
133+
## Dry run
134+
135+
**Input:** `citations = [3,0,6,1,5]`.
136+
137+
```
138+
sorted descending: [6,5,3,1,0]
139+
i=0: 6 >= 1 ok. i=1: 5 >= 2 ok. i=2: 3 >= 3 ok. i=3: 1 >= 4? NO -> return 3 ✓
140+
```
141+
142+
Check the definition against the answer: h=3 means "≥3 papers with ≥3 citations" — papers with citations 6,5,3 (three of them) ✓. And h=4 fails: only 3 papers have ≥4 citations. The first failed check (`1 < 4`) is exactly where the definition stops holding.
143+
144+
## Complexity
145+
146+
**Time.** Sort dominates:
147+
148+
$$
149+
T(n) = O(n \log n)
150+
$$
151+
152+
**Space.** In-place sort:
153+
154+
$$
155+
S(n) = O(1)
156+
$$
157+
158+
## Variants & follow-ups
159+
160+
- **Counting version** — with citations ≤ 1000, count frequencies and walk backward accumulating: $O(n + \text{maxC})$, no sort. Mention when values are bounded.
161+
- **H-Index II** — the *sorted* input version: binary search for the boundary in $O(\log n)$.
162+
- **Interview follow-up:** "Why is the boundary check `citations[i] < i + 1` the whole problem?" After sorting descending, the condition "the first i papers have ≥ i citations" is checked at exactly one position — paper i is the first one *failing* the bar, so the count of passing papers is i. The sort converts a counting question into a boundary scan.
Lines changed: 25 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,25 @@
1+
# Chapter 14 — Sorting & QuickSelect
2+
3+
> **Source:** `src/main/kotlin/sorting/` and `src/main/kotlin/quicksort/`
4+
>
5+
> **Master idea:** sorting is *preprocessing that buys structure* — after a sort, adjacency means "next in order", and every comparison-based algorithm's cost is set by it. This chapter pairs the classic **merge sort** with **quickselect** (the "sort only enough" answer) and the **custom-comparator** problems where "sorted" is redefined.
6+
>
7+
> **Prerequisites:** recursion, arrays, and the heaps from [Chapter 7](../ch07-heaps/index.md) — the top-k problems have both a heap answer and a quickselect answer.
8+
9+
## Problems at a glance (this chapter's core set)
10+
11+
| # | Problem | Pattern | Complexity | Page |
12+
|---|---------|---------|------------|------|
13+
| 14.1 | Merge Sort | divide + merge | $O(n \log n)$ | [→](merge-sort.md) |
14+
| 14.2 | Kth Largest Element | randomized quickselect | $O(n)$ avg | [→](kth-largest-element.md) |
15+
| 14.3 | K Closest Points To Origin | quickselect on distance | $O(n)$ avg | [→](k-closest-points-to-origin.md) |
16+
| 14.4 | Largest Number | custom comparator | $O(n \log n)$ | [→](largest-number.md) |
17+
| 14.5 | H-Index | sort + scan | $O(n \log n)$ | [→](h-index.md) |
18+
| 14.6 | Russian Doll Envelopes | sort + LIS | $O(n \log n)$ | [→](russian-doll-envelopes.md) |
19+
| 14.7 | Top K Frequent (QuickSelect) | quickselect on frequency | $O(n)$ avg | [→](top-k-frequent-elements-quickselect.md) |
20+
21+
## The rest of the sorting/ and quicksort/ directories
22+
23+
`sorting/` also holds `EmployeeFreeTime.kt` and `RankTeamsByVote.kt`; `quicksort/` adds `DualPivotQuickSelect.kt` and `GenericRanrmoizedQuickSelect.kt` (generalized quickselect). The quickselect problems cross-reference the heap versions in [Chapter 7](../ch07-heaps/index.md) — the two answers to the same "top k" question, contrasted.
24+
25+
New pages are appended to the table above as they're written.

0 commit comments

Comments
 (0)