|
| 1 | +# 9.2 Group Anagrams |
| 2 | + |
| 3 | +> **Source:** [`src/main/kotlin/string/GroupAnagrams.kt`](https://github.com/arpanpathak/AdvancedAlgorithmPatterns/blob/main/src/main/kotlin/string/GroupAnagrams.kt) |
| 4 | +> **Pattern:** frequency-vector key · **Core page** |
| 5 | +
|
| 6 | +## The Problem |
| 7 | + |
| 8 | +Given an array of strings, group the **anagrams** together. The answer may be in any order. |
| 9 | + |
| 10 | +- Constraints: $1 \le n \le 10^4$; $0 \le L \le 100$ (word length); lowercase letters only. |
| 11 | + |
| 12 | +## Examples |
| 13 | + |
| 14 | +``` |
| 15 | +Input: strs = ["eat","tea","tan","ate","nat","bat"] |
| 16 | +Output: [["bat"],["nat","tan"],["ate","eat","tea"]] |
| 17 | +
|
| 18 | +Input: strs = [""] -> Output: [[""]] |
| 19 | +Input: strs = ["a"] -> Output: [["a"]] |
| 20 | +``` |
| 21 | + |
| 22 | +## Intuition — "same multiset" needs a canonical key |
| 23 | + |
| 24 | +Two words are anagrams iff they share a *canonical form* — something identical for exactly the anagram class. Two classic choices: |
| 25 | + |
| 26 | +1. **Sorted word** — `"eat"` and `"tea"` both canonicalize to `"aet"`. Simple, but sorting each word costs $O(L \log L)$. |
| 27 | +2. **Frequency vector** — the `int[26]` count from [9.1](valid-anagram.md), used as a *key*: `"eat"` and `"tea"` both produce `{a:1, e:1, t:1}`. Counting is $O(L)$ per word — faster than sorting — and the repo uses exactly this. |
| 28 | + |
| 29 | +The data structure is a **map from canonical form to word list**: for each word, compute its frequency vector, look up (or create) the bucket, append. What the [primer](pattern-primer.md) calls the *encode-then-equate* move: an equivalence question becomes a *hash lookup* question. |
| 30 | + |
| 31 | +**Why a `List<Int>` key and not a `String`?** In Kotlin the repo builds `count.toList()` — a 26-element vector — and uses it directly as the map key (lists have structural equality). Java/C++/Python below use the stringified counts (`"1#0#..."`) or the sorted word; any canonical form that is equal exactly for anagram classes works. |
| 32 | + |
| 33 | +**The early trap:** using the *set of characters* instead of the counts — `"aab"` and `"abb"` have the same character set `{a,b}` but are not anagrams. The frequency vector distinguishes them; a set does not. Say this distinction unprompted. |
| 34 | + |
| 35 | +## Approach 1 — Sort each word, group by sorted form |
| 36 | + |
| 37 | +`map[sorted(word)] += word`: $O(n \cdot L \log L)$ total. Simpler to read; the frequency-vector version trades the log for a constant. |
| 38 | + |
| 39 | +## Approach 2 — Frequency-vector keys (the repo's version, optimal) |
| 40 | + |
| 41 | +```kotlin |
| 42 | +class GroupAnagrams { |
| 43 | + /** |
| 44 | + * @param strs array of words |
| 45 | + * @return words grouped by anagram class |
| 46 | + */ |
| 47 | + fun groupAnagrams(strs: Array<String>): List<List<String>> { |
| 48 | + val anagramMap = mutableMapOf<List<Int>, MutableList<String>>() // frequency vector -> bucket |
| 49 | + |
| 50 | + for (word in strs) { |
| 51 | + val count = IntArray(26) |
| 52 | + word.forEach { count[it - 'a']++ } // canonical form: the counts |
| 53 | + anagramMap.getOrPut(count.toList()) { mutableListOf() }.add(word) |
| 54 | + } |
| 55 | + return anagramMap.values.toList() |
| 56 | + } |
| 57 | +} |
| 58 | +``` |
| 59 | + |
| 60 | +```java |
| 61 | +import java.util.*; |
| 62 | + |
| 63 | +public class GroupAnagrams { |
| 64 | + /** |
| 65 | + * @param strs array of words |
| 66 | + * @return words grouped by anagram class |
| 67 | + */ |
| 68 | + public List<List<String>> groupAnagrams(String[] strs) { |
| 69 | + Map<String, List<String>> map = new HashMap<>(); |
| 70 | + |
| 71 | + for (String word : strs) { |
| 72 | + int[] count = new int[26]; |
| 73 | + for (char c : word.toCharArray()) count[c - 'a']++; |
| 74 | + |
| 75 | + StringBuilder key = new StringBuilder(); // canonical form: "1#2#0#..." |
| 76 | + for (int c : count) key.append(c).append('#'); |
| 77 | + map.computeIfAbsent(key.toString(), k -> new ArrayList<>()).add(word); |
| 78 | + } |
| 79 | + return new ArrayList<>(map.values()); |
| 80 | + } |
| 81 | +} |
| 82 | +``` |
| 83 | + |
| 84 | +```cpp |
| 85 | +#include <string> |
| 86 | +#include <unordered_map> |
| 87 | +#include <vector> |
| 88 | + |
| 89 | +class GroupAnagrams { |
| 90 | +public: |
| 91 | + /** |
| 92 | + * @param strs array of words |
| 93 | + * @return words grouped by anagram class |
| 94 | + */ |
| 95 | + std::vector<std::vector<std::string>> groupAnagrams(std::vector<std::string>& strs) { |
| 96 | + std::unordered_map<std::string, std::vector<std::string>> map; |
| 97 | + |
| 98 | + for (auto& word : strs) { |
| 99 | + int count[26] = {0}; |
| 100 | + for (char c : word) count[c - 'a']++; |
| 101 | + |
| 102 | + std::string key; // canonical form: counts concatenated |
| 103 | + for (int c : count) key += std::to_string(c) + "#"; |
| 104 | + map[key].push_back(word); |
| 105 | + } |
| 106 | + |
| 107 | + std::vector<std::vector<std::string>> result; |
| 108 | + for (auto& [_, bucket] : map) result.push_back(bucket); |
| 109 | + return result; |
| 110 | + } |
| 111 | +}; |
| 112 | +``` |
| 113 | + |
| 114 | +```python |
| 115 | +def group_anagrams(strs: list[str]) -> list[list[str]]: |
| 116 | + """ |
| 117 | + @param strs: array of words |
| 118 | + @return: words grouped by anagram class |
| 119 | + """ |
| 120 | + buckets: dict[tuple[int, ...], list[str]] = {} |
| 121 | + for word in strs: |
| 122 | + count = [0] * 26 |
| 123 | + for c in word: |
| 124 | + count[ord(c) - ord('a')] += 1 |
| 125 | + key = tuple(count) # canonical form: the counts |
| 126 | + buckets.setdefault(key, []).append(word) |
| 127 | + return list(buckets.values()) |
| 128 | +``` |
| 129 | + |
| 130 | +```rust |
| 131 | +use std::collections::HashMap; |
| 132 | + |
| 133 | +impl Solution { |
| 134 | + /// @param strs array of words |
| 135 | + /// @return words grouped by anagram class |
| 136 | + pub fn group_anagrams(strs: Vec<String>) -> Vec<Vec<String>> { |
| 137 | + let mut buckets: HashMap<[i32; 26], Vec<String>> = HashMap::new(); |
| 138 | + |
| 139 | + for word in strs { |
| 140 | + let mut count = [0i32; 26]; |
| 141 | + for b in word.bytes() { |
| 142 | + count[(b - b'a') as usize] += 1; |
| 143 | + } |
| 144 | + buckets.entry(count).or_default().push(word); // canonical form: the counts |
| 145 | + } |
| 146 | + buckets.into_values().collect() |
| 147 | + } |
| 148 | +} |
| 149 | +``` |
| 150 | + |
| 151 | +## Dry run |
| 152 | + |
| 153 | +**Input:** `strs = ["eat","tea","tan","ate","nat","bat"]`. |
| 154 | + |
| 155 | +``` |
| 156 | +"eat": count = {a:1,e:1,t:1} -> key -> bucket["eat"] |
| 157 | +"tea": count = {a:1,e:1,t:1} -> same key -> bucket["eat","tea"] |
| 158 | +"tan": count = {a:1,n:1,t:1} -> new key -> bucket["tan"] |
| 159 | +"ate": count = {a:1,e:1,t:1} -> same as eat -> bucket["eat","tea","ate"] |
| 160 | +"nat": count = {a:1,n:1,t:1} -> same as tan -> bucket["tan","nat"] |
| 161 | +"bat": count = {a:1,b:1,t:1} -> new key -> bucket["bat"] |
| 162 | +
|
| 163 | +buckets: {"a1e1t1": ["eat","tea","ate"], "a1n1t1": ["tan","nat"], "a1b1t1": ["bat"]} |
| 164 | +Output: [["eat","tea","ate"],["tan","nat"],["bat"]] ✓ |
| 165 | +``` |
| 166 | + |
| 167 | +The key insight in action: `"eat"` and `"tea"` never needed to be *compared* — they both produced the identical 26-count vector, so the hash map did the grouping. And `"tan"` vs `"bat"`: both length-3 with `a`,`t`, but different third letters — different vectors, different buckets. |
| 168 | + |
| 169 | +## Complexity |
| 170 | + |
| 171 | +**Time.** Each word counted once ($O(L)$) and hashed: |
| 172 | + |
| 173 | +$$ |
| 174 | +T(n, L) = O(n \cdot L) |
| 175 | +$$ |
| 176 | + |
| 177 | +**Space.** The map holds every word: |
| 178 | + |
| 179 | +$$ |
| 180 | +S(n, L) = O(n \cdot L) |
| 181 | +$$ |
| 182 | + |
| 183 | +## Variants & follow-ups |
| 184 | + |
| 185 | +- **Valid Anagram** ([9.1](valid-anagram.md)) — the single-pair case; this page is the "many pairs at once" generalization. |
| 186 | +- **Count Words With A Given Prefix** (`src/main/kotlin/string/CountWordsWithAGivenPrefix.kt`) — grouping by a *prefix* instead of a frequency class: a trie (`src/main/kotlin/trie/`) is the structured version of the same bucket idea. |
| 187 | +- **Find Duplicate File In System / grouping by content** — any "group by canonical form" problem is this skeleton. |
| 188 | +- **Interview follow-up:** "Why is the frequency vector better than sorting each word?" Sorting is $O(L \log L)$ per word; counting is $O(L)$. Over $10^4$ words of length 100 that is a constant-factor (but real) difference — and counting also makes the "what exactly defines an anagram?" reasoning explicit. For very short words the sort is simpler; say both. |
0 commit comments