Skip to content
negativeview edited this page Jun 14, 2011 · 2 revisions

Tokenizer

Should be able to split a string into bbcode tokens and normal text. Certain bits of whitespace should be treated as special tokens so that we can use whitespace-changes as a measure of impact.

Normalizer

Normalize the token structure so that, for instance, multiple whitespace elements where they won't matter, are treated as a single one.

Longest Substring

Find the longest substring, as the Wikipedia article about diffing algorithms talks about.

Adds and Deletes

Compute the adds and deletes.

Final Structure

We need all nodes in order, with an add, delete, or same flag. These are nodes, not lines, so that we can have sub-line changes.

Clone this wiki locally