SweetLine uses JSON to define syntax highlighting rules. Each JSON file describes the highlighting rules for a programming language. The engine compiles these rules at runtime for text analysis.
- Basic Structure
- Top-Level Fields
- variables - Variable Definitions
- fragments - Rule Fragment Reuse
- states - State Machine Definition
- Pattern Matching Rules
- Style System
- State Transitions
- importSyntax - Syntax Import
- SubStates
- onLineEndState - Line End State
- scopeRules - Scope Rules
- bracketRules - Bracket Pair Rules
- Inline Style Mode
- Complete Example
- Best Practices
A complete syntax rule JSON file has the following structure:
{
"name": "languageName",
"fileNames": ["SpecialFile"],
"fileSuffixes": [".ext1", ".ext2"],
"fileNamePatterns": ["(?:generated|templated)\\.ext1"],
"variables": { ... },
"fragments": { ... },
"styles": [ ... ],
"states": {
"default": [ ... ],
"stateName": [ ... ]
},
"scopeRules": {
"skips": [ ... ],
"rules": [ ... ]
},
"bracketRules": {
"pairs": [ ... ]
}
}| Field | Type | Required | Description |
|---|---|---|---|
name |
string | Yes | Syntax rule name, used for getSyntaxRuleByName() lookup |
fileName |
string | No | Exact basename match, shorthand for a single fileNames entry |
fileNames |
string[] | No | Exact basenames used for getSyntaxRuleByFileName() matching |
fileSuffix |
string | No | Basename suffix match, shorthand for a single fileSuffixes entry |
fileSuffixes |
string[] | No | Basename suffix matches such as .java or .gradle.kts |
fileNamePattern |
string | No | Full-basename regex match, shorthand for a single fileNamePatterns entry |
fileNamePatterns |
string[] | No | Full-basename regex matches used only when exact names and suffixes are not enough |
variables |
object | No | Reusable regex pattern variable definitions |
fragments |
object | No | Reusable rule arrays that can be referenced by include / includes |
styles |
array | No | Inline style definitions (only for inline_style mode) |
states |
object | Yes | State machine definitions containing all states and their matching rules |
scopeRules |
object | No | Scope analysis rules used by indent guides |
bracketRules |
object | No | Bracket pair rules used by rainbow bracket and bracket matching analysis |
At least one of fileName, fileNames, fileSuffix, fileSuffixes, fileNamePattern, or fileNamePatterns is required. Routing is basename-based and case-sensitive. The engine resolves exact names first, then suffixes, then file-name patterns.
variables defines reusable regex fragments that can be referenced in pattern using ${variableName}. Variables can reference other variables (nested expansion is supported).
{
"variables": {
"identifier": "[a-zA-Z_]\\w*",
"whiteSpace": "[ \\t\\f]",
"any": "[\\S\\s]",
"primitiveType": "void|boolean|byte|char|short|int|long|float|double"
}
}Usage in patterns:
{
"pattern": "\\b(${primitiveType})\\b${whiteSpace}+(${identifier})",
"styles": [1, "keyword", 2, "variable"]
}After expansion, this is equivalent to:
\b(void|boolean|byte|char|short|int|long|float|double)\b[ \t\f]+([a-zA-Z_]\w*)
Notes:
- Variable names are case-sensitive
- Variables support nested references (e.g.,
"identifier": "${identifierStart}${identifierPart}") - SweetLine uses Oniguruma regex syntax, supporting
\p{Han}(Unicode properties), lookaheads/lookbehinds, and other advanced features - Backslashes must be double-escaped in JSON: regex
\bis written as\\bin JSON
fragments lets you define reusable rule arrays and expand them into states (or other fragments)
via include / includes. This helps reduce duplicated JSON and keep rule priority predictable.
{
"fragments": {
"commonLiterals": [
{ "pattern": "\"(?:[^\"\\\\]|\\\\.)*\"", "style": "string" },
{ "pattern": "\\b[0-9]+\\b", "style": "number" }
],
"commonComments": [
{ "pattern": "/\\*", "style": "comment", "state": "longComment" },
{ "pattern": "//${any}*", "style": "comment" }
]
},
"states": {
"default": [
{ "include": "commonLiterals" },
{ "includes": ["commonComments"] }
]
}
}Rules:
includetakes one fragment name (string)includestakes multiple fragment names (string[]) and expands in array orderinclude/includesentries are expanded in-place, so normal rule priority is preservedinclude/includesentries cannot contain other fields- Circular fragment references are rejected during compilation
states is the core of syntax rules, defining a Finite State Machine (FSM). Each state contains an ordered list of matching rules. The engine tries to match them in order and uses the first successful match.
default is the initial state — all text parsing starts from this state. Other states are entered via the state field.
{
"states": {
"default": [
{ "pattern": "...", "style": "...", "state": "stringState" }
],
"stringState": [
{ "pattern": "...", "style": "...", "state": "default" }
]
}
}Each state is an array [] of matching rule objects. The engine tries to match each rule at the current position in array order. On a successful match, it applies the style and advances past the matched text.
Each matching rule object contains the following fields:
| Field | Type | Required | Description |
|---|---|---|---|
pattern |
string | Yes* | Oniguruma regex, supports ${variableName} substitution |
style |
string | No | Style name for the entire match (mutually exclusive with styles) |
styles |
array | No | Capture group style mapping (mutually exclusive with style) |
state |
string | No | Target state to transition to after a successful match |
subStates |
array | No | Delegate specific capture group content to a sub-state for processing |
importSyntax |
string | No* | Import another syntax rule into the current state |
#ifdef |
string | No | Used with importSyntax; import only when the macro is defined |
onLineEndState |
string | No* | Auto-transition state at end of line (special rule, no pattern) |
include |
string | No* | Expand one named fragment in-place (special rule, no pattern) |
includes |
string[] | No* | Expand multiple named fragments in-place (special rule, no pattern) |
*Note: Regular matching rules must havepattern;onLineEndState,importSyntax,include, andincludesare special rules withoutpattern.
{
"pattern": "\\b(if|else|while|for|return)\\b",
"styles": [1, "keyword"]
}SweetLine uses Oniguruma as its regex engine, supporting:
- Standard regex:
\b,\w,\d,\s,.,*,+,?,|,(),[], etc. - Unicode properties:
\p{Han}(CJK characters),\p{L}(letters), etc. - Lookaheads/lookbehinds:
(?=...)lookahead,(?<=...)lookbehind,(?!...)negative lookahead - Non-greedy quantifiers:
*?,+?,?? - Character classes:
[^()]*matches any character except parentheses
Oniguruma vs PCRE differences:
- Supports variable-length lookbehind (PCRE does not)
- Some syntax details differ — refer to Oniguruma documentation
Multiple patterns within the same state are compiled into a single combined regex (joined with |). Matching follows Oniguruma's leftmost-first principle. When multiple patterns can match at the same position, rules listed earlier have higher priority.
Therefore, more specific rules should be placed before more general ones. For example:
[
{ "pattern": "\\b(class)\\b${whiteSpace}+(${identifier})", "styles": [1, "keyword", 2, "class"] },
{ "pattern": "\\b(class)\\b", "styles": [1, "keyword"] }
]Applies a single style to the entire matched text:
{
"pattern": "\"(?:[^\"\\\\]|\\\\.)*\"",
"style": "string"
}Maps different capture groups to different styles. Format: [groupNumber, "styleName", groupNumber, "styleName", ...]:
{
"pattern": "\\b(class)\\b${whiteSpace}+(${identifier})",
"styles": [1, "keyword", 2, "class"]
}- Group numbers start from 1 (corresponding to the first
()capture group in the regex) - Matched text not covered by any capture group will not be highlighted
- Gap text between capture groups will not be highlighted
SweetLine does not have a predefined style list — style names are entirely user-defined. Common conventions include:
| Name | Typical Usage | Name | Typical Usage |
|---|---|---|---|
keyword |
Language keywords | string |
String literals |
number |
Numeric literals | comment |
Comments |
class |
Class/type names | method |
Method/function names |
variable |
Variable names | property |
Property names |
punctuation |
Punctuation/operators | annotation |
Annotations/decorators |
builtin |
Built-in constants | preprocessor |
Preprocessor directives |
When using style ID mode, register name-to-ID mappings via engine.registerStyleName("keyword", 1).
Automatically transitions to the specified state after a successful match. This is the core mechanism for handling cross-line syntax structures.
Typical Scenario: String State
{
"states": {
"default": [
{
"pattern": "\"",
"style": "string",
"state": "doubleString"
}
],
"doubleString": [
{
"pattern": "\\\\.",
"style": "string"
},
{
"pattern": "\"",
"style": "string",
"state": "default"
},
{
"pattern": "${any}",
"style": "string"
}
]
}
}Workflow:
- In the
defaultstate, when"is encountered, it is styled asstringand the state transitions todoubleString - In
doubleString,\\.matches escape characters, keeping the string style - When the next
"is encountered, it is styled as string and transitions back todefault ${any}matches any other character, maintaining the string style (including across lines)
SweetLine supports zero-width matches (matches with length 0), useful for "look without consuming" state transitions:
{
"pattern": "^(?=[^ \\t])",
"style": "string",
"state": "default"
}This rule matches the position at the start of a line where the next character is not a space/tab (zero-width). It only performs a state transition without consuming any characters. The engine has a built-in infinite loop prevention mechanism: at most one zero-width match is allowed at the same position.
importSyntax imports another syntax rule into the current state. During import, token rules from the imported syntax's default state are merged, which is useful for host-language plus embedded-language scenarios (e.g., Markdown code blocks).
{
"importSyntax": "python"
}{
"importSyntax": "java",
"#ifdef": "ANDROID"
}The import is applied only when the macro is defined on HighlightEngine (for example, defineMacro("ANDROID")); otherwise the import rule is skipped.
Notes:
importSyntaxis a special rule and does not needpattern/style/styles#ifdefis only effective onimportSyntaxrules
subStates allows delegating the content of specific capture groups to another state for processing, instead of directly assigning a style. This is very useful for handling nested structures such as generic parameters.
"subStates": [groupNumber, "stateName", groupNumber, "stateName", ...]{
"pattern": "(${identifier})${whiteSpace}*(<)([^()]*)(>)${whiteSpace}+(${identifier})",
"styles": [1, "class", 2, "punctuation", 4, "punctuation", 5, "variable"],
"subStates": [3, "genericType"]
}For List<String> names:
- Group 1
List→classstyle - Group 2
<→punctuationstyle - Group 3
String→ delegated togenericTypestate (not styled instyles) - Group 4
>→punctuationstyle - Group 5
names→variablestyle
"genericType": [
{
"pattern": "(${identifier})${whiteSpace}*(<)([^()]*)(>)",
"styles": [1, "class", 2, "punctuation", 4, "punctuation"],
"subStates": [3, "genericType"]
},
{
"pattern": "${identifier}",
"style": "class"
},
{
"pattern": "[,\\[\\]?]",
"style": "punctuation"
}
]subStates can recursively reference its own state, enabling nested generic processing (e.g., Map<String, List<Integer>>).
Note: Capture groups referenced by subStates should not appear in styles, otherwise styles will override the sub-state processing results.
onLineEndState is a special rule that controls state transitions at the end of a line. When a line finishes analysis, if the current state contains an onLineEndState, the next line will begin analysis from the specified state.
{
"onLineEndState": "targetStateName"
}1. State persists to the next line (state preservation):
"classHeader": [
{ "pattern": "${identifier}", "style": "class" },
{ "pattern": "[\\{;]", "style": "punctuation", "state": "default" },
{ "onLineEndState": "classHeader" }
]When a class declaration spans multiple lines, onLineEndState ensures the next line remains in the classHeader state.
2. Line-end state fallback:
"methodParams": [
{ "pattern": "\\)", "style": "punctuation", "state": "default" },
{ "onLineEndState": "default" }
]If ) is not encountered on the current line, the next line automatically returns to the default state.
Note: onLineEndState should be placed as the last element in the state's rule array.
scopeRules defines lexical skip regions and scope markers for indent guides. It scans raw text directly, so it does not require highlight analysis.
{
"scopeRules": {
"skips": [
{ "kind": "lineComment", "start": "//" },
{ "kind": "blockComment", "start": "/*", "end": "*/" },
{ "kind": "string", "start": "\"", "end": "\"", "escape": "\\" }
],
"rules": [
{ "kind": "delimiter", "start": "{", "end": "}", "branches": ["case"] }
]
}
}| Field | Type | Description |
|---|---|---|
kind |
string | lineComment, blockComment, or string |
start |
string | Text that starts the skipped region |
end |
string | Text that ends the skipped region; required except for lineComment, and defaults to start for string |
escape |
string | Optional escape sequence for string-like skips |
multiLine |
boolean | Whether the skip can continue across lines; blockComment always behaves as multiline |
| Field | Type | Description |
|---|---|---|
kind |
string | delimiter, word, or indentStart |
start |
string | Scope start marker |
end |
string | Scope end marker; required for delimiter and word, omitted for indentStart |
branches |
string[] | Optional branch keywords within the block, such as case in switch |
delimiter matches literal marker text. word also requires word boundaries around the marker. indentStart starts a scope at a marker such as : and closes it when indentation falls back.
{
"scopeRules": {
"skips": [
{ "kind": "lineComment", "start": "#" },
{ "kind": "string", "start": "\"", "end": "\"", "escape": "\\" },
{ "kind": "string", "start": "'", "end": "'", "escape": "\\" }
],
"rules": [
{ "kind": "indentStart", "start": ":" }
]
}
}This is a common Python pattern: after :, the following indented block is treated as one scope.
Notes:
start/end/branchesare literal text markers, not regex patterns- Put comment and string delimiters in
skipsso scope markers inside them are ignored
bracketRules defines literal bracket pairs for rainbow bracket rendering and partner lookup. It scans raw text directly and does not depend on highlight spans or style IDs.
{
"bracketRules": {
"inheritScopeSkips": true,
"skips": [
{ "kind": "string", "start": "\"", "end": "\"", "escape": "\\" }
],
"pairs": [
{ "start": "(", "end": ")" },
{ "start": "[", "end": "]" },
{ "start": "{", "end": "}" }
]
}
}| Field | Type | Required | Description |
|---|---|---|---|
pairs |
object[] | Yes | Literal bracket pairs. The array must not be empty. |
inheritScopeSkips |
boolean | No | Whether to reuse scopeRules.skips; defaults to true. |
skips |
object[] | No | Additional skip rules using the same schema as scopeRules.skips. |
Each pair requires non-empty start and end strings. start and end must be different. Longer bracket or skip markers are matched first, so multi-character markers can coexist with shorter ones.
If a syntax has no scopeRules, define bracketRules.skips directly so brackets inside strings and comments are ignored. If a syntax already has complete scopeRules.skips, the common case is to omit skips and rely on the default inheritScopeSkips: true.
In inline_style mode, style definitions are written directly in the JSON. Highlighting results include colors and font attributes directly, without the need for external style registration.
{
"styles": [
{
"name": "keyword",
"foreground": "#FF569CD6",
"background": "#00000000",
"tags": ["bold"]
},
{
"name": "string",
"foreground": "#FFBD63C5"
},
{
"name": "comment",
"foreground": "#FF60AE6F",
"tags": ["italic"]
}
]
}| Field | Type | Description |
|---|---|---|
name |
string | Style name, corresponding to the name referenced in patterns |
foreground |
string | Foreground color in #AARRGGBB (ARGB) format |
background |
string | Background color in #AARRGGBB (ARGB) format, optional |
tags |
string[] | Font attribute tags, optional. Supports "bold", "italic", "strikethrough" |
Note: When using inline_style mode, set inline_style = true when creating the HighlightEngine.
Here is a complete simplified Java syntax rule example:
{
"name": "java",
"fileSuffixes": [".java"],
"variables": {
"identifier": "[a-zA-Z_$][\\w$]*",
"whiteSpace": "[ \\t\\f]",
"any": "[\\S\\s]",
"primitiveType": "void|boolean|byte|char|short|int|long|float|double"
},
"states": {
"default": [
{
"pattern": "\\b(class|interface|enum)\\b${whiteSpace}+(${identifier})",
"styles": [1, "keyword", 2, "class"],
"state": "classHeader"
},
{
"pattern": "\\b(new)\\b${whiteSpace}+(${identifier})${whiteSpace}*(<)([^()]*)(>)${whiteSpace}*(\\()",
"styles": [1, "keyword", 2, "class", 3, "punctuation", 5, "punctuation", 6, "punctuation"],
"subStates": [4, "genericType"]
},
{
"pattern": "\\b(new)\\b${whiteSpace}+(${identifier})${whiteSpace}*(\\()",
"styles": [1, "keyword", 2, "class", 3, "punctuation"]
},
{
"pattern": "\\b(public|private|protected|static|final|abstract|return|if|else|for|while)\\b",
"styles": [1, "keyword"]
},
{
"pattern": "\\b(true|false|null)\\b",
"styles": [1, "builtin"]
},
{
"pattern": "@${identifier}",
"style": "annotation"
},
{
"pattern": "\\b(${primitiveType})\\b${whiteSpace}+(${identifier})${whiteSpace}*(\\()",
"styles": [1, "keyword", 2, "method", 3, "punctuation"]
},
{
"pattern": "\\b(${primitiveType})\\b${whiteSpace}+(${identifier})${whiteSpace}*([;=,)])",
"styles": [1, "keyword", 2, "variable", 3, "punctuation"]
},
{
"pattern": "\\b(${identifier})${whiteSpace}*(<)([^()]*)(>)${whiteSpace}+(${identifier})${whiteSpace}*([;=,)])",
"styles": [1, "class", 2, "punctuation", 4, "punctuation", 5, "variable", 6, "punctuation"],
"subStates": [3, "genericType"]
},
{
"pattern": "\\b(${identifier})${whiteSpace}+(${identifier})${whiteSpace}*([;=,)])",
"styles": [1, "class", 2, "variable", 3, "punctuation"]
},
{
"pattern": "(${identifier})${whiteSpace}*(\\()",
"styles": [1, "method", 2, "punctuation"]
},
{
"pattern": "\"(?:[^\"\\\\]|\\\\.)*\"",
"style": "string"
},
{
"pattern": "\\b[0-9][0-9_]*\\.?[0-9_]*(?:[eE][+-]?[0-9]+)?[fFdDlL]?\\b",
"style": "number"
},
{
"pattern": "/\\*",
"style": "comment",
"state": "blockComment"
},
{
"pattern": "//${any}*",
"style": "comment"
},
{
"pattern": "[.()\\[\\]{}+\\-*/<>=!&|;:,?~^%@]",
"style": "punctuation"
}
],
"classHeader": [
{
"pattern": "(<)([^()]*)(>)",
"styles": [1, "punctuation", 3, "punctuation"],
"subStates": [2, "genericType"]
},
{
"pattern": "\\b(extends|implements)\\b",
"styles": [1, "keyword"]
},
{
"pattern": "${identifier}",
"style": "class"
},
{
"pattern": "[,:]",
"style": "punctuation"
},
{
"pattern": "\\{",
"style": "punctuation",
"state": "default"
},
{ "onLineEndState": "classHeader" }
],
"genericType": [
{
"pattern": "\\b(extends|super)\\b",
"styles": [1, "keyword"]
},
{
"pattern": "(${identifier})${whiteSpace}*(<)([^()]*)(>)",
"styles": [1, "class", 2, "punctuation", 4, "punctuation"],
"subStates": [3, "genericType"]
},
{
"pattern": "${identifier}",
"style": "class"
},
{
"pattern": "[,\\[\\]?]",
"style": "punctuation"
}
],
"blockComment": [
{
"pattern": "\\*/",
"style": "comment",
"state": "default"
},
{
"pattern": "${any}",
"style": "comment"
}
]
},
"scopeRules": {
"rules": [
{ "kind": "delimiter", "start": "{", "end": "}" }
]
},
"bracketRules": {
"pairs": [
{ "start": "(", "end": ")" },
{ "start": "[", "end": "]" },
{ "start": "{", "end": "}" }
]
}
}Rule ordering affects matching priority. The recommended order for default state rules:
- Class/struct declarations (triggers classHeader state transition)
newexpressions (constructor calls)- import/using statements
- Variable declaration keywords (var, let, etc.)
- Preprocessor directives
- Built-in type + method/property/variable declarations
- General keywords
- Built-in constants (true/false/null)
- Annotations/decorators
- Generic type + method/property/variable declarations
- Simple type + method/property/variable declarations
- Method calls (fallback pattern)
- String literals (multi-line strings requiring state transitions)
- Numeric literals
- Comments (block comments require state transitions)
- Operators and punctuation (fallback pattern)
When handling generic parameters, use [^()]* instead of .* to prevent regex backtracking across too much content:
// Recommended
"(${identifier})${whiteSpace}*(<)([^()]*)(>)"
// Avoid
"(${identifier})${whiteSpace}*(<)(.*)(>)"When the same line contains multiple independent <> pairs (e.g., class inheritance declarations), use non-greedy matching (.*?):
"(<)(.*?)(>)"- Multi-line structures (class declarations, function parameter lists) use
onLineEndStateto preserve state - States that only need to persist until line end don't need
onLineEndState(defaults back to default) - Place
onLineEndStateat the end of the state's rule array
Extract frequently used regex fragments as variables:
{
"variables": {
"identifier": "[a-zA-Z_]\\w*",
"whiteSpace": "[ \\t\\f]",
"builtinType": "int|float|double|string|bool|void"
}
}- Every state should have clear entry conditions and exit conditions
- Multi-line syntax structures (strings, comments) must use state transitions
- The last pattern in a state should be a fallback match (e.g.,
${any}) to prevent the engine from getting stuck on certain characters - Avoid creating too many states — most languages can be handled with 5-10 states