Skip to content

Latest commit

 

History

History
822 lines (667 loc) · 25.9 KB

File metadata and controls

822 lines (667 loc) · 25.9 KB

SweetLine Syntax Rule Configuration Guide

SweetLine uses JSON to define syntax highlighting rules. Each JSON file describes the highlighting rules for a programming language. The engine compiles these rules at runtime for text analysis.

Table of Contents


Basic Structure

A complete syntax rule JSON file has the following structure:

{
  "name": "languageName",
  "fileNames": ["SpecialFile"],
  "fileSuffixes": [".ext1", ".ext2"],
  "fileNamePatterns": ["(?:generated|templated)\\.ext1"],
  "variables": { ... },
  "fragments": { ... },
  "styles": [ ... ],
  "states": {
    "default": [ ... ],
    "stateName": [ ... ]
  },
  "scopeRules": {
    "skips": [ ... ],
    "rules": [ ... ]
  },
  "bracketRules": {
    "pairs": [ ... ]
  }
}

Top-Level Fields

Field Type Required Description
name string Yes Syntax rule name, used for getSyntaxRuleByName() lookup
fileName string No Exact basename match, shorthand for a single fileNames entry
fileNames string[] No Exact basenames used for getSyntaxRuleByFileName() matching
fileSuffix string No Basename suffix match, shorthand for a single fileSuffixes entry
fileSuffixes string[] No Basename suffix matches such as .java or .gradle.kts
fileNamePattern string No Full-basename regex match, shorthand for a single fileNamePatterns entry
fileNamePatterns string[] No Full-basename regex matches used only when exact names and suffixes are not enough
variables object No Reusable regex pattern variable definitions
fragments object No Reusable rule arrays that can be referenced by include / includes
styles array No Inline style definitions (only for inline_style mode)
states object Yes State machine definitions containing all states and their matching rules
scopeRules object No Scope analysis rules used by indent guides
bracketRules object No Bracket pair rules used by rainbow bracket and bracket matching analysis

At least one of fileName, fileNames, fileSuffix, fileSuffixes, fileNamePattern, or fileNamePatterns is required. Routing is basename-based and case-sensitive. The engine resolves exact names first, then suffixes, then file-name patterns.

variables - Variable Definitions

variables defines reusable regex fragments that can be referenced in pattern using ${variableName}. Variables can reference other variables (nested expansion is supported).

{
  "variables": {
    "identifier": "[a-zA-Z_]\\w*",
    "whiteSpace": "[ \\t\\f]",
    "any": "[\\S\\s]",
    "primitiveType": "void|boolean|byte|char|short|int|long|float|double"
  }
}

Usage in patterns:

{
  "pattern": "\\b(${primitiveType})\\b${whiteSpace}+(${identifier})",
  "styles": [1, "keyword", 2, "variable"]
}

After expansion, this is equivalent to:

\b(void|boolean|byte|char|short|int|long|float|double)\b[ \t\f]+([a-zA-Z_]\w*)

Notes:

  • Variable names are case-sensitive
  • Variables support nested references (e.g., "identifier": "${identifierStart}${identifierPart}")
  • SweetLine uses Oniguruma regex syntax, supporting \p{Han} (Unicode properties), lookaheads/lookbehinds, and other advanced features
  • Backslashes must be double-escaped in JSON: regex \b is written as \\b in JSON

fragments - Rule Fragment Reuse

fragments lets you define reusable rule arrays and expand them into states (or other fragments) via include / includes. This helps reduce duplicated JSON and keep rule priority predictable.

{
  "fragments": {
    "commonLiterals": [
      { "pattern": "\"(?:[^\"\\\\]|\\\\.)*\"", "style": "string" },
      { "pattern": "\\b[0-9]+\\b", "style": "number" }
    ],
    "commonComments": [
      { "pattern": "/\\*", "style": "comment", "state": "longComment" },
      { "pattern": "//${any}*", "style": "comment" }
    ]
  },
  "states": {
    "default": [
      { "include": "commonLiterals" },
      { "includes": ["commonComments"] }
    ]
  }
}

Rules:

  • include takes one fragment name (string)
  • includes takes multiple fragment names (string[]) and expands in array order
  • include / includes entries are expanded in-place, so normal rule priority is preserved
  • include / includes entries cannot contain other fields
  • Circular fragment references are rejected during compilation

states - State Machine Definition

states is the core of syntax rules, defining a Finite State Machine (FSM). Each state contains an ordered list of matching rules. The engine tries to match them in order and uses the first successful match.

The default State is Required

default is the initial state — all text parsing starts from this state. Other states are entered via the state field.

{
  "states": {
    "default": [
      { "pattern": "...", "style": "...", "state": "stringState" }
    ],
    "stringState": [
      { "pattern": "...", "style": "...", "state": "default" }
    ]
  }
}

Rules Within a State

Each state is an array [] of matching rule objects. The engine tries to match each rule at the current position in array order. On a successful match, it applies the style and advances past the matched text.


Pattern Matching Rules

Each matching rule object contains the following fields:

Field Type Required Description
pattern string Yes* Oniguruma regex, supports ${variableName} substitution
style string No Style name for the entire match (mutually exclusive with styles)
styles array No Capture group style mapping (mutually exclusive with style)
state string No Target state to transition to after a successful match
subStates array No Delegate specific capture group content to a sub-state for processing
importSyntax string No* Import another syntax rule into the current state
#ifdef string No Used with importSyntax; import only when the macro is defined
onLineEndState string No* Auto-transition state at end of line (special rule, no pattern)
include string No* Expand one named fragment in-place (special rule, no pattern)
includes string[] No* Expand multiple named fragments in-place (special rule, no pattern)

* Note: Regular matching rules must have pattern; onLineEndState, importSyntax, include, and includes are special rules without pattern.

Basic Pattern Example

{
  "pattern": "\\b(if|else|while|for|return)\\b",
  "styles": [1, "keyword"]
}

Regex Reference

SweetLine uses Oniguruma as its regex engine, supporting:

  • Standard regex: \b, \w, \d, \s, ., *, +, ?, |, (), [], etc.
  • Unicode properties: \p{Han} (CJK characters), \p{L} (letters), etc.
  • Lookaheads/lookbehinds: (?=...) lookahead, (?<=...) lookbehind, (?!...) negative lookahead
  • Non-greedy quantifiers: *?, +?, ??
  • Character classes: [^()]* matches any character except parentheses

Oniguruma vs PCRE differences:

  • Supports variable-length lookbehind (PCRE does not)
  • Some syntax details differ — refer to Oniguruma documentation

Matching Behavior

Multiple patterns within the same state are compiled into a single combined regex (joined with |). Matching follows Oniguruma's leftmost-first principle. When multiple patterns can match at the same position, rules listed earlier have higher priority.

Therefore, more specific rules should be placed before more general ones. For example:

[
  { "pattern": "\\b(class)\\b${whiteSpace}+(${identifier})", "styles": [1, "keyword", 2, "class"] },
  { "pattern": "\\b(class)\\b", "styles": [1, "keyword"] }
]

Style System

Option 1: style - Whole Match Style

Applies a single style to the entire matched text:

{
  "pattern": "\"(?:[^\"\\\\]|\\\\.)*\"",
  "style": "string"
}

Option 2: styles - Capture Group Style Mapping

Maps different capture groups to different styles. Format: [groupNumber, "styleName", groupNumber, "styleName", ...]:

{
  "pattern": "\\b(class)\\b${whiteSpace}+(${identifier})",
  "styles": [1, "keyword", 2, "class"]
}
  • Group numbers start from 1 (corresponding to the first () capture group in the regex)
  • Matched text not covered by any capture group will not be highlighted
  • Gap text between capture groups will not be highlighted

Style Names

SweetLine does not have a predefined style list — style names are entirely user-defined. Common conventions include:

Name Typical Usage Name Typical Usage
keyword Language keywords string String literals
number Numeric literals comment Comments
class Class/type names method Method/function names
variable Variable names property Property names
punctuation Punctuation/operators annotation Annotations/decorators
builtin Built-in constants preprocessor Preprocessor directives

When using style ID mode, register name-to-ID mappings via engine.registerStyleName("keyword", 1).


State Transitions

state - Transition After Match

Automatically transitions to the specified state after a successful match. This is the core mechanism for handling cross-line syntax structures.

Typical Scenario: String State

{
  "states": {
    "default": [
      {
        "pattern": "\"",
        "style": "string",
        "state": "doubleString"
      }
    ],
    "doubleString": [
      {
        "pattern": "\\\\.",
        "style": "string"
      },
      {
        "pattern": "\"",
        "style": "string",
        "state": "default"
      },
      {
        "pattern": "${any}",
        "style": "string"
      }
    ]
  }
}

Workflow:

  1. In the default state, when " is encountered, it is styled as string and the state transitions to doubleString
  2. In doubleString, \\. matches escape characters, keeping the string style
  3. When the next " is encountered, it is styled as string and transitions back to default
  4. ${any} matches any other character, maintaining the string style (including across lines)

Zero-Width Matches and State Transitions

SweetLine supports zero-width matches (matches with length 0), useful for "look without consuming" state transitions:

{
  "pattern": "^(?=[^ \\t])",
  "style": "string",
  "state": "default"
}

This rule matches the position at the start of a line where the next character is not a space/tab (zero-width). It only performs a state transition without consuming any characters. The engine has a built-in infinite loop prevention mechanism: at most one zero-width match is allowed at the same position.


importSyntax - Syntax Import

importSyntax imports another syntax rule into the current state. During import, token rules from the imported syntax's default state are merged, which is useful for host-language plus embedded-language scenarios (e.g., Markdown code blocks).

Basic Syntax

{
  "importSyntax": "python"
}

Conditional Import (#ifdef)

{
  "importSyntax": "java",
  "#ifdef": "ANDROID"
}

The import is applied only when the macro is defined on HighlightEngine (for example, defineMacro("ANDROID")); otherwise the import rule is skipped.

Notes:

  • importSyntax is a special rule and does not need pattern / style / styles
  • #ifdef is only effective on importSyntax rules

SubStates

subStates allows delegating the content of specific capture groups to another state for processing, instead of directly assigning a style. This is very useful for handling nested structures such as generic parameters.

Format

"subStates": [groupNumber, "stateName", groupNumber, "stateName", ...]

Example: Generic Type Handling

{
  "pattern": "(${identifier})${whiteSpace}*(<)([^()]*)(>)${whiteSpace}+(${identifier})",
  "styles": [1, "class", 2, "punctuation", 4, "punctuation", 5, "variable"],
  "subStates": [3, "genericType"]
}

For List<String> names:

  • Group 1 List → class style
  • Group 2 < → punctuation style
  • Group 3 String → delegated to genericType state (not styled in styles)
  • Group 4 > → punctuation style
  • Group 5 names → variable style

genericType State Definition

"genericType": [
  {
    "pattern": "(${identifier})${whiteSpace}*(<)([^()]*)(>)",
    "styles": [1, "class", 2, "punctuation", 4, "punctuation"],
    "subStates": [3, "genericType"]
  },
  {
    "pattern": "${identifier}",
    "style": "class"
  },
  {
    "pattern": "[,\\[\\]?]",
    "style": "punctuation"
  }
]

subStates can recursively reference its own state, enabling nested generic processing (e.g., Map<String, List<Integer>>).

Note: Capture groups referenced by subStates should not appear in styles, otherwise styles will override the sub-state processing results.


onLineEndState - Line End State

onLineEndState is a special rule that controls state transitions at the end of a line. When a line finishes analysis, if the current state contains an onLineEndState, the next line will begin analysis from the specified state.

Syntax

{
  "onLineEndState": "targetStateName"
}

Typical Use Cases

1. State persists to the next line (state preservation):

"classHeader": [
  { "pattern": "${identifier}", "style": "class" },
  { "pattern": "[\\{;]", "style": "punctuation", "state": "default" },
  { "onLineEndState": "classHeader" }
]

When a class declaration spans multiple lines, onLineEndState ensures the next line remains in the classHeader state.

2. Line-end state fallback:

"methodParams": [
  { "pattern": "\\)", "style": "punctuation", "state": "default" },
  { "onLineEndState": "default" }
]

If ) is not encountered on the current line, the next line automatically returns to the default state.

Note: onLineEndState should be placed as the last element in the state's rule array.


scopeRules - Scope Rules

scopeRules defines lexical skip regions and scope markers for indent guides. It scans raw text directly, so it does not require highlight analysis.

{
  "scopeRules": {
    "skips": [
      { "kind": "lineComment", "start": "//" },
      { "kind": "blockComment", "start": "/*", "end": "*/" },
      { "kind": "string", "start": "\"", "end": "\"", "escape": "\\" }
    ],
    "rules": [
      { "kind": "delimiter", "start": "{", "end": "}", "branches": ["case"] }
    ]
  }
}

Skip Rules

Field Type Description
kind string lineComment, blockComment, or string
start string Text that starts the skipped region
end string Text that ends the skipped region; required except for lineComment, and defaults to start for string
escape string Optional escape sequence for string-like skips
multiLine boolean Whether the skip can continue across lines; blockComment always behaves as multiline

Scope Rules

Field Type Description
kind string delimiter, word, or indentStart
start string Scope start marker
end string Scope end marker; required for delimiter and word, omitted for indentStart
branches string[] Optional branch keywords within the block, such as case in switch

delimiter matches literal marker text. word also requires word boundaries around the marker. indentStart starts a scope at a marker such as : and closes it when indentation falls back.

Indentation Start-Marker Mode

{
  "scopeRules": {
    "skips": [
      { "kind": "lineComment", "start": "#" },
      { "kind": "string", "start": "\"", "end": "\"", "escape": "\\" },
      { "kind": "string", "start": "'", "end": "'", "escape": "\\" }
    ],
    "rules": [
      { "kind": "indentStart", "start": ":" }
    ]
  }
}

This is a common Python pattern: after :, the following indented block is treated as one scope.

Notes:

  • start / end / branches are literal text markers, not regex patterns
  • Put comment and string delimiters in skips so scope markers inside them are ignored

bracketRules - Bracket Pair Rules

bracketRules defines literal bracket pairs for rainbow bracket rendering and partner lookup. It scans raw text directly and does not depend on highlight spans or style IDs.

{
  "bracketRules": {
    "inheritScopeSkips": true,
    "skips": [
      { "kind": "string", "start": "\"", "end": "\"", "escape": "\\" }
    ],
    "pairs": [
      { "start": "(", "end": ")" },
      { "start": "[", "end": "]" },
      { "start": "{", "end": "}" }
    ]
  }
}

Bracket Rule Fields

Field Type Required Description
pairs object[] Yes Literal bracket pairs. The array must not be empty.
inheritScopeSkips boolean No Whether to reuse scopeRules.skips; defaults to true.
skips object[] No Additional skip rules using the same schema as scopeRules.skips.

Each pair requires non-empty start and end strings. start and end must be different. Longer bracket or skip markers are matched first, so multi-character markers can coexist with shorter ones.

If a syntax has no scopeRules, define bracketRules.skips directly so brackets inside strings and comments are ignored. If a syntax already has complete scopeRules.skips, the common case is to omit skips and rely on the default inheritScopeSkips: true.


Inline Style Mode

In inline_style mode, style definitions are written directly in the JSON. Highlighting results include colors and font attributes directly, without the need for external style registration.

styles Definition

{
  "styles": [
    {
      "name": "keyword",
      "foreground": "#FF569CD6",
      "background": "#00000000",
      "tags": ["bold"]
    },
    {
      "name": "string",
      "foreground": "#FFBD63C5"
    },
    {
      "name": "comment",
      "foreground": "#FF60AE6F",
      "tags": ["italic"]
    }
  ]
}
Field Type Description
name string Style name, corresponding to the name referenced in patterns
foreground string Foreground color in #AARRGGBB (ARGB) format
background string Background color in #AARRGGBB (ARGB) format, optional
tags string[] Font attribute tags, optional. Supports "bold", "italic", "strikethrough"

Note: When using inline_style mode, set inline_style = true when creating the HighlightEngine.


Complete Example

Here is a complete simplified Java syntax rule example:

{
  "name": "java",
  "fileSuffixes": [".java"],
  "variables": {
    "identifier": "[a-zA-Z_$][\\w$]*",
    "whiteSpace": "[ \\t\\f]",
    "any": "[\\S\\s]",
    "primitiveType": "void|boolean|byte|char|short|int|long|float|double"
  },
  "states": {
    "default": [
      {
        "pattern": "\\b(class|interface|enum)\\b${whiteSpace}+(${identifier})",
        "styles": [1, "keyword", 2, "class"],
        "state": "classHeader"
      },
      {
        "pattern": "\\b(new)\\b${whiteSpace}+(${identifier})${whiteSpace}*(<)([^()]*)(>)${whiteSpace}*(\\()",
        "styles": [1, "keyword", 2, "class", 3, "punctuation", 5, "punctuation", 6, "punctuation"],
        "subStates": [4, "genericType"]
      },
      {
        "pattern": "\\b(new)\\b${whiteSpace}+(${identifier})${whiteSpace}*(\\()",
        "styles": [1, "keyword", 2, "class", 3, "punctuation"]
      },
      {
        "pattern": "\\b(public|private|protected|static|final|abstract|return|if|else|for|while)\\b",
        "styles": [1, "keyword"]
      },
      {
        "pattern": "\\b(true|false|null)\\b",
        "styles": [1, "builtin"]
      },
      {
        "pattern": "@${identifier}",
        "style": "annotation"
      },
      {
        "pattern": "\\b(${primitiveType})\\b${whiteSpace}+(${identifier})${whiteSpace}*(\\()",
        "styles": [1, "keyword", 2, "method", 3, "punctuation"]
      },
      {
        "pattern": "\\b(${primitiveType})\\b${whiteSpace}+(${identifier})${whiteSpace}*([;=,)])",
        "styles": [1, "keyword", 2, "variable", 3, "punctuation"]
      },
      {
        "pattern": "\\b(${identifier})${whiteSpace}*(<)([^()]*)(>)${whiteSpace}+(${identifier})${whiteSpace}*([;=,)])",
        "styles": [1, "class", 2, "punctuation", 4, "punctuation", 5, "variable", 6, "punctuation"],
        "subStates": [3, "genericType"]
      },
      {
        "pattern": "\\b(${identifier})${whiteSpace}+(${identifier})${whiteSpace}*([;=,)])",
        "styles": [1, "class", 2, "variable", 3, "punctuation"]
      },
      {
        "pattern": "(${identifier})${whiteSpace}*(\\()",
        "styles": [1, "method", 2, "punctuation"]
      },
      {
        "pattern": "\"(?:[^\"\\\\]|\\\\.)*\"",
        "style": "string"
      },
      {
        "pattern": "\\b[0-9][0-9_]*\\.?[0-9_]*(?:[eE][+-]?[0-9]+)?[fFdDlL]?\\b",
        "style": "number"
      },
      {
        "pattern": "/\\*",
        "style": "comment",
        "state": "blockComment"
      },
      {
        "pattern": "//${any}*",
        "style": "comment"
      },
      {
        "pattern": "[.()\\[\\]{}+\\-*/<>=!&|;:,?~^%@]",
        "style": "punctuation"
      }
    ],
    "classHeader": [
      {
        "pattern": "(<)([^()]*)(>)",
        "styles": [1, "punctuation", 3, "punctuation"],
        "subStates": [2, "genericType"]
      },
      {
        "pattern": "\\b(extends|implements)\\b",
        "styles": [1, "keyword"]
      },
      {
        "pattern": "${identifier}",
        "style": "class"
      },
      {
        "pattern": "[,:]",
        "style": "punctuation"
      },
      {
        "pattern": "\\{",
        "style": "punctuation",
        "state": "default"
      },
      { "onLineEndState": "classHeader" }
    ],
    "genericType": [
      {
        "pattern": "\\b(extends|super)\\b",
        "styles": [1, "keyword"]
      },
      {
        "pattern": "(${identifier})${whiteSpace}*(<)([^()]*)(>)",
        "styles": [1, "class", 2, "punctuation", 4, "punctuation"],
        "subStates": [3, "genericType"]
      },
      {
        "pattern": "${identifier}",
        "style": "class"
      },
      {
        "pattern": "[,\\[\\]?]",
        "style": "punctuation"
      }
    ],
    "blockComment": [
      {
        "pattern": "\\*/",
        "style": "comment",
        "state": "default"
      },
      {
        "pattern": "${any}",
        "style": "comment"
      }
    ]
  },
  "scopeRules": {
    "rules": [
      { "kind": "delimiter", "start": "{", "end": "}" }
    ]
  },
  "bracketRules": {
    "pairs": [
      { "start": "(", "end": ")" },
      { "start": "[", "end": "]" },
      { "start": "{", "end": "}" }
    ]
  }
}

Best Practices

1. Pattern Ordering

Rule ordering affects matching priority. The recommended order for default state rules:

  1. Class/struct declarations (triggers classHeader state transition)
  2. new expressions (constructor calls)
  3. import/using statements
  4. Variable declaration keywords (var, let, etc.)
  5. Preprocessor directives
  6. Built-in type + method/property/variable declarations
  7. General keywords
  8. Built-in constants (true/false/null)
  9. Annotations/decorators
  10. Generic type + method/property/variable declarations
  11. Simple type + method/property/variable declarations
  12. Method calls (fallback pattern)
  13. String literals (multi-line strings requiring state transitions)
  14. Numeric literals
  15. Comments (block comments require state transitions)
  16. Operators and punctuation (fallback pattern)

2. Avoid Greedy Matching Pitfalls

When handling generic parameters, use [^()]* instead of .* to prevent regex backtracking across too much content:

// Recommended
"(${identifier})${whiteSpace}*(<)([^()]*)(>)"

// Avoid
"(${identifier})${whiteSpace}*(<)(.*)(>)"

When the same line contains multiple independent <> pairs (e.g., class inheritance declarations), use non-greedy matching (.*?):

"(<)(.*?)(>)"

3. Use onLineEndState Appropriately

  • Multi-line structures (class declarations, function parameter lists) use onLineEndState to preserve state
  • States that only need to persist until line end don't need onLineEndState (defaults back to default)
  • Place onLineEndState at the end of the state's rule array

4. Use Variables to Reduce Repetition

Extract frequently used regex fragments as variables:

{
  "variables": {
    "identifier": "[a-zA-Z_]\\w*",
    "whiteSpace": "[ \\t\\f]",
    "builtinType": "int|float|double|string|bool|void"
  }
}

5. State Design Principles

  • Every state should have clear entry conditions and exit conditions
  • Multi-line syntax structures (strings, comments) must use state transitions
  • The last pattern in a state should be a fallback match (e.g., ${any}) to prevent the engine from getting stuck on certain characters
  • Avoid creating too many states — most languages can be handled with 5-10 states