# Why a Language Plugin Needs a Grammar, Not a Regex

A regex highlighter works right up until a comment shows up inside a string, and the actual fix is two generated parsers and two separate trees, not a smarter pattern.


A regex can highlight a language for about as long as nobody nests anything. The moment a string contains a comment marker, or a comment contains something that looks like a string, or a block comment nests inside another block comment the way Julia’s #= ... =# is allowed to, the pattern either lights up the wrong span or gives up entirely. Everyone’s first instinct for a language plugin is a regex, because a regex is the honest, cheap thing to reach for, and it is also the thing that guarantees you will rewrite the highlighter twice.
What a real language plugin does instead is generate two parsers from a grammar file and build two separate trees on top of them, and almost nobody explains why it needs to be two of anything.
Two generated files, not one On the IntelliJ Platform, the standard tool for this is Grammar-Kit. The workflow is: write a .bnf grammar describing the language’s rules, check it against the live preview as you go, generate the parser, the element types, and the PSI classes from it, then separately generate a JFlex lexer, then wire both into a ParserDefinition that tells the platform how to turn a file into tokens and a tree.
Flexible Julia, my own JetBrains plugin, bundles two of these, completely independent of each other. JuliaLexer.java and JuliaParser.java sit under src/main/gen for Julia itself. Next to them, unrelated, sit DyadLexer.java and DyadParser.java, generated from a second .bnf grammar for Dyad, JuliaHub’s separate modeling language, whose .dyad files the plugin also supports and which Dyad.jl compiles down to ordinary Julia. Two languages, two grammars, two generated lexers, two generated parsers, sharing nothing but a plugin.xml. That is what “language support” actually costs once you commit to doing it properly instead of pattern-matching your way through it.
Why there are two trees, not one Here is the part that took me longer to internalize than I expected. Parsing on the IntelliJ Platform is a two-step process. The generated parser first builds an AST: a raw, generic tree that mirrors exactly what the grammar rules matched, node by node, with no idea yet what any of it means. Then a second tree, the PSI, gets built on top of that AST, and the PSI is where every element actually knows what it is: this node is a function definition, this one is a macro call, this identifier resolves to that other declaration three files away.
The reason for the split is that the two trees are optimized for opposite things. The AST has to be cheap to produce, because it gets rebuilt on practically every keystroke, and it has no business knowing what a symbol resolves to. The PSI is the expensive, typed, IDE-facing layer: it is what completion, navigation, refactoring, and find-usages actually walk, and it is allowed to be slower and smarter because it is not rebuilt from scratch on every character you type. A regex highlighter conflates both jobs into one pass and does neither well. A grammar-based plugin keeps them apart on purpose, and that separation is the entire reason “go to definition” on a Julia function works instead of being a fuzzy string search wearing a keyboard shortcut.
What the grammar looks like, and what it costs you later A slice of the actual Julia grammar, the rule for what counts as an expression:
expr ::= compactFunction | constGlobalStatement | globalStatement | qualifiedMacroOp | applyMacroOp | assignLevel | arrowOp | ternaryOp | quoteLevel | lambda | ... | primaryExpr { implements=['com.ilscipio.language.julia.psi.impl.IJuliaExpr'] mixin='com.ilscipio.language.julia.psi.impl.JuliaExprMixin' } That reads almost like documentation of Julia’s own precedence table, because that is exactly what it is: every alternative is a level of the language’s grammar, ordered the way the reference parser orders them. Grammar-Kit turns this directly into recursive-descent parsing code.
And this is where the honest annoyance shows up. Right above that rule in my own grammar file sits a comment I left for myself: the generated JuliaParser.java has two manual performance edits in it, a lookahead guard on compactFunction and a token-type dispatch that skips alternatives the current token can’t possibly start, and both are hand-written directly into generated code that Grammar-Kit would happily overwrite the next time I regenerate. Code generation sells you the idea that the generated file is disposable, that you only ever touch the grammar. That stops being true the instant the generic parser is too slow somewhere and you go in and fix the generated code by hand, because now regenerating is not a free action anymore, it is a “did I remember to reapply the two edits I made in March” action.
I still think the grammar-and-two-trees approach is the only correct way to build a language plugin, for what it’s worth. I just no longer believe the part where codegen means you never have to look at the generated file again.
